Geospatial Vision-Language Models for Spatial Reasoning and Temporal Change Understanding: A Task-Centered Benchmarking Framework and Evidence Synthesis Discussion

Authors

  • Abeer Hasshen Abdullah Dawood University of Engineering and Technology

DOI:

https://doi.org/10.58776/ijitcsa.v4i2.262

Keywords:

Temporal change understanding, Remote sensing, Urban intelligence, Change captioning, Multimodal benchmark

Abstract

Geospatial vision-language models (VLMs) are increasingly expected to do more than assign scene labels or generate generic captions. In realistic Earth-observation and urban intelligence settings, useful multimodal systems must support fine-grained spatial reasoning, cross-view interpretation, grounded localization, and explicit understanding of change across time. Yet the current literature remains fragmented across remote sensing visual question answering, visual grounding, urban multi-view reasoning, and bi-temporal change captioning. As a result, claims about progress are often task-local, benchmark-specific, and difficult to compare. This paper reconstructs the field around a more defensible technical center: geospatial multimodal intelligence as the joint problem of spatial reasoning and temporal change understanding. Rather than presenting unverifiable new benchmark runs, we develop a rigorous review paper anchored in public benchmark evidence and a reference architecture for reproducible future work. We first formalize a task-centered problem definition that unifies image-level, region-level, cross-view, and bi-temporal reasoning. We then propose a reference GST-VLM architecture consisting of spatial encoding, temporal difference modeling, multimodal fusion, task-specific decoding, and reliability estimation. Next, we synthesize publicly reported evidence from representative datasets and benchmarks including RSVQA, EarthVQA, VRSBench, GeoChat, LEVIR-CD, LEVIR-CC, SECOND-CC, CHOICE, GEOBench-VLM, CityBench, and UrBench. The synthesis shows that recent models are improving rapidly but remain far from robust geospatial reasoning systems: on GEOBench-VLM, the best public model reported only 41.7% multiple-choice accuracy; on UrBench, even GPT-4o still trails human performance by an average 17.4 percentage points; and while specialized systems such as GeoReasoner, GeoChat, GeoLLaVA, and MModalCC outperform generic baselines on targeted tasks, their gains remain strongly benchmark-dependent. Based on this evidence, we identify the principal bottlenecks as benchmark fragmentation, weak temporal grounding, inadequate calibration, scarce cross-region validation, limited deployment reporting, and insufficient integration of geometry with language-conditioned reasoning. The paper concludes with a concrete research agenda for trustworthy geospatial VLMs that is centered on multi-temporal supervision, interactive change analysis, uncertainty-aware outputs, and evaluation protocols that measure not only accuracy but also transfer, calibration, and operational feasibility.

References

. A. Radford et al., “Learning transferable visual models from natural language supervision,” arXiv preprint arXiv:2103.00020, 2021, doi: 10.48550/arXiv.2103.00020.

. J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” arXiv preprint arXiv:2301.12597, 2023, doi: 10.48550/arXiv.2301.12597.

. S. Liu et al., “Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection,” arXiv preprint arXiv:2303.05499, 2023, doi: 10.48550/arXiv.2303.05499.

. B. Xiao et al., “Florence-2: Advancing a unified representation for a variety of vision tasks,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4818–4829, 2024.

. G. Team et al., “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” arXiv preprint arXiv:2403.05530, 2024, doi: 10.48550/arXiv.2403.05530.

. OpenAI et al., “GPT-4 technical report,” arXiv preprint arXiv:2303.08774, 2023, doi: 10.48550/arXiv.2303.08774.

. A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, et al., “Qwen2 technical report,” arXiv preprint arXiv:2407.10671, 2024, doi: 10.48550/arXiv.2407.10671.

. B. Li et al., “LLaVA-OneVision: Easy visual task transfer,” arXiv preprint arXiv:2408.03326, 2024, doi: 10.48550/arXiv.2408.03326.

. X. Li, C. Wen, Y. Hu, Z. Yuan, and X. X. Zhu, “Vision-language models in remote sensing: Current progress and future trends,” IEEE Geoscience and Remote Sensing Magazine, vol. 12, no. 2, pp. 32–66, 2024, doi: 10.1109/MGRS.2024.3383473.

. X. Zhou et al., “Vision language models in autonomous driving: A survey and outlook,” IEEE Transactions on Intelligent Vehicles, pp. 1–20, 2024, doi: 10.1109/TIV.2024.3402136.

. S. Lobry, D. Marcos, J. Murray, and D. Tuia, “RSVQA: Visual question answering for remote sensing data,” IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 12, pp. 8555–8566, 2020, doi: 10.1109/TGRS.2020.2988782.

. C. Chappuis, S. Lobry, and D. Tuia, “Prompt-RSVQA: Prompting visual context to a language model for remote sensing visual question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2022, pp. 2505–2515.

. J. Wang, Z. Zheng, Z. Chen, A. Ma, and Y. Zhong, “EarthVQA: Towards queryable earth via relational reasoning-based remote sensing visual question answering,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 6, pp. 5481–5489, 2024, doi: 10.1609/aaai.v38i6.28357.

. J. Chen, R. Zhao, H. Chen, Z. Zou, and Z. Shi, “Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–20, 2022, doi: 10.1109/TGRS.2022.3218921.

. C. Liu, K. Chen, B. Chen, H. Zhang, Z. Zou, and Z. Shi, “RSCaMa: Remote sensing image change captioning with state space model,” IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024.

. Y. Yang et al., “Remote sensing image change captioning using multi-time-step features of diffusion models,” Remote Sensing, vol. 16, no. 21, p. 4083, 2024, doi: 10.3390/rs16214083.

. A. C. Karaca, E. Ozelbas, S. Berber, O. Karimli, T. Yildirim, and M. F. Amasyali, “Robust change captioning in remote sensing: SECOND-CC dataset and MModalCC framework,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 21494–21513, 2025, doi: 10.1109/JSTARS.2025.3600613.

. K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan, “GeoChat: Grounded large vision-language model for remote sensing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 27831–27840.

. X. Wang, Y. Hu, et al., “RingMoGPT: A unified remote sensing foundation model for vision, language, and grounded tasks,” IEEE Transactions on Geoscience and Remote Sensing, 2024, doi: 10.1109/TGRS.2024.3510833.

. Y. Zhou et al., “GeoGround: A unified large vision-language model for remote sensing visual grounding,” arXiv preprint arXiv:2411.11904, 2024, doi: 10.48550/arXiv.2411.11904.

. M. S. Danish et al., “GEOBench-VLM: Benchmarking vision-language models for geospatial tasks,” in Proceedings of the IEEE/CVF international conference on computer vision, 2025. Available: https://openaccess.thecvf.com/content/ICCV2025/html/Danish_GEOBench-VLM_Benchmarking_Vision-Language_Models_for_Geospatial_Tasks_ICCV_2025_paper.html.

. X. An, J. Sun, Z. Gui, and W. He, “CHOICE: Benchmarking the remote sensing capabilities of large vision-language models,” arXiv preprint arXiv:2411.18145, 2025, doi: 10.48550/arXiv.2411.18145.

. B. Zhou et al., “UrBench: A comprehensive benchmark for evaluating large multimodal models in multi-view urban scenarios,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 10, p. 33163, 2025, doi: 10.1609/aaai.v39i10.33163.

. X. Li, J. Ding, M. Elhoseiny, et al., “VRSBench: A versatile vision-language benchmark dataset for remote sensing image understanding,” Advances in Neural Information Processing Systems, vol. 37, 2024, doi: 10.52202/079017-0106.

. P. Deng, W. Zhou, and H. Wu, “DeltaVLM: Interactive remote sensing image change analysis via instruction-guided difference perception,” arXiv preprint arXiv:2507.22346, 2025, doi: 10.48550/arXiv.2507.22346.

. H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,” Remote Sensing, vol. 12, no. 10, p. 1662, 2020, doi: 10.3390/rs12101662.

. M. Wu, Q. Huang, S. Gao, and Z. Zhang, “Mixed land use measurement and mapping with street view images and spatial context-aware prompts via zero-shot multimodal learning,” International Journal of Applied Earth Observation and Geoinformation, vol. 125, p. 103591, 2023, doi: 10.1016/j.jag.2023.103591.

. Y. Kang, J. Kim, J. Park, and J. Lee, “Assessment of perceived and physical walkability using street view images and deep learning technology,” ISPRS International Journal of Geo-Information, vol. 12, no. 5, p. 186, 2023, doi: 10.3390/ijgi12050186.

. W. Huang, J. Wang, and C. Gao, “Zero-shot urban function inference with street view images through prompting a pretrained vision-language model,” International Journal of Geographical Information Science, vol. 38, no. 7, pp. 1414–1442, 2024, doi: 10.1080/13658816.2024.2347322.

. L. Li, Y. Ye, B. Jiang, and W. Zeng, “GeoReasoner: Geo-localization with reasoning in street views using a large vision-language model,” arXiv preprint arXiv:2406.18572, 2024, doi: 10.48550/arXiv.2406.18572.

. H. Elgendy, A. Sharshar, A. Aboeitta, Y. Ashraf, and M. Guizani, “GeoLLaVA: Efficient fine-tuned vision-language models for temporal change detection in remote sensing,” arXiv preprint arXiv:2410.19552, 2024, doi: 10.48550/arXiv.2410.19552.

. Y. Zhu et al., “Semantic-CC: Boosting remote sensing image change captioning with foundational knowledge and semantic guidance,” arXiv preprint arXiv:2407.14032, 2024, doi: 10.48550/arXiv.2407.14032.

. J. Feng et al., “CityBench: Evaluating the capabilities of large language models for urban tasks,” arXiv preprint arXiv:2406.13945, 2024, doi: 10.48550/arXiv.2406.13945.

Downloads

Published

01-08-2026

How to Cite

Abdullah, A. H. (2026). Geospatial Vision-Language Models for Spatial Reasoning and Temporal Change Understanding: A Task-Centered Benchmarking Framework and Evidence Synthesis Discussion. International Journal of Information Technology and Computer Science Applications, 4(2), 175–184. https://doi.org/10.58776/ijitcsa.v4i2.262

Issue

Section

New Submission