Geospatial Vision-Language Models for Spatial Reasoning and Temporal Change Understanding: A Task-Centered Benchmarking Framework and Evidence Synthesis Discussion
DOI:
https://doi.org/10.58776/ijitcsa.v4i2.262Keywords:
Temporal change understanding, Remote sensing, Urban intelligence, Change captioning, Multimodal benchmarkAbstract
Geospatial vision-language models (VLMs) are increasingly expected to do more than assign scene labels or generate generic captions. In realistic Earth-observation and urban intelligence settings, useful multimodal systems must support fine-grained spatial reasoning, cross-view interpretation, grounded localization, and explicit understanding of change across time. Yet the current literature remains fragmented across remote sensing visual question answering, visual grounding, urban multi-view reasoning, and bi-temporal change captioning. As a result, claims about progress are often task-local, benchmark-specific, and difficult to compare. This paper reconstructs the field around a more defensible technical center: geospatial multimodal intelligence as the joint problem of spatial reasoning and temporal change understanding. Rather than presenting unverifiable new benchmark runs, we develop a rigorous review paper anchored in public benchmark evidence and a reference architecture for reproducible future work. We first formalize a task-centered problem definition that unifies image-level, region-level, cross-view, and bi-temporal reasoning. We then propose a reference GST-VLM architecture consisting of spatial encoding, temporal difference modeling, multimodal fusion, task-specific decoding, and reliability estimation. Next, we synthesize publicly reported evidence from representative datasets and benchmarks including RSVQA, EarthVQA, VRSBench, GeoChat, LEVIR-CD, LEVIR-CC, SECOND-CC, CHOICE, GEOBench-VLM, CityBench, and UrBench. The synthesis shows that recent models are improving rapidly but remain far from robust geospatial reasoning systems: on GEOBench-VLM, the best public model reported only 41.7% multiple-choice accuracy; on UrBench, even GPT-4o still trails human performance by an average 17.4 percentage points; and while specialized systems such as GeoReasoner, GeoChat, GeoLLaVA, and MModalCC outperform generic baselines on targeted tasks, their gains remain strongly benchmark-dependent. Based on this evidence, we identify the principal bottlenecks as benchmark fragmentation, weak temporal grounding, inadequate calibration, scarce cross-region validation, limited deployment reporting, and insufficient integration of geometry with language-conditioned reasoning. The paper concludes with a concrete research agenda for trustworthy geospatial VLMs that is centered on multi-temporal supervision, interactive change analysis, uncertainty-aware outputs, and evaluation protocols that measure not only accuracy but also transfer, calibration, and operational feasibility.
References
. A. Radford et al., “Learning transferable visual models from natural language supervision,” arXiv preprint arXiv:2103.00020, 2021, doi: 10.48550/arXiv.2103.00020.
. J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” arXiv preprint arXiv:2301.12597, 2023, doi: 10.48550/arXiv.2301.12597.
. S. Liu et al., “Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection,” arXiv preprint arXiv:2303.05499, 2023, doi: 10.48550/arXiv.2303.05499.
. B. Xiao et al., “Florence-2: Advancing a unified representation for a variety of vision tasks,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4818–4829, 2024.
. G. Team et al., “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” arXiv preprint arXiv:2403.05530, 2024, doi: 10.48550/arXiv.2403.05530.
. OpenAI et al., “GPT-4 technical report,” arXiv preprint arXiv:2303.08774, 2023, doi: 10.48550/arXiv.2303.08774.
. A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, et al., “Qwen2 technical report,” arXiv preprint arXiv:2407.10671, 2024, doi: 10.48550/arXiv.2407.10671.
. B. Li et al., “LLaVA-OneVision: Easy visual task transfer,” arXiv preprint arXiv:2408.03326, 2024, doi: 10.48550/arXiv.2408.03326.
. X. Li, C. Wen, Y. Hu, Z. Yuan, and X. X. Zhu, “Vision-language models in remote sensing: Current progress and future trends,” IEEE Geoscience and Remote Sensing Magazine, vol. 12, no. 2, pp. 32–66, 2024, doi: 10.1109/MGRS.2024.3383473.
. X. Zhou et al., “Vision language models in autonomous driving: A survey and outlook,” IEEE Transactions on Intelligent Vehicles, pp. 1–20, 2024, doi: 10.1109/TIV.2024.3402136.
. S. Lobry, D. Marcos, J. Murray, and D. Tuia, “RSVQA: Visual question answering for remote sensing data,” IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 12, pp. 8555–8566, 2020, doi: 10.1109/TGRS.2020.2988782.
. C. Chappuis, S. Lobry, and D. Tuia, “Prompt-RSVQA: Prompting visual context to a language model for remote sensing visual question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2022, pp. 2505–2515.
. J. Wang, Z. Zheng, Z. Chen, A. Ma, and Y. Zhong, “EarthVQA: Towards queryable earth via relational reasoning-based remote sensing visual question answering,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 6, pp. 5481–5489, 2024, doi: 10.1609/aaai.v38i6.28357.
. J. Chen, R. Zhao, H. Chen, Z. Zou, and Z. Shi, “Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–20, 2022, doi: 10.1109/TGRS.2022.3218921.
. C. Liu, K. Chen, B. Chen, H. Zhang, Z. Zou, and Z. Shi, “RSCaMa: Remote sensing image change captioning with state space model,” IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024.
. Y. Yang et al., “Remote sensing image change captioning using multi-time-step features of diffusion models,” Remote Sensing, vol. 16, no. 21, p. 4083, 2024, doi: 10.3390/rs16214083.
. A. C. Karaca, E. Ozelbas, S. Berber, O. Karimli, T. Yildirim, and M. F. Amasyali, “Robust change captioning in remote sensing: SECOND-CC dataset and MModalCC framework,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 21494–21513, 2025, doi: 10.1109/JSTARS.2025.3600613.
. K. Kuckreja, M. S. Danish, M. Naseer, A. Das, S. Khan, and F. S. Khan, “GeoChat: Grounded large vision-language model for remote sensing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 27831–27840.
. X. Wang, Y. Hu, et al., “RingMoGPT: A unified remote sensing foundation model for vision, language, and grounded tasks,” IEEE Transactions on Geoscience and Remote Sensing, 2024, doi: 10.1109/TGRS.2024.3510833.
. Y. Zhou et al., “GeoGround: A unified large vision-language model for remote sensing visual grounding,” arXiv preprint arXiv:2411.11904, 2024, doi: 10.48550/arXiv.2411.11904.
. M. S. Danish et al., “GEOBench-VLM: Benchmarking vision-language models for geospatial tasks,” in Proceedings of the IEEE/CVF international conference on computer vision, 2025. Available: https://openaccess.thecvf.com/content/ICCV2025/html/Danish_GEOBench-VLM_Benchmarking_Vision-Language_Models_for_Geospatial_Tasks_ICCV_2025_paper.html.
. X. An, J. Sun, Z. Gui, and W. He, “CHOICE: Benchmarking the remote sensing capabilities of large vision-language models,” arXiv preprint arXiv:2411.18145, 2025, doi: 10.48550/arXiv.2411.18145.
. B. Zhou et al., “UrBench: A comprehensive benchmark for evaluating large multimodal models in multi-view urban scenarios,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 10, p. 33163, 2025, doi: 10.1609/aaai.v39i10.33163.
. X. Li, J. Ding, M. Elhoseiny, et al., “VRSBench: A versatile vision-language benchmark dataset for remote sensing image understanding,” Advances in Neural Information Processing Systems, vol. 37, 2024, doi: 10.52202/079017-0106.
. P. Deng, W. Zhou, and H. Wu, “DeltaVLM: Interactive remote sensing image change analysis via instruction-guided difference perception,” arXiv preprint arXiv:2507.22346, 2025, doi: 10.48550/arXiv.2507.22346.
. H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,” Remote Sensing, vol. 12, no. 10, p. 1662, 2020, doi: 10.3390/rs12101662.
. M. Wu, Q. Huang, S. Gao, and Z. Zhang, “Mixed land use measurement and mapping with street view images and spatial context-aware prompts via zero-shot multimodal learning,” International Journal of Applied Earth Observation and Geoinformation, vol. 125, p. 103591, 2023, doi: 10.1016/j.jag.2023.103591.
. Y. Kang, J. Kim, J. Park, and J. Lee, “Assessment of perceived and physical walkability using street view images and deep learning technology,” ISPRS International Journal of Geo-Information, vol. 12, no. 5, p. 186, 2023, doi: 10.3390/ijgi12050186.
. W. Huang, J. Wang, and C. Gao, “Zero-shot urban function inference with street view images through prompting a pretrained vision-language model,” International Journal of Geographical Information Science, vol. 38, no. 7, pp. 1414–1442, 2024, doi: 10.1080/13658816.2024.2347322.
. L. Li, Y. Ye, B. Jiang, and W. Zeng, “GeoReasoner: Geo-localization with reasoning in street views using a large vision-language model,” arXiv preprint arXiv:2406.18572, 2024, doi: 10.48550/arXiv.2406.18572.
. H. Elgendy, A. Sharshar, A. Aboeitta, Y. Ashraf, and M. Guizani, “GeoLLaVA: Efficient fine-tuned vision-language models for temporal change detection in remote sensing,” arXiv preprint arXiv:2410.19552, 2024, doi: 10.48550/arXiv.2410.19552.
. Y. Zhu et al., “Semantic-CC: Boosting remote sensing image change captioning with foundational knowledge and semantic guidance,” arXiv preprint arXiv:2407.14032, 2024, doi: 10.48550/arXiv.2407.14032.
. J. Feng et al., “CityBench: Evaluating the capabilities of large language models for urban tasks,” arXiv preprint arXiv:2406.13945, 2024, doi: 10.48550/arXiv.2406.13945.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Abeer Hasshen Abdullah

This work is licensed under a Creative Commons Attribution 4.0 International License.
Attribution 4.0 International
You are free to:
- Share — copy and redistribute the material in any medium or format for any purpose, even commercially.
- Adapt — remix, transform, and build upon the material for any purpose, even commercially.
- The licensor cannot revoke these freedoms as long as you follow the license terms.
Under the following terms:
- Attribution — You must give appropriate credit , provide a link to the license, and indicate if changes were made . You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
- No additional restrictions — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits.
Notices:
You do not have to comply with the license for elements of the material in the public domain or where your use is permitted by an applicable exception or limitation .
No warranties are given. The license may not give you all of the permissions necessary for your intended use. For example, other rights such as publicity, privacy, or moral rights may limit how you use the material.


