Research Papers research paper arxiv computer-vision image-recognition

ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks

arXivby [Submitted on 29 May 2025 (v1), last revised 2 Apr 2026 (this version, v3)]April 3, 20262 min read1 views

arXiv:2505.23752v3 Announce Type: replace Abstract: Recent progress in large language models (LLMs) has enabled tool-augmented agents capable of solving complex real-world tasks through step-by-step reasoning. However, existing evaluations often focus on general-purpose or multimodal scenarios, leaving a gap in domain-specific benchmarks that assess tool-use capabilities in complex remote sensing use cases. We present ThinkGeo, an agentic benchmark designed to evaluate LLM-driven agents on remote sensing tasks via structured tool use and multi-step planning. Inspired by tool-interaction paradi — Akashah Shabbir, Muhammad Akhtar Munir, Akshay Dudhane, Muhammad Umer Sheikh, Muhammad Haris Khan, Paolo Fraccaro, Juan Bernabe Moreno, Fahad Shahbaz Khan, Salman Khan

View PDF HTML (experimental)

Abstract:Recent progress in large language models (LLMs) has enabled tool-augmented agents capable of solving complex real-world tasks through step-by-step reasoning. However, existing evaluations often focus on general-purpose or multimodal scenarios, leaving a gap in domain-specific benchmarks that assess tool-use capabilities in complex remote sensing use cases. We present ThinkGeo, an agentic benchmark designed to evaluate LLM-driven agents on remote sensing tasks via structured tool use and multi-step planning. Inspired by tool-interaction paradigms, ThinkGeo includes human-curated queries spanning a wide range of real-world applications such as urban planning, disaster assessment and change analysis, environmental monitoring, transportation analysis, aviation monitoring, recreational infrastructure, and industrial site analysis. Queries are grounded in satellite or aerial imagery, including both optical RGB and SAR data, and require agents to reason through a diverse toolset. We implement a ReAct-style interaction loop and evaluate both open and closed-source LLMs (e.g., GPT-4o, Qwen2.5) on 486 structured agentic tasks with 1,778 expert-verified reasoning steps. The benchmark reports both step-wise execution metrics and final answer correctness. Our analysis reveals notable disparities in tool accuracy and planning consistency across models. ThinkGeo provides the first extensive testbed for evaluating how tool-enabled LLMs handle spatial reasoning in remote sensing.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2505.23752 [cs.CV]

(or arXiv:2505.23752v3 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2505.23752

arXiv-issued DOI via DataCite

Submission history

From: Akashah Shabbir [view email] [v1] Thu, 29 May 2025 17:59:38 UTC (5,006 KB) [v2] Thu, 9 Oct 2025 17:29:59 UTC (8,724 KB) [v3] Thu, 2 Apr 2026 04:47:22 UTC (6,298 KB)

Original source

arXiv

https://arxiv.org/abs/2505.23752

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Building knowledge graph…

Discussion

No comments yet — be the first to share your thoughts!