SPIN Unprocessed September 3, 2026 ai_technology research
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
View original on arxiv.orgOverview
arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart-
SpinGraph analysis pending — check back after processing.
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Artificial Intelligence
View all →- MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
- Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
- HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
- ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
- When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor
- Benchmarking Language Models for Statistical Problem Formulation
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO