Skip to main content
Aggregate arXiv cs.AI 人工智能 3 Sep 2026 - 13:30

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2609.…

  • 02059v1 Announce Type: new Abstract: Multimodal Large Language Models …
  • However, existing benchmarks typically evaluate these domains in isola…
  • We introduce DocHop, a benchmark for integrated chart--context reasoni…

摘要引擎:抽取

正文提要

arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.

来源:https://arxiv.org/abs/2609.02059

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表