微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2609.…
- 02059v1 Announce Type: new Abstract: Multimodal Large Language Models …
- However, existing benchmarks typically evaluate these domains in isola…
- We introduce DocHop, a benchmark for integrated chart--context reasoni…
摘要引擎:抽取
正文提要
arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.