Skip to main content
Aggregate arXiv cs.AI 人工智能 24 Aug 2026 - 13:00

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 20414v1 Announce Type: new Abstract: Vision-language models are increa…
  • Broad benchmarks often combine perception, optical character recogniti…
  • We introduce StateSight, a procedurally generated benchmark for cube-n…

摘要引擎:抽取

正文提要

arXiv:2608.20414v1 Announce Type: new Abstract: Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character recognition, domain knowledge, linguistic priors, and reasoning in the same evaluation. We introduce StateSight, a procedurally generated benchmark for cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting. Each task family contains 300 single-image prompts with deterministic oracle labels and exact-match scoring. OpenAI GPT-5.5, using the API model identifier gpt-5.5, achieved 59.3%, 33.3%, and 28.3% accuracy across the three tasks, while Claude Sonnet 5 achieved 53.3%, 18.7%, and 7.3%. All final direct runs had zero format errors. A 30-participant human baseline on 60 items exceeded both models on every task, with mean accuracies of 80.8%, 68.8%, and 64.3%. Visible-derivation analysis identified recurring errors in image-state reconstruction and reasoning procedure. We also introduce StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states. The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.

来源:https://arxiv.org/abs/2608.20414

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表