Skip to main content
Aggregate arXiv cs.AI 人工智能 28 Aug 2026 - 11:30

PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 25097v1 Announce Type: new Abstract: Understanding how (multimodal) la…
  • Existing physics benchmarks remain limited in the following two import…
  • As a result, model performance on current datasets may not be fully re…

摘要引擎:抽取

正文提要

arXiv:2608.25097v1 Announce Type: new Abstract: Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their ability to solve complex physics problems. To address these issues, we present PhysElite, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning. PhysElite contains 11,586 Olympiad-tier problems. For each problem, we provide corresponding visual diagrams, step-by-step bilingual Chinese-English solution derivations, and the final answer. We benchmark 18 open-source and closed-source MLLMs, and find that even the strongest model reaches only 33.7% answer accuracy. We additionally conduct step-level process evaluation to diagnose where models fail in the reasoning chain. Our datasets are released at https://huggingface.co/datasets/physelite/PhysElite.

来源:https://arxiv.org/abs/2608.25097

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表