微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.…
- 25097v1 Announce Type: new Abstract: Understanding how (multimodal) la…
- Existing physics benchmarks remain limited in the following two import…
- As a result, model performance on current datasets may not be fully re…
摘要引擎:抽取
正文提要
arXiv:2608.25097v1 Announce Type: new Abstract: Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their ability to solve complex physics problems. To address these issues, we present PhysElite, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning. PhysElite contains 11,586 Olympiad-tier problems. For each problem, we provide corresponding visual diagrams, step-by-step bilingual Chinese-English solution derivations, and the final answer. We benchmark 18 open-source and closed-source MLLMs, and find that even the strongest model reaches only 33.7% answer accuracy. We additionally conduct step-level process evaluation to diagnose where models fail in the reasoning chain. Our datasets are released at https://huggingface.co/datasets/physelite/PhysElite.