Skip to main content
Aggregate AI 摘要 arXiv cs.AI 人工智能 31 Aug 2026 - 13:00

Thinking Costs Tokens: When More Structure is Worth the Price

RSS 官方收录 · 可信分层展示

关键摘要

金融推理任务中,验证式搜索架构在1500+ token预算时准确率超单次调用模型

  • 验证式搜索架构在1500+ token预算下准确率(44%)持续高于单次调用(40%)
  • 1000 token时单次调用达18%,验证式因规划开销几乎为0%
  • 两系统在250–42000 token共14档预算下测试,交叉点经统计检验p≤0.001

AI 摘要 · 来源可核验

正文提要

arXiv:2608.27506v1 Announce Type: new Abstract: Adding inference structure to a language model lets it search, verify, and revise, but these actions consume the very budget they are supposed to use well. In this paper, we investigate whether there exists a token-budget threshold, below which the overhead of planning and verification hurts performance and above which it helps. We evaluate two systems on FinQA and TAT-QA financial reasoning tasks, using GPT-5.4 mini across 14 budget tiers ranging from 250 to 42,000 output-equivalent tokens. The first system is a monolith, which is a single LLM call. The second is a verified search architecture that adds planning, label-blind checking, and repair capabilities. We run 1,000 cases for a total of 28,000 completed cells. Both systems score 0% at the two lowest tiers, where neither can fit a complete prompt. At 1,000 tokens, the monolith reaches 18% accuracy while verified search scores near 0%, since the planning overhead leaves no room for an answer. From 1,500 tokens onward, verified search surpasses the monolith and maintains a consistent advantage, reaching approximately 44% at the highest tiers while the monolith reaches approximately 40%. The crossover occurs between 1,000 and 1,500 output-equivalent tokens, confirmed by a strict intersection-union test ($p \le 0.001$ at both endpoints).

来源:https://arxiv.org/abs/2608.27506

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表