Skip to main content
Aggregate AI 摘要 arXiv cs.AI 人工智能 7 Sep 2026 - 14:00

Extremely Sparse Supervision Incentivizes Reasoning Ability

RSS 官方收录 · 可信分层展示

关键摘要

仅用0.05%生成token监督,Qwen3模型推理能力不降反升

  • 极稀疏监督(每推理轨迹仅1-2个token)提升推理能力
  • 在9种师生配置及数学、编程、PPO-RLVR任务中均有效
  • 稀疏监督更贴近人类反思关键步骤的自然学习过程

AI 摘要 · 来源可核验

正文提要

arXiv:2609.04565v1 Announce Type: new Abstract: Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. We revisit this assumption in the on-policy distillation (OPD) setting, which naturally admits dense teacher supervision at every generated token. Using the Qwen3 family, we discover a counter-intuitive phenomenon: reasoning can be effectively incentivized by an extremely small fraction of generated tokens--as few as one or two tokens per reasoning trajectory, corresponding to only 0.05% of all tokens. Surprisingly, this sparse supervision in most cases matches or surpasses full-token training in improving reasoning ability, despite excluding the vast majority of generated tokens from the training objective. This phenomenon is consistently observed across nine teacher--student configurations spanning different model scales on mathematical reasoning tasks, and is further validated on coding reasoning, Llama models and Proximal Policy Optimization (PPO)-based reinforcement learning with verifiable reward (RLVR). Interestingly, such extremely sparse supervision may be closer to the natural learning process: rather than correcting every step word by word, one reflects on a few critical reasoning steps, updates prior understanding, and continues the trial-and-error, avoiding micro-level corrections while remaining remarkably effective. Overall, our results challenge the assumption that effective post-training must be token-intensive and point to a new direction for understanding and designing more efficient post-training algorithms.

来源:https://arxiv.org/abs/2609.04565

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表