微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.…
- 18303v1 Announce Type: new Abstract: LLM-as-judge evaluation reduces r…
- We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a traini…
- On RewardBench (n=1,000), SESSE achieves near-parity with the chain-of…
摘要引擎:抽取
正文提要
arXiv:2608.18303v1 Announce Type: new Abstract: LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate which quality dimensions drove the preference or distinguish model errors from genuine label ambiguity. We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework that decomposes holistic judgment into structured sub-questions mined directly from the judge's own error cases; requiring no oracle responses, task-specific rubrics, or fine-tuning. On RewardBench (n=1,000), SESSE achieves near-parity with the chain-of-thought baseline and is competitive with RISE-Judge-32B (92.7%), a fine-tuned specialist, while remaining fully training-free. Per-criterion vote evidence provides an interpretable audit trail for diagnosing label ambiguity and judge failure modes unavailable from a single holistic output token.