Skip to main content
Aggregate arXiv cs.AI 人工智能 20 Aug 2026 - 14:30

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 18303v1 Announce Type: new Abstract: LLM-as-judge evaluation reduces r…
  • We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a traini…
  • On RewardBench (n=1,000), SESSE achieves near-parity with the chain-of…

摘要引擎:抽取

正文提要

arXiv:2608.18303v1 Announce Type: new Abstract: LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate which quality dimensions drove the preference or distinguish model errors from genuine label ambiguity. We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework that decomposes holistic judgment into structured sub-questions mined directly from the judge's own error cases; requiring no oracle responses, task-specific rubrics, or fine-tuning. On RewardBench (n=1,000), SESSE achieves near-parity with the chain-of-thought baseline and is competitive with RISE-Judge-32B (92.7%), a fine-tuned specialist, while remaining fully training-free. Per-criterion vote evidence provides an interpretable audit trail for diagnosing label ambiguity and judge failure modes unavailable from a single holistic output token.

来源:https://arxiv.org/abs/2608.18303

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表