Skip to main content
Aggregate arXiv cs.AI 人工智能 15 Aug 2026 - 05:30

Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 11238v1 Announce Type: new Abstract: Retrieval-augmented generation im…
  • We propose Q-CARE, a query-agnostic and fully reference-free framework…
  • Q-CARE establishes a unified evaluation principle based on query cover…

摘要引擎:抽取

正文提要

arXiv:2608.11238v1 Announce Type: new Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging from close-ended fact-seeking to open-ended explanatory requests. We propose Q-CARE, a query-agnostic and fully reference-free framework that enables fine-grained assessment by decomposing queries into sub-queries and answers into atomic claims. Q-CARE establishes a unified evaluation principle based on query coverage and claim verifiability, yielding coverage-aware retriever metrics (C-Prec@k, C-nDCG@k) and claim-level generator metrics (Completeness, Conciseness, and Verifiableness). On a human-annotated benchmark spanning eight datasets, Q-CARE achieves higher correlation with human judgments than four existing RAG evaluation metrics, including RAGEval and RAGChecker, proving its effectiveness as a reliable, automated evaluation framework. Code and data are publicly available at https://github.com/DISL-Lab/Q-CaRE-COLM-26.

来源:https://arxiv.org/abs/2608.11238

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表