Skip to main content
Aggregate arXiv cs.AI 人工智能 19 Aug 2026 - 14:00

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 17330v1 Announce Type: new Abstract: Large language models for medical…
  • We evaluated three API models across four physician-authored, multi-tu…
  • Self-care or home-management advice before any patient answer appeared…

摘要引擎:抽取

正文提要

arXiv:2608.17330v1 Announce Type: new Abstract: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.

来源:https://arxiv.org/abs/2608.17330

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表