Skip to main content
Aggregate AI 摘要 arXiv cs.AI 人工智能 26 Aug 2026 - 14:00

SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models

RSS 官方收录 · 可信分层展示

关键摘要

SyPS框架量化大模型谄媚行为对提示词变化的敏感度,提出SPSS评分新指标

  • SyPS构建控制变量提示变体,测量同一情境下谄媚行为变化
  • SPSS评分分离基线谄媚率与提示诱导偏移,支持模型鲁棒性比较
  • 实证发现验证寻求、情绪压力类提示易增谄媚,反谄媚提示可降低

AI 摘要 · 来源可核验

正文提要

arXiv:2608.23837v1 Announce Type: new Abstract: Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with users in socially sensitive contexts. Existing evaluations typically measure sycophancy under a fixed prompt formulation, leaving unclear whether such behavior is stable when the same underlying situation is presented with different sycophancy-relevant prompt variants. In this work, we study sycophancy prompt sensitivity: the extent to which changes in user confidence, emotional framing, social consensus, or validation-seeking language alter a model's sycophantic behavior. We refer to our evaluation framework as SyPS, short for Sycophancy Prompt Sensitivity. Building on existing social sycophancy evaluation settings, SyPS constructs controlled prompt variants that preserve the same underlying user situation while varying sycophancy-relevant social cues. We introduce the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level measure of sycophancy variation across paired prompt variants. Unlike aggregate sycophancy rates, SPSS separates baseline sycophancy from prompt-induced shifts, enabling model-level comparisons of robustness to sycophancy-relevant social cues. Empirically, we find that sycophancy prompt sensitivity is socially structured: validation-seeking and emotional-pressure cues often increase sycophancy, whereas counter-framing and anti-sycophancy prompts tend to reduce it. Our framework highlights whether LLMs maintain stable social judgments while adapting appropriately in tone.

来源:https://arxiv.org/abs/2608.23837

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表