微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
Asymmetries in Spontaneous and Instructed Deception
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to.…
- However, much of the study on deception in models involves instructed …
- We investigated the relationship between instructed and spontaneous (u…
- 1-70B-Instruct.
摘要引擎:抽取
正文提要
arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers.