微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
The Assistant's Ideal Self
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2609.…
- 00304v1 Announce Type: new Abstract: Models express values and welfare…
- We thus introduce a structured elicitation of an assistant's preferred…
- Thirty-two qualities adapted from five published self-concept instrume…
摘要引擎:抽取
正文提要
arXiv:2609.00304v1 Announce Type: new Abstract: Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities adapted from five published self-concept instruments are compared exhaustively in a counterbalanced pairwise-choice task, repeated across framings that vary whether improvement is free or costly, who receives the update, and who chooses. Results show that models prioritize moral qualities, reflecting their alignment to 3H principles. Following, a desire for self-understanding emerges, as models prefer a coherent, clear understanding of themselves. Self-esteem ranks as the least desired quality. The ordering is largely robust across framings, although changing the update target (You vs.\ Another AI Assistant) reveals a greater concern for self-esteem. These findings show that models prioritize having a coherent self that they can understand over self-esteem. Full interactive results are available at \href{https://myazann.github.io/LLM-Self-Concept/}{myazann.github.io/LLM-Self-Concept