微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
AI Alignment through a Game-theoretic Lens: A Survey
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.…
- 27910v1 Announce Type: new Abstract: As large language models and incr…
- Existing alignment methods, while effective in improving helpfulness, …
- This survey reviews AI alignment through a game-theoretic lens.
摘要引擎:抽取
正文提要
arXiv:2608.27910v1 Announce Type: new Abstract: As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, and controllability, often struggle to capture real-world preferences that are context-dependent, non-transitive, and shaped by dynamic multi-party interactions. This survey reviews AI alignment through a game-theoretic lens. Specifically, it organizes recent progress around key game-theoretic elements and synthesizes the literature along three challenges: preference diversity, alignment priority, and temporal dynamics. This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.