微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
SIMGUIDE: Procedurally Grounded Multi-Context Representations for Personalized Agent Planning
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.24888v1 Announce Type: new Abstract: Personalized AI agents overwhelmingly treat users as single entities: a flat profile concatenated into a prompt.…
- This fails when the same person holds different priorities across life…
- The core problem is not that agents lack information about users; it i…
- We introduce SIMGUIDE, a method that structures user context into type…
摘要引擎:抽取
正文提要
arXiv:2608.24888v1 Announce Type: new Abstract: Personalized AI agents overwhelmingly treat users as single entities: a flat profile concatenated into a prompt. This fails when the same person holds different priorities across life contexts -- and fails catastrophically when those priorities conflict. The core problem is not that agents lack information about users; it is that the format of user representations determines whether an agent can act on that information at all. We introduce SIMGUIDE, a method that structures user context into typed, domain-specific blocks called Sims and grounds each constraint with procedural examples drawn from past decisions. To evaluate this, we construct SIMBENCH, a diagnostic suite of 47 preference-conditioned planning tasks where the correct plan depends on which user context is active -- a property no existing benchmark tests. Declarative Sim constraints alone do not outperform retrieval-based personalization (RAG). Procedurally grounded Sims outperform RAG on GPT-4o (+7.9 Preference Adherence points, $p = 0.013$), and this advantage replicates on 100 $\tau$-bench tasks across both GPT-4o and Claude Sonnet~4.5 ($p \leq 0.023$). At the parametric level, the same principle holds: training distribution dominates whether parametric adaptation succeeds at all. Task-matched LoRA fine-tuning improves generation quality by 12.8 ROUGE-L points over the unadapted base model, and routing adapters by Sim type rather than user identity adds a further 7.3 points, robust to 28% routing error. Representation format -- not representation content -- is the first-order design variable.