微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting.…
- Parenting is a critical and socially sensitive domain.
- Thus, evaluating advice provided by LLMs requires indicators beyond ag…
- With a multi-dimensional rubric created by parenting experts, this pap…
摘要引擎:抽取
正文提要
arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional rubric created by parenting experts, this paper evaluates 15 LLMs across 100 parenting scenarios in 2 languages (English and Chinese), using an LLM-as-a-judge method. Results show that aggregate scores can hide rubric item-specific weaknesses, models implicitly encourage different parenting styles, and language influences responses. We highlight the importance of evaluation output auditability and challenges involved in evaluating LLM-generated advice in domains like parenting. Our findings provide important insights for selecting LLMs for direct user engagement and the development of user-facing parenting advice applications.