Skip to main content
Aggregate arXiv cs.AI 人工智能 18 Aug 2026 - 14:00

A Human-Centred Approach to Benchmarking LLMs for Parenting Advice

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting.…

  • Parenting is a critical and socially sensitive domain.
  • Thus, evaluating advice provided by LLMs requires indicators beyond ag…
  • With a multi-dimensional rubric created by parenting experts, this pap…

摘要引擎:抽取

正文提要

arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional rubric created by parenting experts, this paper evaluates 15 LLMs across 100 parenting scenarios in 2 languages (English and Chinese), using an LLM-as-a-judge method. Results show that aggregate scores can hide rubric item-specific weaknesses, models implicitly encourage different parenting styles, and language influences responses. We highlight the importance of evaluation output auditability and challenges involved in evaluating LLM-generated advice in domains like parenting. Our findings provide important insights for selecting LLMs for direct user engagement and the development of user-facing parenting advice applications.

来源:https://arxiv.org/abs/2608.14622

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表