微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接后用系统浏览器访问。
XMT
短闻
信流 · 上滑连读 · 来源可核
1 / 24
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
arXiv:2608.…
- 14631v1 Announce Type: new Abstract: As consumers increasingly turn to…
- We benchmarked 14 LLMs on a structured set of topics related to cosmet…
- Web search was disabled throughout to assess each model's internalized…
RSS 官方收录 · 可信分层展示
Learning Agent Execution for KV-Cache Management in Agentic Serving
arXiv:2608.…
- 14624v1 Announce Type: new Abstract: Multi-agent LLM systems have emer…
- Across these workflows, every agent repeatedly executes a fixed contex…
- Existing LLM serving systems, however, manage KV-cache reactively usin…
RSS 官方收录 · 可信分层展示
Large Language Models and their Awareness of Mechanics and Spatial Geometry
arXiv:2608.…
- 14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perf…
- We present MecEng, a fully automated benchmark that evaluates LLMs on …
- The benchmark comprises 84 generic tasks on three difficulty levels, r…
RSS 官方收录 · 可信分层展示
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting.…
- Parenting is a critical and socially sensitive domain.
- Thus, evaluating advice provided by LLMs requires indicators beyond ag…
- With a multi-dimensional rubric created by parenting experts, this pap…
RSS 官方收录 · 可信分层展示
Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP
arXiv:2608.…
- 14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks incre…
- However, these protocols specify transport and discovery rather than s…
- We introduce a framework that (i) encodes classical negotiation mechan…
RSS 官方收录 · 可信分层展示
文章频道 - 2026上海国际公益广告奖正式启动,重点一篇搞定!
诚邀全球创意力量,以广告之美传递公益温度!上海国际公益广告奖正式启动!上海国际公益广告奖以“海纳百川 创益无界”为主题,以弘扬社会主义核心价值观、传承中华优秀传统文化、践行人民城市理念为目标,面向全球各界创作者征集原创公益广告作品。…
- 主题单元含“美丽世界”“文化中国”“人民城市”三个类别,专项单元设“青少年创意”“高校大学生创意”“AI创新”三个类别,旨在鼓励青年力量与新…
- 作品形式涵盖平面、音视频两类,征集时间截止9月8日,有意者可通过sipsac.
- 优秀作品将获得一定奖金,纳入公益广告作品库并全渠道展播。
RSS 官方收录 · 可信分层展示
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning
arXiv:2608.…
- 14610v1 Announce Type: new Abstract: Legal reasoning tasks such as leg…
- However, whether large language models (LLMs) can reliably perform thi…
- In this paper, we construct a benchmark to evaluate LLMs on temporal a…
RSS 官方收录 · 可信分层展示
Position: Medical AI Neglects Real Treatment Outcomes
arXiv:2608.14598v1 Announce Type: new Abstract: Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions.…
- But understanding of treatment itself is still inadequately trained an…
- This neglect seriously limits the potential of medical AI, and is alre…
- Real treatment outcomes, drawn from sources such as observational data…
RSS 官方收录 · 可信分层展示
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
arXiv:2608.…
- 14590v1 Announce Type: new Abstract: LLM agents increasingly perform i…
- However, no existing system provides formally grounded, task-level saf…
- Research remains fragmented across specification, verification, and en…
RSS 官方收录 · 可信分层展示