微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接后用系统浏览器访问。
XMT
信流 · 上滑连读 · 来源可核
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
arXiv:2608.…
- 14667v1 Announce Type: new Abstract: Large language model-based agents…
- We argue that this overlooks the social aspects of scientific teamwork…
- We establish these points through literature and empirical analysis, a…
RSS 官方收录 · 可信分层展示
Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring
arXiv:2608.…
- 14666v1 Announce Type: new Abstract: Unsupervised fault detection in i…
- This misses coupling faults, where the physical relationship between s…
- Such faults evade marginal monitoring and persist as latent failures, …
RSS 官方收录 · 可信分层展示
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
arXiv:2608.14659v1 Announce Type: new Abstract: Large language models for code generation often produce incorrect solutions without reliable indicators of failure.…
- We study whether uncertainty estimation methods developed for natural …
- We evaluate five uncertainty methods: mean token entropy, verbalized c…
- We find that multi-sample $P(\text{True})$ achieves the strongest corr…
RSS 官方收录 · 可信分层展示
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
arXiv:2608.…
- 14651v1 Announce Type: new Abstract: Effective disaster risk communica…
- Recent advancements in Artificial Intelligence (AI), especially Multi-…
- However, their suitability for deployment rests on a property that rec…
RSS 官方收录 · 可信分层展示
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
arXiv:2608.…
- 14641v1 Announce Type: new Abstract: Agentic systems increasingly dele…
- We present a common measurement protocol and hybrid evaluation of four…
- We evaluate 290 frozen tasks against a locked matrix of 2,610 candidat…
RSS 官方收录 · 可信分层展示
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
arXiv:2608.…
- 14631v1 Announce Type: new Abstract: As consumers increasingly turn to…
- We benchmarked 14 LLMs on a structured set of topics related to cosmet…
- Web search was disabled throughout to assess each model's internalized…
RSS 官方收录 · 可信分层展示
Learning Agent Execution for KV-Cache Management in Agentic Serving
arXiv:2608.…
- 14624v1 Announce Type: new Abstract: Multi-agent LLM systems have emer…
- Across these workflows, every agent repeatedly executes a fixed contex…
- Existing LLM serving systems, however, manage KV-cache reactively usin…
RSS 官方收录 · 可信分层展示
Large Language Models and their Awareness of Mechanics and Spatial Geometry
arXiv:2608.…
- 14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perf…
- We present MecEng, a fully automated benchmark that evaluates LLMs on …
- The benchmark comprises 84 generic tasks on three difficulty levels, r…
RSS 官方收录 · 可信分层展示
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting.…
- Parenting is a critical and socially sensitive domain.
- Thus, evaluating advice provided by LLMs requires indicators beyond ag…
- With a multi-dimensional rubric created by parenting experts, this pap…
RSS 官方收录 · 可信分层展示
Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP
arXiv:2608.…
- 14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks incre…
- However, these protocols specify transport and discovery rather than s…
- We introduce a framework that (i) encodes classical negotiation mechan…
RSS 官方收录 · 可信分层展示
文章频道 - 2026上海国际公益广告奖正式启动,重点一篇搞定!
诚邀全球创意力量,以广告之美传递公益温度!上海国际公益广告奖正式启动!上海国际公益广告奖以“海纳百川 创益无界”为主题,以弘扬社会主义核心价值观、传承中华优秀传统文化、践行人民城市理念为目标,面向全球各界创作者征集原创公益广告作品。…
- 主题单元含“美丽世界”“文化中国”“人民城市”三个类别,专项单元设“青少年创意”“高校大学生创意”“AI创新”三个类别,旨在鼓励青年力量与新…
- 作品形式涵盖平面、音视频两类,征集时间截止9月8日,有意者可通过sipsac.
- 优秀作品将获得一定奖金,纳入公益广告作品库并全渠道展播。
RSS 官方收录 · 可信分层展示