微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开 ,或复制链接后用系统浏览器访问。
综合
官方
企业
汇聚
1 / 24
How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights
arXiv:2608.…
13617v1 Announce Type: new Abstract: Verifying whether clinical care f…
We present an expert-guided pipeline that constrains a large language …
The fuzzy layer encodes eight Surviving Sepsis Campaign bundle rules a…
RSS 官方收录 · 可信分层展示
SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data
arXiv:2608.…
13612v1 Announce Type: new Abstract: Natural-language interfaces to en…
SemPlan Benchmark evaluates this architectural design space with a det…
Four architectures are compared under the same model configuration: di…
RSS 官方收录 · 可信分层展示
Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
arXiv:2608.…
13608v1 Announce Type: new Abstract: Agentic "Continual Learning Harne…
But their value is conventionally measured by gains against labeled be…
Benchmark labels are scarce, stale, and unrepresentative, so a practit…
RSS 官方收录 · 可信分层展示
No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
arXiv:2608.13607v1 Announce Type: new Abstract: Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate.…
But aggregate gains say little about individual samples: an update can…
This paper studies how to predict such regressions from signals availa…
We compare single-model signals (confidence, logit margin, attention e…
RSS 官方收录 · 可信分层展示
MobileMem: Learning from a Year of Mobile Experiences
arXiv:2608.…
13606v1 Announce Type: new Abstract: The next generation of AI agents …
Such assistants require long-term memory to accumulate and leverage us…
We introduce MobileMem, a benchmark and framework for studying on-devi…
RSS 官方收录 · 可信分层展示
调用成功不等于判断正确:KDC 的行动治理主张
点击查看原文>
RSS 官方收录 · 可信分层展示
Vibe check:你的AI产品真的能落地吗
点击查看原文>
RSS 官方收录 · 可信分层展示
拒绝「差不多」,电商模型测评迎来「最严厉的父亲」
测评 AI 这件事,如今越来越难做了。几年前,AI 还是聊天机器人时,测评主要就是做题:写作的、数学的、编程的……最后形成一套跑分,用分数代表模型能力。但当 AI 从聊天框里走出来,变成能够调用工具、操作网页、处理文件、跨系统执行任务的 Agent,情况就变得复杂起来。…
平时我们看到的,是 AI 出现在手机上,能够刷淘宝、加购物车,最后下单。
实际上在商家的那一端,也有 AI 的参与,承担着比选购和下单更复杂的工作。
在专门考察电商环境中模型表现的 RealReplicaBench 测试中,经过 107 个真实商业任务的考核,所有参评模型都没有达到 60 …
RSS 官方收录 · 可信分层展示
Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents
arXiv:2608.…
13604v1 Announce Type: new Abstract: Detection of misunderstanding is …
This shift cuts communicators off from the resources repair depends on…
In this paper we analyse misunderstanding as a layered process in whic…
RSS 官方收录 · 可信分层展示
Active Perception for Embodied Disambiguation
arXiv:2608.…
13605v1 Announce Type: new Abstract: Natural language provides robots …
Existing interactive disambiguation methods primarily obtain additiona…
We propose an active-perception framework for embodied target disambig…
RSS 官方收录 · 可信分层展示
Measuring Cross-Task Behavioral Consistency in Language Model Agents
arXiv:2608.…
13598v1 Announce Type: new Abstract: Agent evaluation relies almost en…
We argue that behavioral consistency across tasks is a distinct and me…
BCM trains a model to predict task success from behavioral features of…
RSS 官方收录 · 可信分层展示
Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors
arXiv:2608.13591v1 Announce Type: new Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference.…
We study a different possibility: stable miscalibration, where a confi…
We combine two diagnostics: a label-aware output-level audit score tha…
On a multi-domain binary factual audit set, this audit score tracks wh…
RSS 官方收录 · 可信分层展示
AI Evaluation Should Work With Humans
arXiv:2608.…
13577v1 Announce Type: new Abstract: This position paper argues that t…
Instead, the AI community should pivot to evaluating the performance o…
We argue that this collaborative shift will foster AI systems that act…
RSS 官方收录 · 可信分层展示
Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents
arXiv:2608.…
13574v1 Announce Type: new Abstract: LLM agents increasingly operate a…
These capabilities make agents useful, but they also introduce risks r…
This paper presents Agentao, a governed local-first runtime for tool-u…
RSS 官方收录 · 可信分层展示
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
arXiv:2608.…
13573v1 Announce Type: new Abstract: Large Language Model (LLM) servin…
However, existing LLM serving workload studies remain limited in scale…
They often observe short time periods and provide limited visibility i…
RSS 官方收录 · 可信分层展示
Modular Cognitive Architecture Emerges in Large Language Models
arXiv:2608.…
13567v1 Announce Type: new Abstract: The human brain exhibits a striki…
Is this modular organization a fundamental principle of how intelligen…
Here, we test whether a similar organization emerges in Large Language…
RSS 官方收录 · 可信分层展示
Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking
arXiv:2608.…
13565v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architec…
Despite their widespread adoption, the relative importance of individu…
This paper presents a systematic layer-wise sensitivity analysis of th…
RSS 官方收录 · 可信分层展示
Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation
arXiv:2608.…
13564v1 Announce Type: new Abstract: Evaluating language-model agents …
Such a judge is a reward-free proxy whose value depends on whether it …
We instead induce the text of an agent-judging rubric from a small set…
RSS 官方收录 · 可信分层展示
闭门、举牌、真话,这场路演有点不一样
作者|Li Yuan 编辑|郑玄 会算命的非遗挂件?专挑膝盖下手的外骨骼?会跳的机器人?8 月 7 日,地瓜机器人办的这场闭门路演「地心引力 DemoDay 2026」,十个项目轮番上台,画风一个比一个野。…
有做具身智能的,也有做智能硬件的,大部分都还初具雏形:一些项目,半年前才刚正式启动。
地瓜选择本场 Demoday 选手的调性,与其他不太一样。
CEO 王丛开场就说:他不太在乎地瓜自己出货量多大、市场占有率多高,他更想看到的画面是,被他们孵化出来的那些五花八门的东西,能被随便拿一个回…
RSS 官方收录 · 可信分层展示
世界机器人大会今年不一定「Wow!」,但有五个问题很值得关注
作者|Li Yuan 编辑|张鹏 8 月 10 日,宇树科技完成科创板申购,发行价 150.80 元,发行市值约 610 亿元,发行市盈率 219 倍,是通用设备行业平均值的近六倍。…
多家机构给出的挂牌后估值直接标到千亿以上,敲钟的日子很可能就落在这个月下旬,几乎与世界机器人大会撞在一起。
对标电动车,意味着一个万亿级的制造业;对标 AI,则意味着根本没有天花板——毕竟具身智能被寄予的,是让 AI 真正落进物理世界。
热的是估值,冷的是订单。
RSS 官方收录 · 可信分层展示
Dropbox 集成 MCP 与 Dash,将安全设计与代码审查连接起来
点击查看原文>
RSS 官方收录 · 可信分层展示
编程能力提高50%!GLM-5.3 满分通过了GPT-5.6给的Coding 测试
点击查看原文>
RSS 官方收录 · 可信分层展示
中行回应“Token贷”:已投放3户800万元;头部AI大厂员工:90小时工作制成常态;宇树科技中签者不敢发朋友圈:怕被嫉妒|AI 周报
点击查看原文>
RSS 官方收录 · 可信分层展示
问界「童车」上市,华为联合设计;DeepSeek 涨价策略今日实行;大学生用 AI 人脸视频盗刷 5 万元被判刑
IPO 前 Anthropic CEO 达里奥 · 阿莫迪罕见发长文回应质疑,预告未来 5-10 年 AI 将治愈多数疾病 Anthropic CEO 达里奥 · 阿莫迪(Dario Amodei)几乎从不碰社交媒体。…
但就在 IPO 估值被推到 2 万亿美元的节骨眼上,他罕见地发了一篇长文,正面回应 Anthropic 如何看待监管。
长文中,达里奥 · 阿莫迪回「我们设计的每一条政策,都让前沿公司承担更多成本,同时有利于规模更小的竞争者。
」他同样支持 Demis Hassabis 提出的、建立一个类似 FINRA(美国金融业监管局)的 AI 监管机构的想法。
RSS 官方收录 · 可信分层展示
上滑下一条
上滑 · j/k · m/u · h 隐藏 · a 稍后 · o 原文 · e 详情 · t 今日 · i 模式 · f 搜索