微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开 ,或复制链接后用系统浏览器访问。
综合
官方
企业
汇聚
1 / 24
CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
arXiv:2608.…
12555v1 Announce Type: new Abstract: Predictive explanation methods at…
We introduce the Causal Attribution Score (CAS), a compact score archi…
CAS starts from an identified interventional coalition game, allocates…
RSS 官方收录 · 可信分层展示
$\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution
arXiv:2608.…
12522v1 Announce Type: new Abstract: LLM-based program evolution syste…
We introduce $\varepsilon$-MemEvo, a framework for cross-task knowledg…
$\varepsilon$-MemEvo stores prior experience as task-agnostic tactic m…
RSS 官方收录 · 可信分层展示
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
arXiv:2608.…
12476v1 Announce Type: new Abstract: Long-term agent memory is usually…
We introduce Governed Persistent Memory (GPM), an auditable bitemporal…
Five executable clauses cover ledger integrity, source binding, confli…
RSS 官方收录 · 可信分层展示
MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
arXiv:2608.…
12428v1 Announce Type: new Abstract: Memory is a core component of AI …
However, existing memory systems often remain fixed after development,…
We present MindMemOS, a portable and self-evolving memory operating la…
RSS 官方收录 · 可信分层展示
Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
arXiv:2608.…
12426v1 Announce Type: new Abstract: Large language models are increas…
Individual constraints are handled proficiently, but the compositional…
We introduce Constraint Saturation Evaluation (CSE), a procedurally ge…
RSS 官方收录 · 可信分层展示
Research Assistant: AstraZeneca's Agentic System for R&D
arXiv:2608.…
12395v1 Announce Type: new Abstract: We describe Research Assistant, a…
The system provides a chat-style interface that brings together eviden…
It supports both a fast mode for direct question answering and a multi…
RSS 官方收录 · 可信分层展示
Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
arXiv:2608.…
12385v1 Announce Type: new Abstract: As large language models serve mo…
The two inference phases stress hardware differently: prompt prefill i…
Conventional width or depth scaling increases both costs together beca…
RSS 官方收录 · 可信分层展示
Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization
arXiv:2608.…
12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot pe…
Existing adaptation methods struggle to calibrate update magnitude und…
To calibrate adaptation to evidence quality, we propose PAC-Bayes-regu…
RSS 官方收录 · 可信分层展示
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
arXiv:2608.…
12373v1 Announce Type: new Abstract: Large language models are increas…
We test nine models from six providers and ask whether the language of…
We use single-turn game-theoretic vignettes in which a model advises a…
RSS 官方收录 · 可信分层展示
Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
arXiv:2608.12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers.…
This position paper argues that in many settings, particularly high-st…
We review evidence that cognitive alignment improves understandability…
We outline the gaps between existing alignment methods and what is nee…
RSS 官方收录 · 可信分层展示
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).…
Yet agreement in final labels does not show that human annotators and …
Two agents may reach the same judgment while appealing to different pr…
We test this distinction using a curated 500-item ETHICS-derived bench…
RSS 官方收录 · 可信分层展示
Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing
arXiv:2608.…
12371v1 Announce Type: new Abstract: Stream-processing systems increas…
This paper proposes \emph{MAS-DecStream}, whose main contribution is \…
Edge-cluster agents refine natural-language offloading proposals from …
RSS 官方收录 · 可信分层展示
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
arXiv:2608.…
12346v1 Announce Type: new Abstract: This position paper argues that m…
By mapping current alignment techniques to the possibility and actual …
We need to discuss this dual-use potential now, as its risk is exacerb…
RSS 官方收录 · 可信分层展示
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
arXiv:2608.…
12345v1 Announce Type: new Abstract: Language models are increasingly …
We introduce IntegrityBench, a benchmark evaluating misconduct classif…
Evaluating 18 frontier model variants, we find that under peak pressur…
RSS 官方收录 · 可信分层展示
Position: Reasoning is a Learnable Rule-Based Process
arXiv:2608.12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today.…
Historically the purview of symbolic AI, recent advances have mainly e…
Despite immense interest and rapid progress, the generative AI communi…
This position contends that definitional ambiguity leaves the construc…
RSS 官方收录 · 可信分层展示
限量手办 + 实景体验,浙江人行NAVIAI2026WRC 福利提前曝光
8 月 19 日至 23 日,2026 世界机器人大会(WRC2026)将在北京亦庄亦创国际会展中心举办。浙江人形机器人创新中心将携旗下 NAVIAI 人形机器人亮相 C-105 展位,集中展示人形机器人在多元场景的落地应用成果。…
本次展会上,NAVIAI 将呈现工业制造、智慧零售、家庭服务、遥操数采多场景的实操能力,现场演示拆垛分拣搬运、货品递送、烹饪清洁、远程数据采…
另有文娱演绎类场景在 8 月 19 日开展当天作为特别环节限时展示。
大会期间,浙江人形机器人创新中心有限公司首席科学家熊蓉教授将于 8 月 21 日站上主论坛,分享 NAVIAI 的技术演进与未来落地规划,多…
RSS 官方收录 · 可信分层展示
智谱发布 GLM-5.3,编程能力更强;传苹果训练国内专用 AI 模型;微信:朋友圈现在、过去、未来都不会有二次编辑功能 | 极客早知道
智谱正式发布 GLM-5.3,拥有更强编程能力 8 月 14 日,据介绍,与 GLM-5.2 相比,GLM-5.3 基座模型未变,但通过极致的后训练 Scaling 大大提高了模型的智能上界。…
3 拥有更强的编程能力,在内部自建体感评测中较 GLM-5.
2 提升 50%,在包括 TerminalBench3.
0、Agents'LastExam(CLI)在内的公开基准测试中取得开源第一。
RSS 官方收录 · 可信分层展示
早报|曝苹果与阿里合作训练AI模型/微信:永不推出朋友圈二次编辑/售价20万,追觅首台手机交付
曝苹果与阿里合作,为中国市场训练自研 AI 模型 微信确认朋友圈永不推出二次编辑功能 售价 20 万元,追觅首台 AURORA 手机交付:24K 足金镶宝石 Google DeepMind 或裁员三分之一以上,资源转向 Flash WorkBuddy 接入 GLM-5.…
3 广州推出「Token 贷」:按算力合同和 Token 消耗额度授信 曝 DeepSeek 正研发情感 AI 模型 调查:美国年轻人普遍不…
3,同一基座靠后训练提升编程与网络安全能力 Ling-3.
0-tiny 与 ASystem AReno 打通单机 Agentic RL 训练闭环 Suno Studio 2.
RSS 官方收录 · 可信分层展示
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs
arXiv:2608.…
11403v1 Announce Type: new Abstract: Self-consistency (SC) via majorit…
On the full GPQA Diamond benchmark (198 graduate-level science questio…
6% of problems for Qwen2.
RSS 官方收录 · 可信分层展示
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
arXiv:2608.…
11381v1 Announce Type: new Abstract: We study whether the localized nu…
Larix maps a 16-lens European listed-real-estate analysis framework to…
Across 19 firms spanning seven regulatory wrappers, decomposition impr…
RSS 官方收录 · 可信分层展示
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
arXiv:2608.…
11343v1 Announce Type: new Abstract: Multimodal retrieval and classifi…
The March 2026 release of Gemini Embedding 2, Google's first natively …
Simultaneously, frontier Large language models (LLMs) have also demons…
RSS 官方收录 · 可信分层展示
Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces
arXiv:2608.…
11354v1 Announce Type: new Abstract: Modern recommender systems treat …
As interfaces evolve from static layouts toward generative UIs and imm…
We propose an Inverse Theory of Mind (IToM) pipeline that reasons back…
RSS 官方收录 · 可信分层展示
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
arXiv:2608.11341v1 Announce Type: new Abstract: Apollo did not reach the Moon merely because its engineers could solve difficult equations.…
It succeeded by turning a distant ambition into a mission architecture…
AI now faces a similar transition: frontier models can solve difficult…
We introduce Apodex Discovery, a framework for building and evaluating…
RSS 官方收录 · 可信分层展示
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
arXiv:2608.11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability.…
We show, across three open agent-trace benchmarks (TheAgentCompany, $\…
Leaderboards rank specialization, not capability.
We arrive at this through a four-facet Generalizability Theory varianc…
RSS 官方收录 · 可信分层展示
上滑下一条
上滑 · j/k · m/u · h 隐藏 · a 稍后 · o 原文 · e 详情 · t 今日 · i 模式 · f 搜索