微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开 ,或复制链接后用系统浏览器访问。
综合
官方
企业
汇聚
当前信源:arXiv cs.AI
· 清除信源筛选
Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
arXiv:2608.…
12892v1 Announce Type: new Abstract: Activation steering turns localiz…
We introduce Predictive Memory Localization (PML), which treats the me…
RSS 官方收录 · 可信分层展示
ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification
arXiv:2608.…
12877v1 Announce Type: new Abstract: Multi-hop fact verification, whic…
Recent methods primarily rely on multi-agent collaboration to decompos…
RSS 官方收录 · 可信分层展示
AI and Consumer Rights in India Working Paper
arXiv:2608.12863v1 Announce Type: new Abstract: As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved.…
This working paper examines whether India's Consumer Protection Act, 2…
The Act's broad definitions of product liability, harm, and deficiency…
RSS 官方收录 · 可信分层展示
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state.…
An unsafe success can thereby become reusable policy after its trigger…
Skill evolution makes this failure measurable by distilling operationa…
RSS 官方收录 · 可信分层展示
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
arXiv:2608.…
12847v1 Announce Type: new Abstract: Retrieval can identify a past tra…
We identify this post-retrieval reuse step as a distinct bottleneck fo…
RSS 官方收录 · 可信分层展示
CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation
arXiv:2608.…
12842v1 Announce Type: new Abstract: Model merging has recently attrac…
However, parameter conflicts and knowledge interference across tasks o…
RSS 官方收录 · 可信分层展示
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
arXiv:2608.…
12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Res…
We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a R…
RSS 官方收录 · 可信分层展示
PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs
arXiv:2608.…
12762v1 Announce Type: new Abstract: Schedulability analysis is essent…
Mechanized verification in PROSA/ROCQ offers a rigorous alternative, y…
RSS 官方收录 · 可信分层展示
Correct Is Not Governed: Provenance Integrity in Agentic Workflows
arXiv:2608.12761v1 Announce Type: new Abstract: Agentic workflows are commonly evaluated by whether they reach the correct outcome.…
That is insufficient in institutional settings, where a correct action…
We define governed execution as work whose decisions, completion, and …
RSS 官方收录 · 可信分层展示
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
arXiv:2608.12743v1 Announce Type: new Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants.…
To improve the spatial reasoning ability of VLM agents, existing work …
One line uses post-training methods, such as supervised fine-tuning an…
RSS 官方收录 · 可信分层展示
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science.…
The usual approach is to present the problem to the model and use its …
However, beyond this best guess, discovery can be enhanced by increasi…
RSS 官方收录 · 可信分层展示
The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
arXiv:2608.…
12677v1 Announce Type: new Abstract: Detecting infection-related behav…
In this study, a YOLO- and Contrastive Language-Image Pre-training (CL…
RSS 官方收录 · 可信分层展示
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
arXiv:2608.…
12675v1 Announce Type: new Abstract: Retrieval-Augmented Generation (R…
Existing privacy research on RAG has focused on preventing unauthorize…
RSS 官方收录 · 可信分层展示
Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
arXiv:2608.12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers.…
However, with catalogs spanning millions of active items, manual gover…
Inconsistent pricing across item variants distorts customer value perc…
RSS 官方收录 · 可信分层展示
On the Expressive Power of Transformers
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.…
Because of their ubiquity and computational capability, there is a rap…
In this endeavor, circuit complexity has by and large emerged as the "…
RSS 官方收录 · 可信分层展示
Designing AI Pipelines for Decision-Ready ITSM Intelligence
arXiv:2608.…
12670v1 Announce Type: new Abstract: IT service management (ITSM) syst…
This paper presents a sociotechnical AI pipeline, designed and evaluat…
RSS 官方收录 · 可信分层展示
General Probabilities of Causation with Causal Knowledge
arXiv:2608.…
12657v1 Announce Type: new Abstract: Probabilities of causation (PoCs)…
Tian and Pearl first derived theoretically sharp bounds for binary PoC…
RSS 官方收录 · 可信分层展示
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment.…
The steering decision is the pre-commit choice at that boundary: proce…
We introduce SteerBench-Work, an incident-anchored, bidirectional benc…
RSS 官方收录 · 可信分层展示
@skills: Attention is all you have
arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately.…
The dominant delivery model is installation: once installed, a skill's…
This leaves the long tail with no practical path to use and forces tea…
RSS 官方收录 · 可信分层展示
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.…
Judges are typically validated by accuracy on golden data, but accurac…
We introduce the \emph{Wiggle Framework}, a unified stress test for ep…
RSS 官方收录 · 可信分层展示
Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
arXiv:2608.…
12599v1 Announce Type: new Abstract: Multi-turn dialogues let users re…
No existing instrument measures this influence per clause, predicts it…
RSS 官方收录 · 可信分层展示
DiG-bench: Discovery in Games
arXiv:2608.12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process.…
Despite its importance, there is a gap in the current AI benchmark lan…
To address this gap, we release a new benchmark: DiG-bench (Discovery …
RSS 官方收录 · 可信分层展示
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
arXiv:2608.…
12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requ…
We present ThyroidXAgent, a clinician-interactive agentic AI system th…
RSS 官方收录 · 可信分层展示
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
arXiv:2608.…
12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires…
Additionally, surfacing reasoning mistakes that the model makes would …
RSS 官方收录 · 可信分层展示
Trie Automata for Constrained Decoding over Large Finite Sets
arXiv:2608.…
12574v1 Announce Type: new Abstract: Large language models increasingl…
Current constrained decoding systems handle this through general-purpo…
RSS 官方收录 · 可信分层展示
CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
arXiv:2608.…
12555v1 Announce Type: new Abstract: Predictive explanation methods at…
We introduce the Causal Attribution Score (CAS), a compact score archi…
RSS 官方收录 · 可信分层展示
$\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution
arXiv:2608.…
12522v1 Announce Type: new Abstract: LLM-based program evolution syste…
We introduce $\varepsilon$-MemEvo, a framework for cross-task knowledg…
RSS 官方收录 · 可信分层展示
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
arXiv:2608.…
12476v1 Announce Type: new Abstract: Long-term agent memory is usually…
We introduce Governed Persistent Memory (GPM), an auditable bitemporal…
RSS 官方收录 · 可信分层展示
MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
arXiv:2608.…
12428v1 Announce Type: new Abstract: Memory is a core component of AI …
However, existing memory systems often remain fixed after development,…
RSS 官方收录 · 可信分层展示
Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
arXiv:2608.…
12426v1 Announce Type: new Abstract: Large language models are increas…
Individual constraints are handled proficiently, but the compositional…
RSS 官方收录 · 可信分层展示
下滑 · j/k · m/u · i 信流 · a 稍后 · o 原文 · e 详情 · t 今日 · f 搜索