微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接后用系统浏览器访问。
XMT
信流 · 上滑连读 · 来源可核
Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
arXiv:2608.…
- 12892v1 Announce Type: new Abstract: Activation steering turns localiz…
- We introduce Predictive Memory Localization (PML), which treats the me…
- PML separates random-calibrated target movement from semantic-neighbor…
RSS 官方收录 · 可信分层展示
AI and Consumer Rights in India Working Paper
arXiv:2608.12863v1 Announce Type: new Abstract: As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved.…
- This working paper examines whether India's Consumer Protection Act, 2…
- The Act's broad definitions of product liability, harm, and deficiency…
- However, significant gaps remain.
RSS 官方收录 · 可信分层展示
ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification
arXiv:2608.…
- 12877v1 Announce Type: new Abstract: Multi-hop fact verification, whic…
- Recent methods primarily rely on multi-agent collaboration to decompos…
- However, these methods face two critical limitations: (1) agents may p…
RSS 官方收录 · 可信分层展示
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state.…
- An unsafe success can thereby become reusable policy after its trigger…
- Skill evolution makes this failure measurable by distilling operationa…
- Because evolution optimizes task outcomes rather than procedure safety…
RSS 官方收录 · 可信分层展示
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
arXiv:2608.…
- 12847v1 Announce Type: new Abstract: Retrieval can identify a past tra…
- We identify this post-retrieval reuse step as a distinct bottleneck fo…
- We instantiate the framework with query-conditioned reuse (QCR), a del…
RSS 官方收录 · 可信分层展示
CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation
arXiv:2608.…
- 12842v1 Announce Type: new Abstract: Model merging has recently attrac…
- However, parameter conflicts and knowledge interference across tasks o…
- Prior work introduced Conflict-Aware and Balanced Sparsification (CABS…
RSS 官方收录 · 可信分层展示
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
arXiv:2608.…
- 12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Res…
- We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a R…
- The framework operates through two synergistic components: the Academi…
RSS 官方收录 · 可信分层展示
PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs
arXiv:2608.…
- 12762v1 Announce Type: new Abstract: Schedulability analysis is essent…
- Mechanized verification in PROSA/ROCQ offers a rigorous alternative, y…
- Recent successes of large language models (LLMs) across a wide range o…
RSS 官方收录 · 可信分层展示
Correct Is Not Governed: Provenance Integrity in Agentic Workflows
arXiv:2608.12761v1 Announce Type: new Abstract: Agentic workflows are commonly evaluated by whether they reach the correct outcome.…
- That is insufficient in institutional settings, where a correct action…
- We define governed execution as work whose decisions, completion, and …
- We present Matrix, a deterministic causal-state layer that records aut…
RSS 官方收录 · 可信分层展示
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
arXiv:2608.12743v1 Announce Type: new Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants.…
- To improve the spatial reasoning ability of VLM agents, existing work …
- One line uses post-training methods, such as supervised fine-tuning an…
- Another line adopts an agentic paradigm in which the model calls exter…
RSS 官方收录 · 可信分层展示
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science.…
- The usual approach is to present the problem to the model and use its …
- However, beyond this best guess, discovery can be enhanced by increasi…
- In a process called pass@k, the model is allowed to explore the soluti…
RSS 官方收录 · 可信分层展示
The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
arXiv:2608.…
- 12677v1 Announce Type: new Abstract: Detecting infection-related behav…
- In this study, a YOLO- and Contrastive Language-Image Pre-training (CL…
- First, YOLO is used to isolate mosquito regions from the background.
RSS 官方收录 · 可信分层展示
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
arXiv:2608.…
- 12675v1 Announce Type: new Abstract: Retrieval-Augmented Generation (R…
- Existing privacy research on RAG has focused on preventing unauthorize…
- However, another important problem that is often overlooked in RAG pri…
RSS 官方收录 · 可信分层展示
Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
arXiv:2608.12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers.…
- However, with catalogs spanning millions of active items, manual gover…
- Inconsistent pricing across item variants distorts customer value perc…
- To address this, we present a scalable, context-aware Multi-Agent Fram…
RSS 官方收录 · 可信分层展示
On the Expressive Power of Transformers
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.…
- Because of their ubiquity and computational capability, there is a rap…
- In this endeavor, circuit complexity has by and large emerged as the "…
- Here, we present an overview of selected results that delineate the ex…
RSS 官方收录 · 可信分层展示
Designing AI Pipelines for Decision-Ready ITSM Intelligence
arXiv:2608.…
- 12670v1 Announce Type: new Abstract: IT service management (ITSM) syst…
- This paper presents a sociotechnical AI pipeline, designed and evaluat…
- The pipeline combines LLM-based schema normalization, HDBSCAN sub-topi…
RSS 官方收录 · 可信分层展示
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment.…
- The steering decision is the pre-commit choice at that boundary: proce…
- We introduce SteerBench-Work, an incident-anchored, bidirectional benc…
- Release v2026-05 contains 106 scenarios anchored in public incidents, …
RSS 官方收录 · 可信分层展示
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.…
- Judges are typically validated by accuracy on golden data, but accurac…
- We introduce the \emph{Wiggle Framework}, a unified stress test for ep…
- The framework decomposes judge robustness along three dimensions: Mech…
RSS 官方收录 · 可信分层展示
@skills: Attention is all you have
arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately.…
- The dominant delivery model is installation: once installed, a skill's…
- This leaves the long tail with no practical path to use and forces tea…
- We observe that installation bundles three separable functions: conten…
RSS 官方收录 · 可信分层展示
Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
arXiv:2608.…
- 12599v1 Announce Type: new Abstract: Multi-turn dialogues let users re…
- No existing instrument measures this influence per clause, predicts it…
- \sysname{} closes the three gaps through the model API alone: a contra…
RSS 官方收录 · 可信分层展示
DiG-bench: Discovery in Games
arXiv:2608.12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process.…
- Despite its importance, there is a gap in the current AI benchmark lan…
- To address this gap, we release a new benchmark: DiG-bench (Discovery …
- DiG-bench consists of a set of 70 independent games.
RSS 官方收录 · 可信分层展示