微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开 ,或复制链接后用系统浏览器访问。
综合
官方
企业
汇聚
当前信源:arXiv cs.AI
· 清除信源筛选
CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation
arXiv:2608.…
12842v1 Announce Type: new Abstract: Model merging has recently attrac…
However, parameter conflicts and knowledge interference across tasks o…
RSS 官方收录 · 可信分层展示
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
arXiv:2608.…
12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Res…
We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a R…
RSS 官方收录 · 可信分层展示
PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs
arXiv:2608.…
12762v1 Announce Type: new Abstract: Schedulability analysis is essent…
Mechanized verification in PROSA/ROCQ offers a rigorous alternative, y…
RSS 官方收录 · 可信分层展示
Correct Is Not Governed: Provenance Integrity in Agentic Workflows
arXiv:2608.12761v1 Announce Type: new Abstract: Agentic workflows are commonly evaluated by whether they reach the correct outcome.…
That is insufficient in institutional settings, where a correct action…
We define governed execution as work whose decisions, completion, and …
RSS 官方收录 · 可信分层展示
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
arXiv:2608.12743v1 Announce Type: new Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants.…
To improve the spatial reasoning ability of VLM agents, existing work …
One line uses post-training methods, such as supervised fine-tuning an…
RSS 官方收录 · 可信分层展示
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science.…
The usual approach is to present the problem to the model and use its …
However, beyond this best guess, discovery can be enhanced by increasi…
RSS 官方收录 · 可信分层展示
The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
arXiv:2608.…
12677v1 Announce Type: new Abstract: Detecting infection-related behav…
In this study, a YOLO- and Contrastive Language-Image Pre-training (CL…
RSS 官方收录 · 可信分层展示
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
arXiv:2608.…
12675v1 Announce Type: new Abstract: Retrieval-Augmented Generation (R…
Existing privacy research on RAG has focused on preventing unauthorize…
RSS 官方收录 · 可信分层展示
Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
arXiv:2608.12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers.…
However, with catalogs spanning millions of active items, manual gover…
Inconsistent pricing across item variants distorts customer value perc…
RSS 官方收录 · 可信分层展示
On the Expressive Power of Transformers
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.…
Because of their ubiquity and computational capability, there is a rap…
In this endeavor, circuit complexity has by and large emerged as the "…
RSS 官方收录 · 可信分层展示
Designing AI Pipelines for Decision-Ready ITSM Intelligence
arXiv:2608.…
12670v1 Announce Type: new Abstract: IT service management (ITSM) syst…
This paper presents a sociotechnical AI pipeline, designed and evaluat…
RSS 官方收录 · 可信分层展示
General Probabilities of Causation with Causal Knowledge
arXiv:2608.…
12657v1 Announce Type: new Abstract: Probabilities of causation (PoCs)…
Tian and Pearl first derived theoretically sharp bounds for binary PoC…
RSS 官方收录 · 可信分层展示
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment.…
The steering decision is the pre-commit choice at that boundary: proce…
We introduce SteerBench-Work, an incident-anchored, bidirectional benc…
RSS 官方收录 · 可信分层展示
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.…
Judges are typically validated by accuracy on golden data, but accurac…
We introduce the \emph{Wiggle Framework}, a unified stress test for ep…
RSS 官方收录 · 可信分层展示
@skills: Attention is all you have
arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately.…
The dominant delivery model is installation: once installed, a skill's…
This leaves the long tail with no practical path to use and forces tea…
RSS 官方收录 · 可信分层展示
Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
arXiv:2608.…
12599v1 Announce Type: new Abstract: Multi-turn dialogues let users re…
No existing instrument measures this influence per clause, predicts it…
RSS 官方收录 · 可信分层展示
DiG-bench: Discovery in Games
arXiv:2608.12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process.…
Despite its importance, there is a gap in the current AI benchmark lan…
To address this gap, we release a new benchmark: DiG-bench (Discovery …
RSS 官方收录 · 可信分层展示
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
arXiv:2608.…
12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requ…
We present ThyroidXAgent, a clinician-interactive agentic AI system th…
RSS 官方收录 · 可信分层展示
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
arXiv:2608.…
12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires…
Additionally, surfacing reasoning mistakes that the model makes would …
RSS 官方收录 · 可信分层展示
Trie Automata for Constrained Decoding over Large Finite Sets
arXiv:2608.…
12574v1 Announce Type: new Abstract: Large language models increasingl…
Current constrained decoding systems handle this through general-purpo…
RSS 官方收录 · 可信分层展示
CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
arXiv:2608.…
12555v1 Announce Type: new Abstract: Predictive explanation methods at…
We introduce the Causal Attribution Score (CAS), a compact score archi…
RSS 官方收录 · 可信分层展示
$\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution
arXiv:2608.…
12522v1 Announce Type: new Abstract: LLM-based program evolution syste…
We introduce $\varepsilon$-MemEvo, a framework for cross-task knowledg…
RSS 官方收录 · 可信分层展示
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
arXiv:2608.…
12476v1 Announce Type: new Abstract: Long-term agent memory is usually…
We introduce Governed Persistent Memory (GPM), an auditable bitemporal…
RSS 官方收录 · 可信分层展示
MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
arXiv:2608.…
12428v1 Announce Type: new Abstract: Memory is a core component of AI …
However, existing memory systems often remain fixed after development,…
RSS 官方收录 · 可信分层展示
Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
arXiv:2608.…
12426v1 Announce Type: new Abstract: Large language models are increas…
Individual constraints are handled proficiently, but the compositional…
RSS 官方收录 · 可信分层展示
Research Assistant: AstraZeneca's Agentic System for R&D
arXiv:2608.…
12395v1 Announce Type: new Abstract: We describe Research Assistant, a…
The system provides a chat-style interface that brings together eviden…
RSS 官方收录 · 可信分层展示
Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
arXiv:2608.…
12385v1 Announce Type: new Abstract: As large language models serve mo…
The two inference phases stress hardware differently: prompt prefill i…
RSS 官方收录 · 可信分层展示
Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization
arXiv:2608.…
12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot pe…
Existing adaptation methods struggle to calibrate update magnitude und…
RSS 官方收录 · 可信分层展示
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
arXiv:2608.…
12373v1 Announce Type: new Abstract: Large language models are increas…
We test nine models from six providers and ask whether the language of…
RSS 官方收录 · 可信分层展示
Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
arXiv:2608.12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers.…
This position paper argues that in many settings, particularly high-st…
We review evidence that cognitive alignment improves understandability…
RSS 官方收录 · 可信分层展示
下滑 · j/k · m/u · i 信流 · a 稍后 · o 原文 · e 详情 · t 今日 · f 搜索