微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接后用系统浏览器访问。
XMT
信源:arXiv cs.AI · 短平快可信阅读。
抖音式连刷节奏 · 视频号式正能量克制 · 第三种:可信信流——短平快,可核验。
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
arXiv:2608.11341v1 Announce Type: new Abstract: Apollo did not reach the Moon merely because its engineers could solve difficult equations.…
- It succeeded by turning a distant ambition into a mission architecture…
- AI now faces a similar transition: frontier models can solve difficult…
RSS 官方收录 · 可信分层展示
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
arXiv:2608.11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability.…
- We show, across three open agent-trace benchmarks (TheAgentCompany, $\…
- Leaderboards rank specialization, not capability.
RSS 官方收录 · 可信分层展示
Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning
arXiv:2608.11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals.…
- Existing approaches exhibit a "when-what" dissociation: traditional DN…
- We attribute this to the absence of a unified reasoning paradigm.
RSS 官方收录 · 可信分层展示
EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
arXiv:2608.11248v1 Announce Type: new Abstract: Long-term memory is essential for language agents operating across extended interactions and evolving tasks.…
- Existing memory-augmented agents mainly focus on storing and retrievin…
- In particular, previously distilled insights can become outdated, over…
RSS 官方收录 · 可信分层展示
BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model
arXiv:2608.11244v1 Announce Type: new Abstract: Construction standards are critical for building safety and sustainability.…
- Existing standard application workflows rely on keyword-based document…
- To address these limitations, this study develops a multimodal knowled…
RSS 官方收录 · 可信分层展示
The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification
arXiv:2608.…
- 11243v1 Announce Type: new Abstract: We argue that a single structural…
- , the agent does not escape its sandbox) is an off-support object.
RSS 官方收录 · 可信分层展示
VQ-bench: A Composable Vector Quantization Framework
arXiv:2608.11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure.…
- It is therefore experiencing a surge of renewed engineering and resear…
- This paper provides a unified framework for developing and benchmarkin…
RSS 官方收录 · 可信分层展示
CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference
arXiv:2608.11235v1 Announce Type: new Abstract: Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon.…
- Many predictions stabilize early, but blockwise decoding continues unt…
- Existing accelerators often rely on learned filters, modified scores, …
RSS 官方收录 · 可信分层展示
Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction
arXiv:2608.11237v1 Announce Type: new Abstract: Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs).…
- However, long-horizon autoregressive prediction remains challenging: l…
- Existing methods mainly improve state representations and operator bac…
RSS 官方收录 · 可信分层展示
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity.…
- Recent advances in AI agents create a timely opportunity to automate i…
- We present InfraBench, a benchmark suite for evaluating AI agents on r…
RSS 官方收录 · 可信分层展示
LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs
arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC).…
- Existing PIC methods, however, are built for full-attention models, wh…
- Hybrid LLMs break these primitives---they replace most attention layer…
RSS 官方收录 · 可信分层展示
Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet
arXiv:2608.…
- 11226v1 Announce Type: new Abstract: Reinforcement-learning post-train…
- We instrument GRPO training with half-second power telemetry at 7B, 14…
RSS 官方收录 · 可信分层展示
Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones
arXiv:2608.11225v1 Announce Type: new Abstract: AI "personality clones" force a re-examination of personal identity in operational terms.…
- Setting aside the hard problem of consciousness, we approach identity …
- We distinguish three criteria that "identity" conflates: fidelity to a…
RSS 官方收录 · 可信分层展示