Skip to main content

XMT

短闻

信流 · 上滑连读 · 来源可核

今日 稍后 搜索 RSS

当前信源:arXiv cs.AI · 清除信源筛选

1 / 24
Aggregate arXiv cs.AI 人工智能 45″

When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

arXiv:2608.…

  • 11403v1 Announce Type: new Abstract: Self-consistency (SC) via majorit…
  • On the full GPQA Diamond benchmark (198 graduate-level science questio…
  • 6% of problems for Qwen2.

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate

arXiv:2608.…

  • 11381v1 Announce Type: new Abstract: We study whether the localized nu…
  • Larix maps a 16-lens European listed-real-estate analysis framework to…
  • Across 19 firms spanning seven regulatory wrappers, decomposition impr…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

arXiv:2608.…

  • 11343v1 Announce Type: new Abstract: Multimodal retrieval and classifi…
  • The March 2026 release of Gemini Embedding 2, Google's first natively …
  • Simultaneously, frontier Large language models (LLMs) have also demons…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

arXiv:2608.…

  • 11354v1 Announce Type: new Abstract: Modern recommender systems treat …
  • As interfaces evolve from static layouts toward generative UIs and imm…
  • We propose an Inverse Theory of Mind (IToM) pipeline that reasons back…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

arXiv:2608.11341v1 Announce Type: new Abstract: Apollo did not reach the Moon merely because its engineers could solve difficult equations.…

  • It succeeded by turning a distant ambition into a mission architecture…
  • AI now faces a similar transition: frontier models can solve difficult…
  • We introduce Apodex Discovery, a framework for building and evaluating…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

arXiv:2608.11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability.…

  • We show, across three open agent-trace benchmarks (TheAgentCompany, $\…
  • Leaderboards rank specialization, not capability.
  • We arrive at this through a four-facet Generalizability Theory varianc…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

arXiv:2608.11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals.…

  • Existing approaches exhibit a "when-what" dissociation: traditional DN…
  • We attribute this to the absence of a unified reasoning paradigm.
  • Inspired by how humans inspect surveillance videos - glancing globally…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Adaptive Hybrid Particle Swarm Optimization with Gradient Descent

arXiv:2608.…

  • 11258v1 Announce Type: new Abstract: Gradient injection helps Particle…
  • We propose Adaptive Hybrid PSO (AHPSO), which uses a sigmoid function …
  • Under budget-normalized comparison (PSO given equivalent total functio…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures

arXiv:2608.…

  • 11255v1 Announce Type: new Abstract: Accurate prediction of vapor--liq…
  • While deep learning models can provide accurate predictions, they ofte…
  • In this work, we propose a symbolic machine learning approach to disco…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning

arXiv:2608.…

  • 11252v1 Announce Type: new Abstract: Agentic AI systems routinely tran…
  • We prove this class of safeguard is structurally incomplete.
  • Modelling a covering of context space by its nerve and evidence by a r…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search

arXiv:2608.…

  • 11250v1 Announce Type: new Abstract: Language models can propose many …
  • We present AgonAlpha, an architecture that searches over frozen resear…
  • To our knowledge, AgonAlpha is the first alpha-mining system to combin…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents

arXiv:2608.11248v1 Announce Type: new Abstract: Long-term memory is essential for language agents operating across extended interactions and evolving tasks.…

  • Existing memory-augmented agents mainly focus on storing and retrievin…
  • In particular, previously distilled insights can become outdated, over…
  • To address this issue, we study insight-level memory maintenance for l…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

arXiv:2608.…

  • 11247v1 Announce Type: new Abstract: Recent advances in language model…
  • Each agent sees what the others assert before it answers, so peer opin…
  • We measure that displacement in 23 open-weight models, 19 conditions, …

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Towards the Harness of Embodied Agents

arXiv:2608.…

  • 11246v1 Announce Type: new Abstract: The success of coding agents has …
  • We ask whether the same paradigm extends to embodied agents in the phy…
  • We present Thea, a harness in which an agentic loop orchestrates robot…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach

arXiv:2608.…

  • 11245v1 Announce Type: new Abstract: Online education offers unprecede…
  • To address these challenges, we introduce AI Tutor, a reinforcement le…
  • In the short term, AI-Tutor draws on cognitive theory to guide learner…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model

arXiv:2608.11244v1 Announce Type: new Abstract: Construction standards are critical for building safety and sustainability.…

  • Existing standard application workflows rely on keyword-based document…
  • To address these limitations, this study develops a multimodal knowled…
  • The framework introduces 1) a multimodal knowledge graph (MKG) for uni…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification

arXiv:2608.…

  • 11243v1 Announce Type: new Abstract: We argue that a single structural…
  • , the agent does not escape its sandbox) is an off-support object.
  • Formally, if q is the data distribution and \(p(\cdot\mid w)\) the mod…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle

arXiv:2608.…

  • 11241v1 Announce Type: new Abstract: Deploying LLM agents into industr…
  • Any two can be maximized against the third.
  • We present RecSys Factory, an LLM-agent platform deployed for 78 days …

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

VQ-bench: A Composable Vector Quantization Framework

arXiv:2608.11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure.…

  • It is therefore experiencing a surge of renewed engineering and resear…
  • This paper provides a unified framework for developing and benchmarkin…
  • We describe 7 common conceptual quantization primitives and show how t…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

arXiv:2608.…

  • 11238v1 Announce Type: new Abstract: Retrieval-augmented generation im…
  • We propose Q-CARE, a query-agnostic and fully reference-free framework…
  • Q-CARE establishes a unified evaluation principle based on query cover…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference

arXiv:2608.11235v1 Announce Type: new Abstract: Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon.…

  • Many predictions stabilize early, but blockwise decoding continues unt…
  • Existing accelerators often rely on learned filters, modified scores, …
  • We ask whether native trajectory signals can identify residual positio…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction

arXiv:2608.11237v1 Announce Type: new Abstract: Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs).…

  • However, long-horizon autoregressive prediction remains challenging: l…
  • Existing methods mainly improve state representations and operator bac…
  • To address these issues, we propose a geometry-aware incremental neura…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity.…

  • Recent advances in AI agents create a timely opportunity to automate i…
  • We present InfraBench, a benchmark suite for evaluating AI agents on r…
  • Experiments with 15 agent-model configurations show that even the stro…

RSS 官方收录 · 可信分层展示

详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC).…

  • Existing PIC methods, however, are built for full-attention models, wh…
  • Hybrid LLMs break these primitives---they replace most attention layer…
  • This raises a natural question: can PIC benefit hybrid models, and wha…

RSS 官方收录 · 可信分层展示

详情 原文 分享图