微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接后用系统浏览器访问。
XMT
信流 · 上滑连读 · 来源可核
Canada Needs Industrial Policy That Lets The Market Say No
Mark Carney’s industrial-policy agenda has reopened a debate Canada never really escaped: when should government build, finance or shape productive capacity, and when should it leave the result to markets?…
- The proposed C$25 billion Canada Strong Fund makes the question concre…
- [continued] The post Canada Needs Industrial Policy That Lets The Mark…
RSS 官方收录 · 可信分层展示
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning
arXiv:2608.…
- 20717v1 Announce Type: new Abstract: Reliable confidence estimation is…
- When the same problem is queried under multiple confidence-steering pr…
- Existing black-box uncertainty methods often rely on answer agreement,…
RSS 官方收录 · 可信分层展示
VortexChat: An agentic framework for autonomous multi-objective integrated photonic design
arXiv:2608.…
- 20688v1 Announce Type: new Abstract: The advancement of modern integra…
- While inverse design offers an alternative, it remains constrained by …
- To address these issues, we present VortexChat, an agentic framework f…
RSS 官方收录 · 可信分层展示
CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery
arXiv:2608.20686v1 Announce Type: new Abstract: Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints.…
- Reinforcement learning (RL) offers a promising approach, but existing …
- We introduce Certification-Driven Reinforcement Learning (CDRL), a fra…
- When a candidate violates domain constraints, these tools produce cert…
RSS 官方收录 · 可信分层展示
Why2Speak: Faithful Reasoning for Abstaining Action Policies
arXiv:2608.…
- 20670v1 Announce Type: new Abstract: Many agentic systems must repeate…
- We study this problem through intervention timing in multi-party conve…
- This setting exposes class imbalance, asymmetric action costs, and the…
RSS 官方收录 · 可信分层展示
DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents
arXiv:2608.…
- 20664v1 Announce Type: new Abstract: DreamBench-SWE is a multi-session…
- We report the original scaled v2 fold and a separately preregistered v2.
- 1 successor audit designed after that study but frozen before successo…
RSS 官方收录 · 可信分层展示
Pokémon Unveils 30th Anniversary Pikachu Figurine Set
SummaryPokémon has officially announced the 30th Anniversary Pikachu Figurine Set from its Dream Art seriesEach blind-box set includes a 3D Pikachu figure a matching foil trading card and a six-card 30th Anniversary Celebration booster packThe collection features 13 total figurine variations including a rare hidden secret designPokémon has officially announced the 30th Anniversary Pikachu Figurine Set as part of its ongoing "Dream Art" series.…
- The commemorative release celebrates three decades of the franchise by…
- Designed to recreate the color layers and brushstrokes of card illustr…
- Each box contains one random Pikachu figure alongside its matching foi…
RSS 官方收录 · 可信分层展示
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance
arXiv:2608.…
- 20661v1 Announce Type: new Abstract: Enterprise adoption of large lang…
- This paper argues that retrieval-augmented generation for enterprise f…
- An evaluation on FinanceBench (145 questions) compares KDAF against ze…
RSS 官方收录 · 可信分层展示
Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions
arXiv:2608.…
- 20649v1 Announce Type: new Abstract: Designers and policymakers in soc…
- Prior work tends to evaluate interventions individually and mostly alo…
- We present a multi-criteria framework for evaluating sociotechnical in…
RSS 官方收录 · 可信分层展示
SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL
arXiv:2608.…
- 20630v1 Announce Type: new Abstract: SQL systems increasingly expose A…
- Despite their diverse APIs, these functions play only three relational…
- We present SAGE (Self-Adaptive Generative Execution), a unified logica…
RSS 官方收录 · 可信分层展示
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents
arXiv:2608.…
- 20631v1 Announce Type: new Abstract: Large language model (LLM) agents…
- Existing memory approaches organize or compress execution histories bu…
- We introduce the, a hierarchical memory system that organizes executio…
RSS 官方收录 · 可信分层展示
Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work
arXiv:2608.…
- 20622v1 Announce Type: new Abstract: Frontier models have collapsed th…
- The cost of reviewing and maintaining that code hasn't collapsed.
- Each solution drifts from the next; understanding one means reading it…
RSS 官方收录 · 可信分层展示
Dual-Cache Latent Space Communication between Heterogeneous Language Models
arXiv:2608.…
- 20617v1 Announce Type: new Abstract: Multi-agent LLM systems split wor…
- They usually communicate by exchanging text, which puts autoregressive…
- Recent latent protocols instead translate the sharer's key-value (KV) …
RSS 官方收录 · 可信分层展示
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
arXiv:2608.…
- 20614v1 Announce Type: new Abstract: Enterprise agent programs are mov…
- Current gates often scan these artifacts for structure, style, and sec…
- We present ACES (Agentic Continuous Evaluation of Skills), a repositor…
RSS 官方收录 · 可信分层展示
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
arXiv:2608.…
- 20611v1 Announce Type: new Abstract: Semantic-ID-based generative reco…
- A common recipe is SFT followed by GRPO, yet vanilla GRPO is poorly ma…
- Under the frozen SFT checkpoint, the exact target is absent from the f…
RSS 官方收录 · 可信分层展示
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
arXiv:2608.20574v1 Announce Type: new Abstract: Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key.…
- We introduce FlavourBench, an automated benchmark in which a versioned…
- Each task presents eight ingredients and asks for a three-ingredient p…
- We evaluate 27 frontier endpoints on an identical 534-task core spanni…
RSS 官方收录 · 可信分层展示
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states?…
- Recent work suggests that under certain conditions a complex enough mo…
- We tested that claim on eight open-weight models from seven families a…
- To test it we built Open-Weight Masked Introspection (OWMI), a framewo…
RSS 官方收录 · 可信分层展示