arXiv:2608.12863v1 Announce Type: new Abstract: As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved.…
This working paper examines whether India's Consumer Protection Act, 2…
The Act's broad definitions of product liability, harm, and deficiency…
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
arXiv:2608.12743v1 Announce Type: new Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants.…
To improve the spatial reasoning ability of VLM agents, existing work …
One line uses post-training methods, such as supervised fine-tuning an…
Another line adopts an agentic paradigm in which the model calls exter…
Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
arXiv:2608.12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers.…
However, with catalogs spanning millions of active items, manual gover…
Inconsistent pricing across item variants distorts customer value perc…
To address this, we present a scalable, context-aware Multi-Agent Fram…
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.…
Because of their ubiquity and computational capability, there is a rap…
In this endeavor, circuit complexity has by and large emerged as the "…
Here, we present an overview of selected results that delineate the ex…
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment.…
The steering decision is the pre-commit choice at that boundary: proce…
We introduce SteerBench-Work, an incident-anchored, bidirectional benc…
Release v2026-05 contains 106 scenarios anchored in public incidents, …
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.…
Judges are typically validated by accuracy on golden data, but accurac…
We introduce the \emph{Wiggle Framework}, a unified stress test for ep…
The framework decomposes judge robustness along three dimensions: Mech…