微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
MoNe: Modular Neural Memory for Efficient Long Context Inference
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2608.…
- 17616v1 Announce Type: new Abstract: We present MoNe, a lightweight mo…
- MoNe reads context in fixed-size segments via test-time learning of fa…
- This two-phase design decouples inference cost from context length, ac…
摘要引擎:抽取
正文提要
arXiv:2608.17616v1 Announce Type: new Abstract: We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.