Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating

2026-07-27Artificial Intelligence

Artificial Intelligence
AI summary

The authors study how language models decide which pieces of information to keep in their limited memory. Traditional methods make decisions immediately when new information arrives or try to guess the future use of that information. They introduce a new approach, called fixed-lag smoothing, that waits a little before deciding, letting the model show which information is actually useful. Their approach, RMM, works better than older methods in controlled tests but does not improve much on standard benchmarks because typical text doesn’t clearly show which memory is truly useful. The main contribution is a new way to think about memory decisions, showing when waiting to decide helps or doesn’t.

language modelworking memoryfixed-lag smoothingcommit lagattentionBelady's algorithmRMMmemory managementstreaming modelsquestion answering
Authors
Maruthi Vemula, Neeraj Praneeth Gajula
Abstract
A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag $H$: online filters and learned predictors commit at $H=0$, while Belady's offline optimum sits where the whole future is known. The missing regime in between, fixed-lag smoothing, waits a bounded number of steps, observes which items a correct near-future prediction attended to, and only then commits. This measurement, demonstrated utility, turns Belady's unobservable future request into something we read off the model itself. We instantiate it as a training-free policy, RMM, a strict generalization of H2O that reduces to it exactly when the measurement is uniform. In controlled settings where reuse is endogenous and separated in time, demonstrated utility identifies used memory far better than accumulated attention, and a small bounded memory behaves like a much larger one. But on independent third-party benchmarks, run inside NVIDIA's KVPress harness against its own SnapKV, H2O, and StreamingLLM implementations, the advantage mostly disappears: RMM is on par with H2O for single-turn question answering and loses to both H2O and SnapKV in a streaming multi-turn setting. The cause is simple: on natural text the model is correct about most tokens, so weighting attention by correctness barely changes it, and demonstrated utility collapses onto accumulated attention unless reuse is sharp and endogenous, which standard benchmarks do not exercise. Our contribution is the framework and an honest map of when measuring beats accumulating, not a new state of the art.