Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to keep GPU memory constant for a given trace, but most strategies rely on prefilling the LLM context many times over, which hinders training throughput. This paper proposes KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance.

Streaming the KV cache forward

KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. In typical compaction, the cache is discarded and rebuilt from scratch, requiring the LLM to re-read and re-process all previous tokens. KV-streams keep the cache alive across compaction steps, streaming it forward through the model. This eliminates the redundant prefilling overhead while maintaining the same memory footprint.

Three compaction strategies enabled

The authors show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. The strategies differ in how aggressively they prune the cached key-value pairs, but all benefit from the streaming approach that avoids repeated context reconstruction. The speedup range depends on the task and model scale, with larger contexts seeing greater relative benefit.

Streamed KV cache as recurrent state

Beyond efficiency, the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the active context. In a controlled setting, the authors show that RL alone is all that is needed for this behavior to emerge—contrary to prior work that suggested additional supervision or architectural modifications were required. The cached activations encode task-relevant state that the model can reuse across episodes or time steps.

Plug-and-play post-training addition

KV-streams are compatible with any existing compaction strategy and require no model retraining, no architectural changes, and no additional data. The approach is lightweight enough to integrate into any post-training pipeline, making it immediately applicable to deployed agentic LLM systems. The code and configurations are available for download, enabling researchers and engineers to experiment with compaction throughput improvements right away.

Read the paper on arXiv