KV-streams for Efficient Compaction in Agentic Reinforcement Learning

Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.


KV-streams for Efficient Compaction in Agentic Reinforcement Learning - Technical Figure
Figure 1: Architectural diagram and empirical setup from the original research paper (arXiv:2609.35750v1).

2. Document Verification & Archival Data

  • Contributing Researchers: Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessandro Sordoni, Minseon Kim, Marc-Alexandre Côté, Laurent Charlin, Guillaume Lajoie
  • Submission Date: September 28, 2026
  • Full Preprint Document: Download Official PDF
  • Permanent Archive Record: arXiv:2609.35750v1

3. Academic Citation Reference

Standard Reference (APA Format):

Emiliano Penaloza, et al. (2026). KV-streams for Efficient Compaction in Agentic Reinforcement Learning. arXiv:2609.35750v1. https://arxiv.org/abs/2609.35750v1

Academic Field: Machine Learning | Document Identifier: arXiv:2609.35750v1


BibTeX Entry:

Kod
@article{arxiv_2609.35750v1,
  author    = {Emiliano Penaloza and Dane Malenfant and Dheeraj Vattikonda and Roger Creus Castanyer and Siddarth Venkatraman and Abhay Puri and Jonathan Light and Matthew James Sargent and Augustine N. Mavor-Parker and Massimo Caccia and Lucas Caccia and Glen Berseth and Esmeralda S. Whitammer and Alessandro Sordoni and Minseon Kim and Marc-Alexandre Côté and Laurent Charlin and Guillaume Lajoie},
  title     = {{KV-streams for Efficient Compaction in Agentic Reinforcement Learning}},
  journal   = {arXiv preprint arXiv:2609.35750v1},
  year      = {2026},
  url       = {https://arxiv.org/abs/2609.35750v1}
}

Yorumlar (0)

Henüz yorum yapılmamış. İlk yorumu siz yapın!

Yorum Bırakın