arXiv · 2604.13349
When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration
Abstract
Multi-agent LLM systems are moving beyond discrete-token messages toward richer relays that preserve internal state. Recent work such as LatentMAS transmits full key-value (KV) caches between agents but pays a high memory and communication cost. We adapt KV-cache eviction to this setting and introduce \textbf{Orthogonal BackFill (OBF)}, which injects a low-rank residual from the discarded KV states back into the retained ones, orthogonal to what is already kept. With only $9.9\%$-$20.2\%$ of the prompt KV retained, compressed relay cuts bandwidth by $4.7\times$ and GPU memory by $8\%$ at under $5\%$ wall-clock overhead, and stays close to full relay in accuracy across nine benchmarks, ahead of it on several. OBF matches or improves over headwise eviction on all nine, and its gain is proportional to the accuracy gap eviction opens against full relay ($r{=}0.78$ across three model scales), so it gives back part of what eviction takes. Code is available at https://github.com/markli404/When-Less-Latent-Leads-to-Better-Relay.
Explore related subjects
Keep this discovery
Yiping Li, Zhiyu An, Wan Du. 2026-04-14. When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration. https://arxiv.org/abs/2604.13349
Cite the original work for its findings. Save a collection to share your selection of sources.