arXiv · 2610.02713
WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds
Abstract
Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime combinations and find that most heads change their reading behavior at least once during generation. We introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir rather than freezing or permanently evicting their state. At matched memory or budget, WakeKV consistently improves miss rate over frozen classification and destructive eviction, evaluated across five model-regime combinations and over three cited baselines (SnapKV, uniform R-KV, and ReasonAlloc) across four eligible combinations. A FlexiCache/vLLM implementation on Mistral-7B confirms the benefit on real hardware, improving throughput while retaining LongBench quality.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Utkarsh Ranjan. 2026-10-02. WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds. https://arxiv.org/abs/2610.02713
Cite the original work for its findings. Save a collection to share your selection of sources.