arXiv Science⌕ Search

arXiv · 2610.11201

PageWeaver: KV-Guided Query Unions for Sparse Attention

Abstract

Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work. Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together. We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership. A bounded GPU search produces query IDs, and an ID-aware two-CTA kernel consumes them without materializing reordered Q tensors or cross-page partial outputs. A direct KV-page union implementation provides a complementary design study of nonlocal reuse and reduction cost. With FP8 KV throughout, the H200 Union8 implementation achieves a 1.70x geometric-mean complete-call speedup over the measured FlashInfer path on six captures. Online regrouping further lowers latency by 3.26-7.66% on five selected 64K-context captures. Whole-model prefill throughput is 7.88-14.36% above the tested native path; the incremental regrouping benefit is smaller, with observed median gains of 0.47-0.73% at 32K/64K and regressions at 8K. A B300 comparison identifies cases where preparation cost and a stronger native kernel remove the advantage. These results separate execution-group reuse from the complete cost of exploiting it online.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhiyuan Li, Zihan Li, Zefang Yuan, Lei Wang, Hao Wang. 2026-10-08. PageWeaver: KV-Guided Query Unions for Sparse Attention. https://arxiv.org/abs/2610.11201

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

eAVID: Asynchronous Verifiable Information Dispersal with Post-Dissemination Pruning

The well-known Asynchronous Verifiable Information Dispersal (AVID) problem lets a sender disperse a message across $N=3F+1$ nodes such that it remains recoverable despite up to $F$ Byzantine failures and unbounded message delays. An optimal AVID scheme requires $3\times$ the message size in communication and storage: up to $F$ nodes may be delayed in responding, and among the remaining $2F+1$ respondents, up to $F$ may be Byzantine. This paper presents eAVID, an elastic AVID scheme that provides two key improvements. First, after dissemination, nodes can prune up to half of their stored information without requiring anyone to reconstruct or recode the message. Second, pruning is enabled by an extremely simple protocol. Nodes collect acknowledgments from one another confirming receipt of their coded pieces, after which each node locally prunes its stored information. eAVID achieves this $2\times$ reduction while making two tradeoffs: the sender generates $2N$ fragments rather than $N$, and pruning may necessitate contacting $2F+1$ nodes for message retrieval rather than $F+1$. eAVID was implemented in DispersedSimplex, a simple-to-understand BFT consensus protocol that erasure-codes its blocks. The implementation demonstrates two features. First, eAVID achieves these improvements using a flat erasure-coding scheme that requires no metadata or bookkeeping at the nodes during reconstruction. Second, it delivers the storage savings without a performance cost: with all nodes responsive, pruning fires on over $99\%$ of blocks and steady-state per-node storage falls by $42$-$46\%$ for committees of $10$ to $22$ nodes. Throughput and latency match the unmodified protocol showing that we can achieve these storage savings without any performance costs.

cs.DC↗

IAPRepair: In-Network Aggregation Enhanced Proactive Repair for Erasure-Coded Storage System

Erasure-coded storage provides fault tolerance with substantially lower storage overhead than full replication, but repairing lost or at-risk blocks requires intensive cross-node data transfer. Existing reactive repair schemes start only after a failure, while proactive schemes can move data before failure but commonly treat migration, reconstruction, and network aggregation as loosely coupled operations. As a result, receiver?side bottlenecks, heterogeneous available bandwidth, and limited programmable-switch state continue to constrain repair paral?lelism. This paper presents IAPRepair, an in-network aggre?gation enhanced proactive repair framework for erasure-coded storage. IAPRepair jointly constructs each repair batch, assigns reconstruction providers and replacement nodes according to normalized transmission loads, schedules migration around the remaining receive capacity, and selectively enables in-network aggregation for reconstruction blocks that would otherwise over?load healthy nodes. The selective design reduces receiver-side traffic while retaining migration parallelism and respecting a configurable switch-resource budget. We implement IAPRepair with a Tofino programmable switch and 16 storage nodes, and evaluate it using both a prototype testbed and large-scale simu?lations. Across coding parameters, block sizes, node populations, and multiple STF-node scenarios, IAPRepair reduces repair time by at least 47.83% compared with the evaluated state-of-the-art methods. The results demonstrate that coordinating proactive repair decisions with in-network processing is an effective way to improve repair efficiency under bandwidth heterogeneity.

cs.DC↗

Zepp: Accelerating Distributed MoE Serving under Relaxed Balance Constraints

As Mixture-of-Experts (MoE) models continue to scale, serving them increasingly relies on expert parallelism (EP) across a growing number of devices. Yet skewed expert workloads create imbalance across computation, communication, and memory, making load balancing a central optimization objective in distributed MoE serving. We observe that balance is not free: operations introduced to balance one dimension can themselves be expensive or imbalanced. This motivates us to rethink balance as a constraint rather than an optimization objective. We present Zepp, which directly optimizes the bottleneck communication in distributed MoE serving subject to simplified balance constraints on physical resources, i.e., GPUs and NICs. Zepp progressively optimizes inter-node communication across placement, routing, and execution. It first places expert replicas to reduce token communication under GPU constraints, then reshapes communication flows through split and merge primitives under NIC constraints, and finally partitions and schedules ex- pert computation to overlap the resulting communication. To adapt to dynamic workloads, Zepp jointly coordinates computation, token communication, and expert-weight movement at each iteration. Together, these designs allow Zepp to pursue the most efficient execution rather than a single-dimension balanced one. We implement Zepp and evaluate it against 7 state-of-the-art MoE serving systems, achieving up to 6.68$\times$ MoE layer speedup and a geometric mean speedup of 1.86$\times$ over the fastest competing baseline.

cs.DC↗