arXiv · 2601.23278
FOCUS: DLLMs Know How to Tame Their Compute Bound
Abstract
Diffusion Large Language Models (DLLMs) offer a compelling alternative to Auto-Regressive models, but their deployment is constrained by high decoding cost. In this work, we identify a key inefficiency in DLLM decoding: while computation is parallelized over token blocks, only a small subset of tokens is decodable at each diffusion step, causing most compute to be wasted on non-decodable tokens. We further observe a strong correlation between attention-derived token importance and token-wise decoding probability. Based on this insight, we propose FOCUS, an inference system designed for DLLMs. By dynamically focusing computation on decodable tokens and evicting non-decodable ones on-the-fly, FOCUS increases the effective batch size, alleviating compute limitations and enabling scalable throughput. Empirical evaluations demonstrate that FOCUS achieves up to 3.52$\times$ throughput improvement over the production-grade engine LMDeploy in large-batch settings, while preserving or improving generation quality across multiple benchmarks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kaihua Liang, Xin Tan, An Zhong, Hong Xu, Marco Canini. 2026-01-30. FOCUS: DLLMs Know How to Tame Their Compute Bound. https://arxiv.org/abs/2601.23278
Cite the original work for its findings. Save a collection to share your selection of sources.