arXiv · 2601.21647
ILRR: Inference-Time Steering Method for Masked Diffusion Language Models
Abstract
Discrete Diffusion Language Models (DLMs) offer a promising non-autoregressive alternative for text generation, yet effective mechanisms for inference-time control remain relatively underexplored. Existing approaches include sampling-level guidance or trajectory optimization mechanisms. In this work, we study the paradigm of reference-based latent steering for DLMs. We introduce Iterative Latent Representation Refinement (ILRR), an efficient framework for steering DLMs using a reference text as a high-level semantic blueprint. ILRR extracts and injects reference-derived semantic signals into the evolving activations of the generated sequence, enabling tunable transfer of coarse properties such as sentiment. We further introduce Spatially Modulated Steering, an extension that enables long-form generation to be guided by shorter references by regulating intensity across the sequence. Empirically, we demonstrate that ILRR achieves effective control on LLaDA and MDLM architectures with low computational overhead, requiring only one additional parallel forward pass per denoising step. Under comparable compute budgets, ILRR improves attribute accuracy over baselines by 10% to 60% points. Our results suggest that the iterative, global denoising process makes DLMs a natural substrate for effective sequence-wide activation-level control.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Eden Avrahami, Eliya Nachmani. 2026-09-20. ILRR: Inference-Time Steering Method for Masked Diffusion Language Models. https://arxiv.org/abs/2601.21647
Cite the original work for its findings. Save a collection to share your selection of sources.