arXiv Science⌕ Search

arXiv · 2610.11284

DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution

Abstract

Diffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation. However, their parallel iterative refinement mismatches AR accelerators optimized for sequential decoding and their discrete token generation differs from DiT accelerators designed for continuous denoising. Recent dLLM accelerators have explored workload-specific optimizations to reduce vocabulary processing overhead and redundant computation across denoising iterations. However, these approaches retain all tokens in parallel execution, despite varying token refinement utility and execution requirements. This paper presents DynaTE, a hardware--software co-design architecture that dynamically adapts accelerator execution to evolving token states during dLLM decoding. DynaTE first enables adaptive token execution by skipping low-utility token computation, while a dimension-reconfigurable PE array maintains high utilization under varying active-token patterns. Second, DynaTE exploits dynamic token dependencies through FLDD to refine a small number of locally dependent tokens within the current iteration, reducing the overall number of denoising iterations, while a Merge--Split--Merge dataflow hides the resulting serial overhead. Third, a streaming vocabulary engine interleaves multiple token streams from the LM head to accommodate irregular output variations caused by selective token computation and uneven vocabulary-selection demands. Evaluated on two representative dLLMs, DynaTE achieves 2.05--2.78$\times$ speedup and 2.99--3.93$\times$ higher energy efficiency over state-of-the-art dLLM accelerators, while delivering 2.55$\times$ speedup and 6.07$\times$ higher energy efficiency over Jetson AGX Orin.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Minghan Jiang, Jiayi Wang, Shuaiting Li, Haibin Shen, Kejie Huang. 2026-10-08. DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution. https://arxiv.org/abs/2610.11284

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration

The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerators are strictly constrained by area and power, they require end-to-end quantized models. However, the extreme dynamic range gap between multi-modal tokens causes standard block formats to suffer "microscaling collapse," where a single massive outlier hijacks the shared exponent, underflowing surrounding elements and destroying attention maps. To break this bottleneck, we propose Micro-Inverted-Scaling (MiX), a novel format that mathematically inverts the microscaling paradigm: rather than grouping multiple mantissas under one shared exponent, MiX groups private, per-element exponents under a single shared mantissa. To handle asymmetric VLM outlier topologies, we introduce an adaptive dual-format (MiX-MX) inference framework. By algebraically factoring out the shared MiX mantissa, this framework maps to a custom accelerator, replacing multipliers with efficient shifters. Evaluated end-to-end on multiple VLMs, our 4.5-bit MiX formulation exhibits equivalent or superior accuracy on multi-modal benchmarks compared to NVFP4. Simultaneously, the MiX accelerator delivers a 25% improvement in area efficiency over the NVFP4 baseline and a 2.3-4.5x speedup with 1.4-2.9x energy reduction across models compared to the state-of-the-art accelerator Focus, proving the inverted-scaling datapath is physically superior for efficient VLM deployment.

cs.AR↗

Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing

LLM-based agents are increasingly applied to the "last mile" of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC) violations and converging Power-Performance-Area (PPA) targets after tool runs. Existing EDA-LLM benchmarks, however, omit DRC fixing entirely and rely on flat hierarchies tied to a single toolchain. We introduce PostEDA-Bench, a hierarchical benchmark with 145 tasks across DRC-Essential, DRC-Reasoning, PPA-Mono, and PPA-Multi, supported by EDA toolchains with machine-checkable evaluation. Across eight commercial and open-source LLMs under multiple agent scaffolds, we find that agents handle synthetic DRC-Essential and single-objective PPA-Mono reasonably well but degrade sharply on the more practical DRC-Reasoning, where the best success rate is 36.66%, and PPA-Multi, where the best success rate is 20.00%; vision augmentation consistently enhances DRC-Bench; and trade-off reasoning, rather than knob knowledge, is the dominant PPA-Multi bottleneck.

cs.AR↗

Predictive Software Scheduling as an Early-Warning Hint Layer for Optical Engine Thermal Drift in Heterogeneous SoIC Packaging

As semiconductor scaling approaches the A16 / 2 nm node, the integration of co-packaged optics (CPO) through TSMC's Compact Universal Photonic Engine (COUPE) architecture introduces critical thermal-optical coupling challenges. Micro-ring resonators embedded in the Photonic Integrated Circuit (PIC) layer are highly sensitive to temperature, with a wafer-level center wavelength deviation of merely +/-1.7 nm across a 300 mm wafer representing the platform's manufacturing control limit. To address this, we propose XRM-SSD V24, a physics-aware scheduling layer that models inference-load density 20-50 ms before execution and issues early-warning hints to the COUPE bias-control firmware, enabling pre-emptive thermal compensation. Simulation-based validation on a software emulation platform (physical characterization pending TSMC tape-out) over 90,000 inference steps yields a simulator-internal thermal-load correlation of R^2 = 0.9911 across a workload density range of pv24 in [0.9, 2.7] (a 3x span), with wavelength drift below 0.354 nm - equivalent to 21% of the +/-1.7 nm wafer-level wavelength control budget and 71% of the tighter +/-0.5 nm per-channel spectral specification. A full Thermal Resistance Fingerprint characterization further confirms Rth = 0.45 deg C/W, a thermal time constant tau = 80 ms, and a thermo-optic coefficient of 0.0852 nm/deg C across five discrete load states (Idle to Peak). Memory stability is reported as zero leakage in the current simulation run; long-duration soak testing to confirm sustained stability remains future work. We establish a formal domain separation between deterministic software scheduling and continuous physical thermal dynamics, ensuring physics-consistent claims suitable for peer review.

cs.AR↗