arXiv · 2609.12551
RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems
Abstract
AI is beginning to make substantive contributions to LLM inference optimization. Existing AI optimizations are predominantly profiling-based. Profiling-bound feedback confines the search to the capabilities and performance of an existing software stack, preventing a fundamentally better architecture of LLM inference systems from being identified. To enable the AI-driven LLM inference system architecting loop, we argue that a general workload representation, a verifiable mutation space, and an implementation-independent evaluator are required. We present the RoofLang domain-specific language (DSL) that provides these features. In our evaluation, RoofLang reveals that DeepSeek V4-series models could achieve 3.5-39.5$\times$ higher peak decode throughput than other representative models. This gap is disproportionate to their total parameter counts and arises largely from compact KV-cache designs that support larger batches and reduce memory traffic. A persistent optimizer agent further discovered several new architectures that improved both throughput and interactivity of DeepSeek V4 Pro on NVIDIA B300 by 6.23-50.1%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ziyue Yang, Yuting Jiang, Lei Qu, Peng Cheng. 2026-09-11. RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems. https://arxiv.org/abs/2609.12551
Cite the original work for its findings. Save a collection to share your selection of sources.