arXiv · 2609.14864
GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems
Abstract
We predict single-sequence model throughput from GGUF metadata using roofline-shaped predictors with quantization-specific scale factors fitted on reference models. The scored cohort comprises 318 phase-depth measurements from 53 host-file configurations on two Apple M4 Max systems and an NVIDIA RTX 5080. On host-specific held-out sets of four, five, and two configurations, an active-parameter decode model obtains 13.1%, 14.4%, and 36.1% mean absolute percentage error (MAPE), versus 49.4%, 55.3%, and 51.9% when charging total parameters. Leave-one-host-out coefficients fitted on the other two systems yield 11.6%, 16.8%, and 36.0% test MAPE. A low-bit model ladder changes ordering across runtime stacks. The P2 prefill baseline gives 18.7%, 22.2%, and 108.2% test MAPE. GGUF structure helps on all three systems, but fitted efficiencies are not universal.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xinyu Qiu, Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai. 2026-09-14. GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems. https://arxiv.org/abs/2609.14864
Cite the original work for its findings. Save a collection to share your selection of sources.