arXiv · 2609.27926
The Joule Point: an Energy-Optimal Operating Point for AI Inference
Abstract
Data centers serving AI inference strive to maximize GPU utilization, running their cards at full power by default: this maximizes throughput and holds latencies down, but it is energy-inefficient, spending more energy per inference than the same work needs at a better operating point. That inefficiency is a scheduling choice, not a hardware limit. The cause is physical: for a given workload, a GPU's board power rises superlinearly along its power-performance curve, on top of a fixed power floor that is paid for as long as the job runs, so energy per inference is U-shaped in the operating point (GPU, power cap). The minimum, which we name the Joule Point, is a cap at 43 to 46 per cent of a large GPU's rated power; capping to it cuts energy per inference by roughly a quarter to a third at a modest cost in capital and latency: each request runs about 1.2 times slower, and holding aggregate throughput takes that same factor more cards. Under load, the Joule Point is nearly a per-card-type constant, so a single static cap per card type captures nearly all the saving (mean penalty under one per cent), turning the online per-job search that prior systems run into a one-time characterization. We ground this in ELF, a dense power-cap dataset that sweeps 20 inference models across four GPUs; replaying it in a fleet simulation under a power budget, capping each job to the least power meeting its deadline spends 18 to 45 per cent less energy per served job. These results recast data-center energy as a schedulable resource an operator can adapt to a budget, an electricity price, or a carbon signal while meeting its service targets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alexander Apartsin, Yehudit Aperstein. 2026-08-23. The Joule Point: an Energy-Optimal Operating Point for AI Inference. https://arxiv.org/abs/2609.27926
Cite the original work for its findings. Save a collection to share your selection of sources.