arXiv · 2503.09980
Thermodynamic cost of inference and learning in physical neural networks
Abstract
How much of the energy consumed by artificial neural networks is set by physics rather than by implementation? For irreversible digital hardware the reference is Landauer's principle, which charges $k_B T\ln 2$ per erased bit. We map a generic feedforward network onto a physical Hamiltonian in which each layer relation is an elastic compatibility constraint, and obtain two exact bounds. First, its equilibrium free energy is independent of the input and of every weight and bias, at all temperatures, so quasi-static inference requires no work whatsoever: no thermodynamic cost attaches to computation itself. Second, at finite speed the work exceeds the squared Wasserstein-2 distance between the initial and final thermal states, divided by the protocol duration. Relaxed to an entropic measure of distinguishability, this identifies the cost with the information separating successive inputs, about $k_B T$ per dimension of the widest layer at the fastest usable speed. Learning is fundamentally different: writing the parameters carries an irreducible cost of a few $k_B T$ each that survives the quasi-static limit. The thermodynamic price of a neural network is therefore set by its memory rather than its arithmetic. Simulations confirm both bounds: the inference work saturates the transport bound to within four percent, and accuracy collapses once the dissipated work falls below the thermal scale.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alexei V. Tkachenko. 2025-03-13. Thermodynamic cost of inference and learning in physical neural networks. https://arxiv.org/abs/2503.09980
Cite the original work for its findings. Save a collection to share your selection of sources.