arXiv Science⌕ Search

arXiv · 2609.37693

FP64 Is All You Want, INT8 Is All You Need, FP4/6/8 Is All You Have

Abstract

Ozaki scheme II emulates FP64 matrix products with INT8 ones through residues modulo pairwise coprime moduli, and variants for FP8 and FP4 have followed. We treat these schemes as one family and pose the choice of a scheme as a combinatorial program that minimizes the number of low-precision GEMMs. Given, for each modulus, a finite set of ways to compute products modulo it from low-precision GEMMs, we find the choice of moduli and ways with the fewest GEMMs, for any format, accumulator and inner dimension, and derive lower bounds on the GEMM count over the whole family. Applied to the formats of current GPUs, the method gives the first FP6 schemes, an FP8 scheme with fewer GEMMs than any previous one, and an FP4 scheme that the bounds show needs the fewest GEMMs of any scheme in the family whose moduli lie in a stated range. Implemented on three Blackwell GPUs, the INT8, FP8 and FP4 schemes run faster than native FP64, up to 83x on B300.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pratyai Mazumder, Alexandru Calotoiu, Torsten Hoefler. 2026-09-29. FP64 Is All You Want, INT8 Is All You Need, FP4/6/8 Is All You Have. https://arxiv.org/abs/2609.37693

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Recolorable Graph Exploration by an Oblivious Agent with Fewer Colors

Recently, Böckenhauer, Frei, Unger, and Wehner (SIROCCO 2023) introduced a novel variant of the graph exploration problem in which a single memoryless agent must visit all nodes of an unknown, undirected, and connected graph before returning to its starting node. Unlike the standard model for mobile agents, edges are not labeled with port numbers. Instead, the agent can color its current node and observe the color of each neighboring node. To move, it specifies a target color and then moves to an adversarially chosen neighbor of that color. Böckenhauer~et al.~analyzed the minimum number of colors required for successful exploration and proposed an elegant algorithm that enables the agent to explore an arbitrary graph using only eight colors. In this paper, we present a novel graph exploration algorithm that requires only six colors. Furthermore, we prove that five colors are sufficient if we consider only a restricted class of graphs, which we call the $φ$-free graphs, a class that includes every graph with maximum degree at most three and every cactus.

cs.DC↗

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines

Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic workflows onto a GPU cluster. Scepsy exploits the insight that, while the end-to-end latency of an agentic workflow is unpredictable, each LLM's fraction of execution time is comparatively stable across requests. Scepsy profiles each LLM under different parallelism degrees and combines the profiles with these fractions into an Aggregate LLM Pipeline, a lightweight throughput and latency predictor for allocations. To minimize latency at a target throughput, Scepsy uses the Aggregate LLM Pipeline to search over fractional GPU shares, tensor parallelism degrees, and replica counts. A hierarchical heuristic then places the chosen allocation onto the cluster, minimizing fragmentation and respecting network topology. On realistic agentic workflows, Scepsy achieves up to 2.5x higher throughput before saturation and 1.0-3.3x lower latency than systems that optimize LLMs independently or rely on user-specified allocations.

cs.DC↗

The World's Fastest Matching Engine Algorithm

We drove 247 matching engines through one C ABI harness on one identical workload: every open-source FIFO implementation we could find, deduplicated, and our own, on the same gate. The workload doubles as a byte-identical correctness oracle, 1,000,000,000+ order messages per engine, replayed against an independent-engine consensus. Only 47 are correct as shipped; we filed 181 GitHub issues upstream, 28 already fixed by their maintainers, none declined. Our engine leads the 160 that reproduce the consensus by ~95 M/s (12.6x the second best) on worst-case throughput. One core sustains 103.4 million order messages per second (122.09 million on AMD EPYC processors) worst-case, and in the engine's production configuration its wire pipeline keeps single-book P99 host-path latency, OUCH parsing and OUCH/ITCH encoding included, under a microsecond through 82 M msgs/s; a 96-core server (~$1,630/month, 3-year reserved) reaches ~1.3 billion/s across over 10,000 symbols, for scale, over 45x the CTA consolidated quote feed's provisioned capacity. The lead is structural: the 73 engines written inside the trading industry sit under the same 8.19 M/s ceiling as the rest of the field, and it is one of them that sets it.

cs.DC↗