arXiv Science⌕ Search

arXiv · 2610.05059

You Only Convert Twice: Ozaki Scheme II for Chained Tensor Mode Products

Abstract

This work proposes OzII-RescaleBE, a method that enables persistent execution of chained tensor mode products in Ozaki scheme II, keeping the intermediate tensors in residue form. By performing rescaling and base extension directly on the residue representation, it requires conversion to residues only at the beginning of the chain and reconstruction only at the end. An implementation using CuTe DSL and INT8 Tensor Cores is provided for verification and evaluation. The method is compared with per-mode composition of Ozaki scheme~II (GEMMul8-Composed) and a dense Kronecker-product formulation (GEMMul8-KRON), both implemented using the GEMMul8 library. Across three datasets with different input distributions, OzII-RescaleBE achieves normwise relative errors ranging from $2.5\times10^{-16}$ to $2.6\times10^{-15}$ for chain depths $d=3$ through $7$. The corresponding errors range from $10^{-16}$ to $10^{-15}$ for GEMMul8-KRON and from $1\times10^{-16}$ to $3\times10^{-16}$ for GEMMul8-Composed and the FP64 chain. At $d=8$, the errors of OzII-RescaleBE increase to between $7\times10^{-14}$ and $6\times10^{-13}$. For mode sizes $n=8$ to $256$, OzII-RescaleBE achieves throughputs ranging from 1.47 to 4.42\,GDoF/s. Its throughput is within 12\,\% of GEMMul8-KRON at $n=8$ and $12$ and exceeds it at larger tested sizes. Compared with GEMMul8-Composed, OzII-RescaleBE achieves $2.1$--$60\times$ the throughput for $n\leq128$, but has 6--12\,\% lower throughput at $n=256$. It also uses less device memory than the GEMMul8 baselines in the measured memory comparisons. These results demonstrate that rescaling and base extension enable accurate residue-domain tensor chains while avoiding repeated intermediate conversions to floating point.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xuanzhengbo Ren. 2026-10-04. You Only Convert Twice: Ozaki Scheme II for Chained Tensor Mode Products. https://arxiv.org/abs/2610.05059

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

VeriNC: Finding Design Risks of In-Network Computing Systems

The emergence of programmable switches has brought in-network computing (INC) into the spotlight in recent years. By offloading computation directly onto the data transmission process, INC improves network utilization, reduces latency to sub-RTT levels, saves link bandwidth, and maintains throughput. However, INC disrupts the transparency of traditional networks, forcing developers to consider network exceptions like packet loss and out-of-order. If not properly handled, these exceptions can lead to violations of application properties, such as cache consistency and lock exclusion. Usual testing cannot exhaustively cover these exceptions, raising doubts about the correctness of INC systems and hindering their deployment in the industry. This paper presents VeriNC, the first general-purpose tool for verifying INC systems. VeriNC provides a high-level specification language and saves developers 67.2% lines of code on average. To help better understand the behavior of the system, VeriNC offers configurable network environments. VeriNC enables developers to express INC-specific correctness properties. VeriNC translates developer-specified systems into state transition representations, performs model checking to detect potential design risks, and reports violation traces to developers. We propose optimizations for INC-specific scenarios to address the challenge of state space explosion. We modeled INC systems across four application domains and identified design risks with VeriNC in seconds. VeriNC has also been adopted to guide the design of a new INC protocol. Based on our verification experience, we summarize lessons that help develop a correct INC protocol. We further reproduce them in real systems to confirm the validity of our verification result.

cs.DC↗

TorchGWAS 1.0: GPU-accelerated GWAS at scale

Imaging, molecular, and machine-learning workflows can generate thousands of quantitative phenotypes in a single cohort, creating substantial computational and output bottlenecks when testing traits individually. TorchGWAS is a GPU-accelerated framework that uses batched operations for high-throughput, covariate-adjusted linear association testing across large panels of quantitative phenotypes. Across 500,036 allele-harmonized tests, TorchGWAS t statistics agreed with PLINK 2.0. On an NVIDIA H100 80-GB GPU with a 48-core Intel Xeon Gold 6442Y host and measured disk read and write rates of 5.98 and 1.49 GB/s, respectively, median end-to-end times for 4.57 billion associations (8,931,083 variants by 512 phenotypes in 35,365 samples) were 28.46 s for BED, 29.13 s for hard-call PGEN, 51.48 s for BGEN, and 58.95 s for dosage PGEN, including writing 36.7 GB of binary summary statistics. TorchGWAS provides an efficient Python-based framework for parallel fixed-effect association screening at biobank scale.TorchGWAS is implemented in Python and distributed as a documented source repository at https://github.com/ZhiGroup/TorchGWAS.

cs.DC↗

AI-Assisted Computational Reproducibility on the FABRIC Testbed

Computational reproducibility remains difficult despite being central to scientific research. In this paper, we show how the international FABRIC testbed, combined with a large language model (LLM) coding agent through LoomAI, can simplify reproducing published experiments across multiple domains. We reproduced three case studies on FABRIC, covering BBR-family congestion-control evaluations, LAMMPS molecular dynamics scaling benchmarks on a CPU-only MPI cluster, and stress protein homeostasis genomics pipelines. Rather than focusing only on matching numerical outputs, we evaluate whether the reproduced experiments support the same scientific conclusions as the original studies. The AI assistant was effective in setting up the environment, adapting code, and debugging, but struggled with the analysis stages that lacked clearly defined workflows, which required human guidance to establish execution order and data dependencies. Across the case studies, the AI-assisted workflow reduced reproduction effort by roughly 4--6 times. We conclude with practical recommendations for improving AI-assisted reproducibility on research testbeds.

cs.DC↗