arXiv Science⌕ Search

arXiv · 2609.40239

Two bits about lossy compression: On the limits of compression in cosmology

Abstract

Astronomy is in an era of enormous sky surveys and of the large simulation suites needed to interpret them, both costly to store and slow to share. Simulation outputs are stored as 32-bit floats, yet numerical noise and astrophysical uncertainties make better than percent-level pixel accuracy unnecessary. Because cosmological fields are statistically homogeneous with nearly Gaussian mode amplitudes, classic rate-distortion results apply directly, and non-Gaussian structure permits further compression. We use the scientific compression package SZ3, and a neural compressor that fixes the quantization and learns the probability of each quantization bin with an autoregressive transformer. SZ3 beats the Gaussian approach only where the field is smooth on the pixel scale or its values pile up at a single point, while the neural approach matches or beats SZ3 on every field we consider and is 9 to 27\% below the Gaussian-optimal coder at fixed distortion, with the largest margin where the field is most non-Gaussian. Because the quantizer, not the network, sets the error, a poorly trained model can waste bits but never cost accuracy. A model trained only on weak lensing convergence maps transfers to $N$-body density fields without retraining, suggesting it has learned generic properties of cosmological structure rather than features of one dataset. Eulerian grids need only 1-4 bits per pixel (bpp) at percent-level accuracy, and particle displacements and velocities 5-6 at the precisions they require. The approach carries over to observational data: on Rubin Observatory Data Preview 1 coadds, with the quantization step set to a quarter of the background noise, the fine-tuned transformer needs 4.0 bpp, 27\% fewer than the Gaussian coder.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hurum Maksora Tohfa, Matthew McQuinn. 2026-09-30. Two bits about lossy compression: On the limits of compression in cosmology. https://arxiv.org/abs/2609.40239

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Reconstruction of cosmic-ray direction and energy in radio arrays using deep ensemble graph neural networks

Using advanced machine learning techniques, we developed a method to reconstruct the arrival direction and energy of ultra-high-energy cosmic rays from the voltage traces they induce on ground-based radio detector arrays. In our approach, triggered antennas are represented as a graph structure, which serves as input for a graph neural network (GNN). By incorporating physical knowledge into both the GNN architecture and the input data, we improve the precision and reduce the required size of the training set with respect to a fully data-driven approach. This method achieves an angular resolution of 0.092 degrees and an electromagnetic energy reconstruction resolution of 16.4% on simulated data with realistic noise conditions. We also employ uncertainty estimation methods to enhance the reliability of our predictions, quantifying the confidence of the GNN's outputs and providing confidence intervals for both direction and energy reconstruction. Finally, we investigate strategies to verify the model's consistency and robustness under real-life variations, with the goal of identifying scenarios in which predictions remain reliable despite domain shifts between simulation and reality.

astro-ph.IM↗

Hybrid neural denoising for resource-efficient near- and sub-threshold radio triggering of extensive air showers

Autonomous radio self-triggering for extensive air showers requires strong rejection of radio-frequency interference while preserving weak pulses within station-level hardware constraints. We present a hybrid neural trigger comprising a compact waveform denoiser followed by a classifier. The method is evaluated using experimentally measured high-interference background traces and detector-folded simulated air-shower pulses produced with the Pierre Auger Offline chain, with the benchmark concentrated in the near-threshold regime. Hardware constraints are incorporated through hyperparameter optimisation, quantisation-aware training, and high-granularity fixed-point quantisation. Applying the conventional peak-envelope trigger after denoising increases its area under the receiver-operating-characteristic curve from 0.63 to 0.98. At a false-positive rate of 10^-4, the full denoiser-classifier chain retains about 41% of the held-out signal traces, compared with 27% for the classifier acting on the raw waveform and about 2% for the peak-envelope reference. The fixed-point firmware is synthesised, placed, and routed on representative field-programmable gate array targets, where it achieves timing closure with microsecond-scale latency and compact arithmetic-resource usage. These results establish hybrid neural denoising as a practical FPGA-compatible route toward radio-only triggering of weak and inclined air-shower signals in noisy environments.

astro-ph.IM↗

Characterization of the RF Board for microwave SQUID multiplexing readout electronics

Microwave SQUID multiplexing ($μ$MUX) is a widely used readout technique for large-scale transition-edge sensor (TES) arrays. It uses radio-frequency (RF) probe tones to interrogate cryogenic resonators, requiring frequency conversion between the baseband electronics and the cryogenic RF signal chain. This work describes the RF Board, a room-temperature frequency-conversion board deployed in the AliCPT $μ$MUX readout system. The board up-converts 0-4 GHz baseband I/Q signals to the 4-8 GHz RF band for injection into the cryogenic chain and down-converts returned RF signals to baseband I/Q for ADC digitization. For the current 1000-tone operation, with a DAC output tone power of -30 dBm/tone, the required power windows are -35 to -25 dBm/tone for the RF tones transmitted to the cryostat and -45 to -35 dBm/tone at the ADC input for the returned tones. The RF Board is characterized using swept single-tone measurements covering up-conversion, down-conversion, and RF loopback. Based on these measurements, the RF output power is calculated to be -31.06 to -25.53 dBm, satisfying the RF output window. Assuming a representative cryogenic-chain transmission of -40 dB, the loopback result gives an estimated returned power of -45 to -35 dBm, within the target range. These results show that the RF Board meets the wideband frequency-conversion and tone-power requirements for the $μ$MUX readout system.

astro-ph.IM↗