arXiv ScienceSearch

arXiv · 2609.25988

Options for Compression of radio interferometry data: lossy compression of visibilities and lossless compression of uv-visibility grids for the MHONGOOSE survey

Abstract

Next generation radio astronomy telescopes are challenging existing data reduction paradigms. With ever more antennas, larger bandwidths, and sometimes multiple primary beams, they often generate more observed data products than can readily be stored long-term. Thus, data storage becomes a major cost driver and processing constraint. In this paper, we test two methods of addressing this problem: grid-stacking, a two-stage lossless compression solution; and the lossy compression of the raw visibilities before traditional processing. To demonstrate these solutions we utilised a deep imaging pipeline based on software for the ASKAP telescope, ASKAPSoft, but applied to a strong source (NGC1566) from the deep MeerKAT HI spectral line project, MHONGOOSE. The grid-stacking solution reproduces the spectrum from traditional processing to within better than 0.7%, and also allows for the reconstruction of other weighting scales without significant computing costs. In comparison, image-stacking also reproduces the spectrum from the traditional processing, to within better than 3% but with worse image residuals in the cube. The lossy compression, even at a near ten-fold reduction in file size, reproduces the spectra almost perfectly (to better than ~0.01% in all cases). Thus both compression methods are promising solutions, and we discuss considerations for their application.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Richard Dodson, Jonghwan Rhee, W. J. G. de Blok, Martin Meyer, Alexander Williamson, Daniel Mitchell, Kristóf Rozgonyi, María J. Rioja, Pascal J. Elahi. 2026-09-22. Options for Compression of radio interferometry data: lossy compression of visibilities and lossless compression of uv-visibility grids for the MHONGOOSE survey. https://arxiv.org/abs/2609.25988

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Simons Observatory: Development of a Pipeline to Detect Rapid Transients in Time-Ordered Data

We introduce a method for detecting astrophysical transients evolving on timescales of milliseconds to minutes using cosmic microwave background (CMB) survey telescopes. While previous transient searches in CMB data operate in map space, our pipeline directly processes the raw time-ordered data, enabling sensitivity to fast, dynamic signals. We integrate our detection approach into the Simons Observatory time-domain pipeline and assess the performance by injecting symmetric, stellar flare-like light curves into simulated observations. For events flaring with a timescale of 0.5 s, the pipeline detects $\gtrsim90$ % of events at flux densities of 800, 1150, 1650, and 4250\,mJy when measured in the 93, 145, 225, and 280 GHz bands respectively. At a fixed peak flux density, the pipeline more readily detects longer flares. The limiting flux density for 90 % completeness is four times lower for a $\ge5$ s flare than for a 0.5 s flare, while the flux density limits for $\gtrsim50$ % detection efficiency are comparable to the rms noise of the time-ordered data. We are able to determine the position of detected events in each observing band, with a positional uncertainty at the detection threshold comparable to the telescope resolution at that band. These results demonstrate the readiness of this pipeline for incorporation into upcoming Simons Observatory data analyses.

astro-ph.IM

Fitting Moving Objects in Up-The-Ramp Data with Applications to the Roman Space Telescope and JWST

A moving object breaks the fundamental property of constant per-pixel count rates in an astronomical image read out up-the-ramp. In this paper, we show how to fit a moving object's path across a detector as that detector is read out nondestructively. We write the full likelihood function for every pixel subject to a constant count rate plus a time-dependent count rate due to a moving source. Assuming the moving source to be point-like and assuming the effective point-spread function to be known, we are left with four parameters that enter the likelihood nonlinearly: two for position and two for velocity. All remaining parameters can be optimized using closed-form expressions. Our approach extracts maximal information on a moving source's position and speed and enables the source to be accurately removed from the image. We investigate the dependence of flux, position, and velocity precision on the target's speed and the readout pattern. We also find a small, positive bias on the recovered flux due to the need to fit for an uncertain position and speed. Our approach can be used for space-based images with minor Solar system bodies in the foreground, e.g.~from Roman and JWST, or for ground-based observations with satellites in the foreground. We demonstrate the promise of our method with a fit to an asteroid track observed serendipitously by the NIRISS instrument on JWST, comparing it to the performance of the JWST pipeline. Python code implementing our approach is available at https://github.com/t-brandt/moving_source. The total computational cost to fit the track of a moving object is $\sim$1 second on a 2023 Macbook Pro.

astro-ph.IM

An FPCA-enhanced Ensemble Learning Framework for Photometric Identification of Type Ia Supernovae

Type Ia supernovae (SNe Ia) are essential tools for addressing key cosmic questions, including the Hubble tension and the nature of dark energy. Modern surveys are predominantly photometry-based, making the construction of a clean photometric SNe Ia sample crucial. In this study, we investigate whether functional principal component analysis (FPCA) scores derived from photometric light curves, combined with ensemble learning, can reliably distinguish SNe Ia from other transients using the PLAsTiCC dataset. FPCA provides a data-driven, flexible characterization of light curves without relying on rigid theoretical model assumptions. Light curves are fitted by minimizing residuals with penalty terms from clean samples, making the method robust to outliers or poorly sampled bands. The first two FPCA scores and peak magnitudes across the five LSST bands are used as classification features. We implement two complementary binary classifiers: an ensemble boosting model (CatBoost) and a statistical probabilistic method based on Euclidean distances. CatBoost slightly outperforms the statistical method, achieving 98.5% accuracy and 97.8% precision. Performance remains robust (>90%) under typical photometric redshift uncertainties (σ = 0.1). On the spectroscopic DES Y5 sample, both methods reach approximately 90% accuracy and 95% precision, demonstrating strong out-of-domain generalization compared to state-of-the-art methods with limited cross-survey applicability. Applied to DECam DDF and DESIRT transients, the predictions strongly agree, and their intersection provides a high-confidence SNe Ia sample for cosmological analyses. Overall, this FPCA-based framework offers a powerful, flexible tool for classifying transients in upcoming large-scale surveys such as LSST and Roman.

astro-ph.IM