arXiv ScienceSearch

arXiv subjects

Chris Finlay

Publications and source records attributed to Chris Finlay.

At least 19 recordsLinked to original sources

TABASCAL: Removing multi-satellite interference from radio interferometry observations

In the first trajectory-based radio frequency interference (RFI) subtraction and calibration (TABASCAL) paper, we showed how to calibrate radio interferometers in the presence of RFI sources by simultaneously isolating the trajectories and signals of the RFI sources. In this paper, we show that we can accurately remove RFI (i.e. recover the astronomical signal) from simulated MeerKAT radio interferometry target data. We are able to do so for a single frequency channel, corrupted by up to nine simultaneous satellites, with average RFI amplitudes varying from weak to very strong (1-1000 Jy). Additionally, TABASCAL also manages to leverage the signal-to-noise ratio (S/N) of the RFI to phase-calibrate the astronomical signal. TABASCAL, effectively performs a suitably phased up fringe filter for each RFI source, which essentially allows for an ideal removal of RFI across all RFI strengths. As a result, TABASCAL is able to reach image noises equivalent to the uncorrupted, no-RFI, case. For larger RFI amplitudes, the resulting image noise is 10x - 100x smaller than those from traditional RFI flagging methods such as AOFLAGGER. As a specific application, we show that point-source science with TABASCAL almost matches the no-RFI case with near perfect completeness for all RFI amplitudes. In contrast, the completeness of AOFLAGGER and idealised $3\sigma$ flagging drops below 40% for strong RFI amplitudes, where recovered flux errors are approximately 10x - 100x worse than those from TABASCAL. Finally, we note that TABASCAL works for astronomical sources with both static and varying fluxes.

astro-ph.IM

Deep learning approach for identification of HII regions during reionization in 21-cm observations -- III. image recovery

The low-frequency component of the upcoming Square Kilometre Array Observatory (SKA-Low) will be sensitive enough to construct 3D tomographic images of the 21-cm signal distribution during reionisation. However, foreground contamination poses challenges for detecting this signal, and image recovery will heavily rely on effective mitigation methods. We introduce \texttt{SERENEt}, a deep-learning framework designed to recover the 21-cm signal from SKA-Low's foreground-contaminated observations, enabling the detection of ionised (HII) and neutral (HI) regions during reionisation. \texttt{SERENEt} can recover the signal distribution with an average accuracy of 75 per cent at the early stages ($\overline{x}_\mathrm{HI}\simeq0.9$) and up to 90 per cent at the late stages of reionisation ($\overline{x}_\mathrm{HI}\simeq0.1$). Conversely, HI region detection starts at 92 per cent accuracy, decreasing to 73 per cent as reionisation progresses. Beyond improving image recovery, \texttt{SERENEt} provides cylindrical power spectra with an average accuracy exceeding 93 per cent throughout the reionisation period. We tested \texttt{SERENEt} on a 10-degree field-of-view simulation, consistently achieving better and more stable results when prior maps were provided. Notably, including prior information about HII region locations improved 21-cm signal recovery by approximately 10 per cent. This capability was demonstrated by supplying \texttt{SERENEt} with ionising source distribution measurements, showing that high-redshift galaxy surveys of similar observation fields can optimise foreground mitigation and enhance 21-cm image construction.

astro-ph.CO

Trajectory Based RFI Subtraction and Calibration for Radio Interferometry

Radio interferometry calibration and Radio Frequency Interference (RFI) removal are usually done separately. Here we show that jointly modelling the antenna gains and RFI has significant benefits when the RFI follows precise trajectories, such as for satellites. One surprising benefit is improved calibration solutions, by leveraging the RFI signal itself. We present tabascal (TrAjectory BAsed RFI Subtraction and CALibration), a new algorithm that jointly models the RFI and calibration parameters in visibilities. We test tabascal on simulated MeerKAT calibration observations contaminated by satellite-based RFI. We obtain gain estimates that are both unbiased and up to an order of magnitude better constrained compared to uncontaminated data. When combined with an ad hoc RFI subtraction scheme, tabascal solutions can be further applied to an adjacent target observation: 5 minutes of calibration data results in an image with about a third the noise achieved when using flagging alone. The recovered flux distribution of RFI subtracted data was on par with uncontaminated data. In contrast, RFI flagging alone resulted in a higher detection threshold and consistent underestimation of source fluxes. For a mean RFI amplitude of 17 Jy, using RFI subtraction leads to less than 1% loss of data compared to 75% data loss from an ideal $3\sigma$ flagging algorithm, a very significant increase in data available for science analysis. Although we have examined the case of satellite RFI, tabascal should work for any RFI moving on parameterizable trajectories, relative to the phase centre, such as planes and/or objects fixed to the ground.

astro-ph.IM

Multi-Resolution Continuous Normalizing Flows

Recent work has shown that Neural Ordinary Differential Equations (ODEs) can serve as generative models of images using the perspective of Continuous Normalizing Flows (CNFs). Such models offer exact likelihood calculation, and invertible generation/density estimation. In this work we introduce a Multi-Resolution variant of such models (MRCNF), by characterizing the conditional distribution over the additional information required to generate a fine image that is consistent with the coarse image. We introduce a transformation between resolutions that allows for no change in the log likelihood. We show that this approach yields comparable likelihood values for various image datasets, with improved performance at higher resolutions, with fewer parameters, using only 1 GPU. Further, we examine the out-of-distribution properties of (Multi-Resolution) Continuous Normalizing Flows, and find that they are similar to those of other likelihood-based generative models.

cs.CV

Adversarial Boot Camp: label free certified robustness in one epoch

Machine learning models are vulnerable to adversarial attacks. One approach to addressing this vulnerability is certification, which focuses on models that are guaranteed to be robust for a given perturbation size. A drawback of recent certified models is that they are stochastic: they require multiple computationally expensive model evaluations with random noise added to a given input. In our work, we present a deterministic certification approach which results in a certifiably robust model. This approach is based on an equivalence between training with a particular regularized loss, and the expected values of Gaussian averages. We achieve certified models on ImageNet-1k by retraining a model with this loss for one epoch without the use of label information.

cs.LG

Climate & BCG: Effects on COVID-19 Death Growth Rates

Multiple studies have suggested the spread of COVID-19 is affected by factors such as climate, BCG vaccinations, pollution and blood type. We perform a joint study of these factors using the death growth rates of 40 regions worldwide with both machine learning and Bayesian methods. We find weak, non-significant (< 3$\sigma$) evidence for temperature and relative humidity as factors in the spread of COVID-19 but little or no evidence for BCG vaccination prevalence or $\text{PM}_{2.5}$ pollution. The only variable detected at a statistically significant level (>3$\sigma$) is the rate of positive COVID-19 tests, with higher positive rates correlating with higher daily growth of deaths.

q-bio.PE

Learning normalizing flows from Entropy-Kantorovich potentials

We approach the problem of learning continuous normalizing flows from a dual perspective motivated by entropy-regularized optimal transport, in which continuous normalizing flows are cast as gradients of scalar potential functions. This formulation allows us to train a dual objective comprised only of the scalar potential functions, and removes the burden of explicitly computing normalizing flows during training. After training, the normalizing flow is easily recovered from the potential functions.

cs.LG

Deterministic Gaussian Averaged Neural Networks

We present a deterministic method to compute the Gaussian average of neural networks used in regression and classification. Our method is based on an equivalence between training with a particular regularized loss, and the expected values of Gaussian averages. We use this equivalence to certify models which perform well on clean data but are not robust to adversarial perturbations. In terms of certified accuracy and adversarial robustness, our method is comparable to known stochastic methods such as randomized smoothing, but requires only a single model evaluation during inference.

cs.LG

Deep Learning improves identification of Radio Frequency Interference

Flagging of Radio Frequency Interference (RFI) is an increasingly important challenge in radio astronomy. We present R-Net, a deep convolutional ResNet architecture that significantly outperforms existing algorithms -- including the default MeerKAT RFI flagger, and deep U-Net architectures -- across all metrics including AUC, F1-score and MCC. We demonstrate the robustness of this improvement on both single dish and interferometric simulations and, using transfer learning, on real data. Our R-Net model's precision is approximately $90\%$ better than the current MeerKAT flagger at $80\%$ recall and has a 35\% higher F1-score with no additional performance cost. We further highlight the effectiveness of transfer learning from a model initially trained on simulated MeerKAT data and fine-tuned on real, human-flagged, KAT-7 data. Despite the wide differences in the nature of the two telescope arrays, the model achieves an AUC of 0.91, while the best model without transfer learning only reaches an AUC of 0.67. We consider the use of phase information in our models but find that without calibration the phase adds almost no extra information relative to amplitude data only. Our results strongly suggest that deep learning on simulations, boosted by transfer learning on real data, will likely play a key role in the future of RFI flagging of radio astronomy data.

astro-ph.IM

How to train your neural ODE: the world of Jacobian and kinetic regularization

Training neural ODEs on large datasets has not been tractable due to the necessity of allowing the adaptive numerical ODE solver to refine its step size to very small values. In practice this leads to dynamics equivalent to many hundreds or even thousands of layers. In this paper, we overcome this apparent difficulty by introducing a theoretically-grounded combination of both optimal transport and stability regularizations which encourage neural ODEs to prefer simpler dynamics out of all the dynamics that solve a problem well. Simpler dynamics lead to faster convergence and to fewer discretizations of the solver, considerably decreasing wall-clock time without loss in performance. Our approach allows us to train neural ODE-based generative models to the same performance as the unregularized dynamics, with significant reductions in training time. This brings neural ODEs closer to practical relevance in large-scale applications.

stat.ML

Farkas layers: don't shift the data, fix the geometry

Successfully training deep neural networks often requires either batch normalization, appropriate weight initialization, both of which come with their own challenges. We propose an alternative, geometrically motivated method for training. Using elementary results from linear programming, we introduce Farkas layers: a method that ensures at least one neuron is active at a given layer. Focusing on residual networks with ReLU activation, we empirically demonstrate a significant improvement in training capacity in the absence of batch normalization or methods of initialization across a broad range of network sizes on benchmark datasets.

cs.LG

A principled approach for generating adversarial images under non-smooth dissimilarity metrics

Deep neural networks perform well on real world data but are prone to adversarial perturbations: small changes in the input easily lead to misclassification. In this work, we propose an attack methodology not only for cases where the perturbations are measured by $\ell_p$ norms, but in fact any adversarial dissimilarity metric with a closed proximal form. This includes, but is not limited to, $\ell_1, \ell_2$, and $\ell_\infty$ perturbations; the $\ell_0$ counting "norm" (i.e. true sparseness); and the total variation seminorm, which is a (non-$\ell_p$) convolutional dissimilarity measuring local pixel changes. Our approach is a natural extension of a recent adversarial attack method, and eliminates the differentiability requirement of the metric. We demonstrate our algorithm, ProxLogBarrier, on the MNIST, CIFAR10, and ImageNet-1k datasets. We consider undefended and defended models, and show that our algorithm easily transfers to various datasets. We observe that ProxLogBarrier outperforms a host of modern adversarial attacks specialized for the $\ell_0$ case. Moreover, by altering images in the total variation seminorm, we shed light on a new class of perturbations that exploit neighboring pixel information.

cs.LG

Scaleable input gradient regularization for adversarial robustness

In this work we revisit gradient regularization for adversarial robustness with some new ingredients. First, we derive new per-image theoretical robustness bounds based on local gradient information. These bounds strongly motivate input gradient regularization. Second, we implement a scaleable version of input gradient regularization which avoids double backpropagation: adversarially robust ImageNet models are trained in 33 hours on four consumer grade GPUs. Finally, we show experimentally and through theoretical certification that input gradient regularization is competitive with adversarial training. Moreover we demonstrate that gradient regularization does not lead to gradient obfuscation or gradient masking.

stat.ML

The LogBarrier adversarial attack: making effective use of decision boundary information

Adversarial attacks for image classification are small perturbations to images that are designed to cause misclassification by a model. Adversarial attacks formally correspond to an optimization problem: find a minimum norm image perturbation, constrained to cause misclassification. A number of effective attacks have been developed. However, to date, no gradient-based attacks have used best practices from the optimization literature to solve this constrained minimization problem. We design a new untargeted attack, based on these best practices, using the established logarithmic barrier method. On average, our attack distance is similar or better than all state-of-the-art attacks on benchmark datasets (MNIST, CIFAR10, ImageNet-1K). In addition, our method performs significantly better on the most challenging images, those which normally require larger perturbations for misclassification. We employ the LogBarrier attack on several adversarially defended models, and show that it adversarially perturbs all images more efficiently than other attacks: the distance needed to perturb all images is significantly smaller with the LogBarrier attack than with other state-of-the-art attacks.

cs.LG

Calibrated Top-1 Uncertainty estimates for classification by score based models

While the accuracy of modern deep learning models has significantly improved in recent years, the ability of these models to generate uncertainty estimates has not progressed to the same degree. Uncertainty methods are designed to provide an estimate of class probabilities when predicting class assignment. While there are a number of proposed methods for estimating uncertainty, they all suffer from a lack of calibration: predicted probabilities can be off from empirical ones by a few percent or more. By restricting the scope of our predictions to only the probability of Top-1 error, we can decrease the calibration error of existing methods to less than one percent. As a result, the scores of the methods also improve significantly over benchmarks.

stat.ML

Improved robustness to adversarial examples using Lipschitz regularization of the loss

We augment adversarial training (AT) with worst case adversarial training (WCAT) which improves adversarial robustness by 11% over the current state-of-the-art result in the $\ell_2$ norm on CIFAR-10. We obtain verifiable average case and worst case robustness guarantees, based on the expected and maximum values of the norm of the gradient of the loss. We interpret adversarial training as Total Variation Regularization, which is a fundamental tool in mathematical image processing, and WCAT as Lipschitz regularization.

cs.LG

Lipschitz regularized Deep Neural Networks generalize and are adversarially robust

In this work we study input gradient regularization of deep neural networks, and demonstrate that such regularization leads to generalization proofs and improved adversarial robustness. The proof of generalization does not overcome the curse of dimensionality, but it is independent of the number of layers in the networks. The adversarial robustness regularization combines adversarial training, which we show to be equivalent to Total Variation regularization, with Lipschitz regularization. We demonstrate empirically that the regularized models are more robust, and that gradient norms of images can be used for attack detection.

cs.LG

Improved accuracy of monotone finite difference schemes on point clouds and regular grids

Finite difference schemes are the method of choice for solving nonlinear, degenerate elliptic PDEs, because the Barles-Sougandis convergence framework [Barles and Sougandidis, Asymptotic Analysis, 4(3):271-283, 1991] provides sufficient conditions for convergence to the unique viscosity solution [Crandall, Ishii and Lions, Bull. Amer. Math Soc., 27(1):1-67, 1992]. For anisotropic operators, such as the Monge-Ampere equation, wide stencil schemes are needed [Oberman, SIAM J. Numer. Anal., 44(2):879-895]. The accuracy of these schemes depends on both the distances to neighbors, $R$, and the angular resolution, $d\theta$. On uniform grids, the accuracy is $\mathcal O(R^2 + d\theta)$. On point clouds, the most accurate schemes are of $\mathcal O(R + d\theta)$, by Froese [Numerische Mathematik, 138(1):75-99, 2018]. In this work, we construct geometrically motivated schemes of higher accuracy in both cases: order $\mathcal O(R + d\theta^2)$ on point clouds, and $\mathcal O(R^2 + d\theta^2)$ on uniform grids.

math.NA