arXiv ScienceSearch

arXiv subjects

David Thomas

Publications and source records attributed to David Thomas.

At least 19 recordsLinked to original sources

Priming: Hybrid State Space Models From Pre-trained Transformers

Hybrid State-Space models combine Attention with recurrent State-Space Model (SSM) layers, balancing eidetic memory from Attention with compressed fading memory from SSMs. This yields smaller Key-Value caches and faster decoding than Transformers, along with a richer architectural design space. Exploring that design space at scale has so far required training from scratch, a barrier that has kept most large-model Hybrid research within a narrow range of architectures. We introduce Priming, a method that turns Hybrid architecture design from a pre-training problem into a knowledge transfer one. Priming initializes a Hybrid model from a pre-trained Transformer and, through short alignment and post-training phases, recovers downstream quality using less than 0.5% of the source model's pre-training token budget. Priming is agnostic to the source Transformer family (e.g., Qwen, Llama, Mistral), model class (dense or Mixture-of-Experts), and model scale. Priming enables us to run the first controlled comparison of SSM layer types at scale under identical conditions. We evaluate, Gated KalmaNet (GKA), Gated DeltaNet (GDN), and Mamba-2, and show that their expressiveness hierarchy, GKA>GDN>Mamba-2, directly predicts downstream performance on long-context reasoning tasks. We scale Priming to 8B/32B reasoning models with native 128K contexts. Our Hybrid GKA 32B improves over its source Qwen3-32B by +3.8 average reasoning points, while staying within 1% of a Transformer post-trained on the same data and enabling up to 2.3x higher decode throughput. To foster research on Hybrid architectures, we release a model zoo of primed Hybrid models for long-context reasoning and instruction following, together with the Priming training and inference code (Sequence Parallelism algorithms for long-context training, optimized GKA kernels, and vLLM serving plugin), all under Apache~2.0 License.

cs.LG

A 2D antiscatter grid and scatter sampling based CBCT method for online dose calculations during CBCT guided radiation therapy of pelvis

Online dose calculations before radiation treatment have applications in dose delivery verification, plan adaptation, and treatment planning. We propose a novel CBCT imaging pipeline to enhance accuracy. Our approach aims to improve HU accuracy in CBCT images for more precise dose calculations. A quantitative CBCT pipeline was implemented, combining data correction strategies and scatter rejection, achieving high CT number accuracy. We evaluated the pipeline's effect using pelvis anatomy phantoms and found that dosimetric errors in quantitative CBCT-based dose calculations were minimal. In contrast, clinical CBCT and high-performance ASG CBCT-based plans showed significant errors. The proposed quantitative CBCT pipeline offers comparable dose calculation accuracy to the gold-standard planning CT, eliminating the need for density overrides and enabling precise dose delivery monitoring or online plan adaptations in radiation therapy.

physics.med-ph

An Event-Driven Approach To Genotype Imputation On A Custom RISC-V FPGA Cluster

This paper proposes an event-driven solution to genotype imputation, a technique used to statistically infer missing genetic markers in DNA. The work implements the widely accepted Li and Stephens model, primary contributor to the computational complexity of modern x86 solutions, in an attempt to determine whether further investigation of the application is warranted in the event-driven domain. The model is implemented using graph-based Hidden Markov Modeling and executed as a customized forward/backward dynamic programming algorithm. The solution uses an event-driven paradigm to map the algorithm to thousands of concurrent cores, where events are small messages that carry both control and data within the algorithm. The design of a single processing element is discussed. This is then extended across multiple FPGAs and executed on a custom RISC-V NoC FPGA cluster called POETS. Results demonstrate how the algorithm scales over increasing hardware resources and a 48 FPGA run demonstrates a 270X reduction in wall-clock processing time when compared to a single-threaded x86 solution. Optimisation of the algorithm via linear interpolation is then introduced and tested, with results demonstrating a wall-clock reduction time of approx. 5 orders of magnitude when compared to a similarly optimised x86 solution.

cs.DC

2x2 MIMO Prototype for BER and EVM Measurements in Metal Enclosure

In this work, we present a 2x2 near-field multi-input multiple-output (MIMO) prototype for bit-error-rate (BER) and error vector magnitude (EVM) measurements in a metal enclosure. The near-field MIMO prototype is developed using software-defined-radios (SDRs) for over-the-air transmission of QPSK modulated baseband waveforms. We check the near-field MIMO BER and EVM measurements in three different scenarios in a highly reflecting metal enclosure environment. In the first scenario, the line-of-sight (LOS) communication link is investigated when the mode-stirrer is stationary. In stationary channel conditions near-field MIMO BER and EVM measurements are performed. In the second scenario, BER and EVM measurements are performed in dynamic channel conditions when the mode-stirrer is set to move continuously. In the third scenario, LOS communication near-field MIMO BER and EVM measurements are performed in stationary channel conditions but now in the presence of MIMO interference. In three different scenarios, near-field MIMO BER and EVM measurements are investigated at different Tx USRP gain values and in the presence of varying levels of MIMO interference.

cs.IT

Acquisition-invariant brain MRI segmentation with informative uncertainties

Combining multi-site data can strengthen and uncover trends, but is a task that is marred by the influence of site-specific covariates that can bias the data and therefore any downstream analyses. Post-hoc multi-site correction methods exist but have strong assumptions that often do not hold in real-world scenarios. Algorithms should be designed in a way that can account for site-specific effects, such as those that arise from sequence parameter choices, and in instances where generalisation fails, should be able to identify such a failure by means of explicit uncertainty modelling. This body of work showcases such an algorithm, that can become robust to the physics of acquisition in the context of segmentation tasks, while simultaneously modelling uncertainty. We demonstrate that our method not only generalises to complete holdout datasets, preserving segmentation quality, but does so while also accounting for site-specific sequence choices, which also allows it to perform as a harmonisation tool.

eess.IV

The role of MRI physics in brain segmentation CNNs: achieving acquisition invariance and instructive uncertainties

Being able to adequately process and combine data arising from different sites is crucial in neuroimaging, but is difficult, owing to site, sequence and acquisition-parameter dependent biases. It is important therefore to design algorithms that are not only robust to images of differing contrasts, but also be able to generalise well to unseen ones, with a quantifiable measure of uncertainty. In this paper we demonstrate the efficacy of a physics-informed, uncertainty-aware, segmentation network that employs augmentation-time MR simulations and homogeneous batch feature stratification to achieve acquisition invariance. We show that the proposed approach also accurately extrapolates to out-of-distribution sequence samples, providing well calibrated volumetric bounds on these. We demonstrate a significant improvement in terms of coefficients of variation, backed by uncertainty based volumetric validation.

eess.IV

PER Measurement of BLE in RF Interference and Harsh Electromagnetic Environment

Bluetooth Low Energy (BLE) is a short-range data transmission technology that is used for multimedia file sharing, home automation, and internet-of-things application. In this work, we perform packet error rate (PER) measurement and RF testing of BLE receiver in the harsh electromagnetic environment and in presence of RF interference. We check the PER performance in the line-of-sight (LOS) and non-line-of-sight (NLOS) scenario in absence of any interfering signal and in presence of wideband WLAN interference. The BLE PER measurements are conducted in a large reverberation chamber which is a rich scattering environment. Software-defined-radio has been used to create BLE communication link for PER measurement in LOS and NLOS configuration. The BLE PER is measured both in the presence and in absence of WLAN interference. Our measurement results show a higher PER for uncoded BLE PHY modes in NLOS channel condition and in presence of wideband interference. Whereas coded BLE PHY modes i.e. LE500K and LE125K are robust to interference with lower PER measurements.

eess.SP

Statistical Characterization of Wireless MIMO Channels in Mode-Stirred Enclosures

We present the statistical characterization of a 2x2 Multiple-Input Multiple-Output wireless link operated in a mode-stirred enclosure, with channel state information available only at the receiver (agnostic transmitter). Our wireless channel measurements are conducted in absence of line of sight and varying the inter-element spacing between the two antenna elements in both the transmit and receive array. The mode-stirred cavity is operated: i) at a low number of stirrer positions to create statistical inhomogeneity; ii) at two different loading conditions, empty and with absorbers, in order to mimic a wide range of realistic equipment level enclosures. Our results show that two parallel channels are obtained within the confined space at both the operating conditions. The statistical characterization of the wireless channel is presented in terms of coherence bandwidth, path loss, delay spread and Rician factor, and wideband channel capacity. It is found that the severe multipath fading supported by a highly reflecting environment creates unbalance between the two Multiple-Input Multiple-Output channels, even in presence of substantial losses. Furthermore, the channel capacity has a multi-modal distribution whose average and variance scale monotonically with the number of absorbers. Results are of interest in IoT devices, including wireless chip-to-chip and device-to-device communications, operating in highly reflective environments.

cs.IT

Near-field Image Transmission and EVM Measurements in Rich Scattering Environment

In this work, we present near-field image transmission and error vector magnitude measurement in a rich scattering environment in a metal enclosure. We check the effect of loading metal enclosure on the performance of SDR based near-field communication link. We focus on the key communication receiver parameters to observe the effect of near-field link in presence of rich-scattering and in presence of loading with RF absorber cones. The near-field performance is measured by transmitting wideband OFDM-modulated packets containing image information. Our finding suggests that the performance of OFDM based wideband near-field communication improves when the metal enclosure is loaded with RF absorbers. Near-field EVM improves when the enclosure is loaded with RF absorber cones. Loading of the metal enclosure has the effect of increased coherence bandwidth. Frequency selectivity was observed in an empty enclosure which suggests coherence bandwidth less than the signal bandwidth.

eess.SP

Physics-informed brain MRI segmentation

Magnetic Resonance Imaging (MRI) is one of the most flexible and powerful medical imaging modalities. This flexibility does however come at a cost; MRI images acquired at different sites and with different parameters exhibit significant differences in contrast and tissue appearance, resulting in downstream issues when quantifying brain anatomy or the presence of pathology. In this work, we propose to combine multiparametric MRI-based static-equation sequence simulations with segmentation convolutional neural networks (CNN), to make these networks robust to variations in acquisition parameters. Results demonstrate that, when given both the image and their associated physics acquisition parameters, CNNs can produce segmentations that exhibit robustness to acquisition variations. We also show that the proposed physics-informed methods can be used to bridge multi-centre and longitudinal imaging studies where imaging acquisition varies across a site or in time.

physics.med-ph

Unveiling the Rich and Diverse Universe of Subsecond Astrophysics through LSST Star Trails

We present a unique method that allows the LSST to scan the sky for stellar variability on short timescales. The operational component of the strategy requires LSST to take star trail images. The image processing component uses deep learning to sift for transient events on timescales down to 10 ms. We advocate for enabling this observing mode with LSST, as coupling this capability with the LSST's tremendous 319.5 m$^2$deg$^2$ etendue will produce the first wide area optical survey of the universe on these timescales. We explain how these data will advance both planned lines of investigation and enable new research in the areas of stellar flares, cataclysmic variables, active galactic nuclei, Kuiper Belt objects, gamma-ray bursts, and fast radio bursts.

astro-ph.IM

Binding energies of excitonic complexes in type-II quantum rings from diffusion quantum Monte Carlo calculations

Excitonic complexes in type-II quantum-ring heterostructures may be considered as artificial atoms due to the confinement of only one charge-carrier type in an artificial nucleus. Binding energies of excitons, trions, and biexcitons in these nanostructures are then effectively ionization energies of these artificial atoms. The binding energies reported here are calculated within the effective-mass approximation using the diffusion quantum Monte Carlo method and realistic geometries for gallium antimonide rings in gallium arsenide. The electrons form a halo outside the ring, with very little charge density inside the central cavity of the ring. The de-excitonization and binding energies of the complexes are relatively independent of the precise shape of the ring.

cond-mat.mes-hall

Searching for Sub-Second Stellar Variability with Wide-Field Star Trails and Deep Learning

We present a method that enables wide field ground-based telescopes to scan the sky for sub-second stellar variability. The method has operational and image processing components. The operational component is to take star trail images. Each trail serves as a light curve for its corresponding source and facilitates sub-exposure photometry. We train a deep neural network to identify stellar variability in wide-field star trail images. We use the Large Synoptic Survey Telescope (LSST) Photon Simulator to generate simulated star trail images and include transient bursts as a proxy for variability. The network identifies transient bursts on timescales down to 10 milliseconds. We argue that there are multiple fields of astrophysics that can be advanced by the unique combination of time resolution and observing throughput that our method offers.

astro-ph.IM

A Phase-Space Approach for Propagating Field-Field Correlation Functions

We show that radiation from complex and inherently random but correlated wave sources can be modelled efficiently by using an approach based on the Wigner distribution function. Our method exploits the connection between correlation functions and theWigner function and admits in its simplest approximation a direct representation in terms of the evolution of ray densities in phase space. We show that next leading order corrections to the ray-tracing approximation lead to Airy-function type phase space propagators. By exploiting the exact Wigner function propagator, inherently wave-like effects such as evanescent decay or radiation from more heterogeneous sources as well as diffraction and reflections can be included and analysed. We discuss in particular the role of evanescent waves in the near-field of non-paraxial sources and give explicit expressions for the growth rate of the correlation length as function of the distance from the source. Furthermore, results for the reflection of partially coherent sources from flat mirrors are given. We focus here on electromagnetic sources at microwave frequencies and modelling efforts in the context of electromagnetic compatibility.

nlin.CD

A Domain Specific Approach to Heterogeneous Computing: From Availability to Accessibility

We advocate a domain specific software development methodology for heterogeneous computing platforms such as Multicore CPUs, GPUs and FPGAs. We argue that three specific benefits are realised from adopting such an approach: portable, efficient implementations across heterogeneous platforms; domain specific metrics of quality that characterise platforms in a form software developers will understand; automatic, optimal partitioning across the available computing resources. These three benefits allow a development methodology for software developers where they describe their computational problems in a single, easy to understand form, and after a modeling procedure on the available resources, select how they would like to trade between various domain specific metrics. Our work on the Forward Financial Framework ($F^3$) demonstrates this methodology in practise. We are able to execute a range of computational finance option pricing tasks efficiently upon a wide range of CPU, GPU and FPGA computing platforms. We can also create accurate financial domain metric models of walltime latency and statistical confidence. Furthermore, we believe that we can support automatic, optimal partitioning using this execution and modelling capability.

cs.CE

An Automatic Mixed Software Hardware Pipeline Builder for CPU-FPGA Platforms

Our toolchain for accelerating application called Courier-FPGA, is designed for utilize the processing power of CPU-FPGA platforms for software programmers and non-expert users. It automatically gathers runtime information of library functions from a running target binary, and constructs the function call graph including input-output data. Then, it uses corresponding predefined hardware modules if these are ready for FPGA and prepares software functions on CPU by using Pipeline Generator. The Pipeline Generator builds a pipeline control program by using Intel Threading Building Block to run both hardware modules and software functions in parallel. Finally, Courier-FPGA dynamically replaces the original functions in the binary and accelerates it by using the built pipeline. Courier-FPGA performs these acceleration processes without user intervention, source code tweaks or re-compilations of the binary. We describe the technical details of this mixed software hardware pipeline on CPU-FPGA platforms in this paper. In our case study, Courier-FPGA was used to accelerate a corner detection using the Harris-Stephens method application binary on the Zynq platform. A series of functions were off-loaded, and speed up 15.36 times was achieved by using the built pipeline.

cs.DC

Projection and Galaxy Clustering Fourier Spectra

Second order perturbation theory predicts a specific dependence of the bispectrum, or three-point correlation function in the Fourier transform domain, on the shape of the configuration of its three wave vector arguments, which can be taken as a signature of structure formed by gravitational instability. Comparing this known dependence on configuration shape with the weak shape dependence of the galaxy bispectrum has been suggested as an indication of bias in the galaxy distribution. However, to interpret results obtained from projected catalogs, we must first understand the effects of projection on this shape dependence. We present expressions for the projected power spectrum and bispectrum in both Cartesian and spherical geometries, and we examine the effects of projection on the predicted bispectrum with particular attention to the dependence on configuration shape. Except for an overall numerical factor, for Cartesian projection with characteristic depth $ \Dstar $ there is little effect on the shape dependence of the bispectrum for wavelengths small compared to $ \Dstar $ or projected wavenumbers $ q \Dstar \gg 1 $. For angular projection, a scaling law is found for spherical harmonic index $ \ell \gg 1 $, but there is always a mixing of scales over the range of the selection function. For large $ \ell $ it is sufficient to examine a small portion of the sky.

astro-ph

Generalized Limits to the Number of Light Particle Degrees of Freedom from Big Bang Nucleosynthesis

We compute the big bang nucleosynthesis limit on the number of light neutrino degrees of freedom in a model-independent likelihood analysis based on the abundances of He4 and Li7. We use the two-dimensional likelihood functions to simultaneously constrain the baryon-to-photon ratio and the number of light neutrinos for a range of He4 abundances $Y_p$ = 0.225 -- 0.250, as well as a range in primordial Li7 abundances from (1.6 to 4.1) $ \times 10^{-10}$. For (Li7/H)$_p = 1.6 \times 10^{-10}$, as can be inferred from the Li7 data from Population II halo stars, the upper limit to $N_ν$ based on the current best estimate of the primordial He4 abundance of $Y_p = 0.238$, is $N_ν< 4.3$ and varies from $N_ν< 3.3$ (at 95% C.L.) when $Y_p=0.225$ to $N_ν< 5.3$ when $Y_p=0.250$. If Li7 is depleted in these stars the upper limit to $N_ν$ is relaxed. Taking (Li7/H)$_p = 4.1 \times 10^{-10}$, the limit varies from $N_ν< 3.9$ when $Y_p = 0.225$ to $N_ν\la 6$ when $Y_p = 0.250$. We also consider the consequences on the upper limit to $N_ν$ if recent observations of deuterium in high-redshift quasar absorption-line systems are confirmed.

hep-ph