arXiv ScienceSearch

arXiv subjects

Kevin Wilson

Publications and source records attributed to Kevin Wilson.

At least 19 recordsLinked to original sources

Private Vertical Federated Inference for Time-Series

Institutions may benefit from collaborative inference on time-series data. In settings where privacy is necessary, multi-party computation (MPC) is a straightforward approach to providing strong guarantees, yet it remains prohibitively expensive and scales poorly with modern transformer architectures. Vertical Federated Learning (VFL) offers efficiency but suffers from privacy leakage at the embedding level, and securing the entire VFL model head via MPC remains prohibitively slow and communication-heavy for larger models. To enable practical, secure inference at scale, we propose "Public/Private Hybrid Head-VFL" (PPHH-VFL). This hybrid architecture splits the model head into an efficient plaintext public head and a secure, lightweight MPC private head. By applying adversarial training to the public embeddings, we mitigate privacy leakage; concurrently, the small private head securely preserves the flow of sensitive information needed for high downstream utility. Empirical evaluations on models ranging up to 86 million parameters demonstrate that PPHH-VFL accelerates inference by up to six orders of magnitude compared to end-to-end MPC. Compared to a standard VFL+MPC baseline, our approach scales significantly better, achieving a speedup of up to 44.4x in WAN and a 91.2x reduction in communication costs (dropping from 1.7 GB to 19 MB per batch), while simultaneously improving downstream classification accuracy by 2.50% and regression RMSE by 40.7%.

cs.LG

Conditional Copula models using loss-based Bayesian Additive Regression Trees

The study of dependence between random variables under external influences is a challenging problem in multivariate analysis. We address this by proposing a novel semi-parametric approach for conditional copula models using Bayesian additive regression trees (BART) models. BART is becoming a popular approach in statistical modelling due to its simple ensemble type formulation complemented by its ability to provide inferential insights. Although BART allows us to model complex functional relationships, it tends to suffer from overfitting. In this article, we exploit a loss-based prior for the tree topology that is designed to reduce the tree complexity. In addition, we propose a novel adaptive Reversible Jump Markov Chain Monte Carlo algorithm that is ergodic in nature and requires very few assumptions allowing us to model complex and non-smooth likelihood functions with ease. Moreover, we show that our method can efficiently recover the true tree structure and approximate a complex conditional copula parameter, and that our adaptive routine can explore the true likelihood region under a sub-optimal proposal variance. Lastly, we provide case studies concerning the effect of gross domestic product on the dependence between the life expectancies and literacy rates of the male and female populations of different countries.

stat.ME

Bayesian design and analysis of two-arm cluster randomised trials using assurance: extension to binary outcomes and comparison of MCMC and INLA

The paper considers two different designs; a two-arm superiority cluster randomised controlled trial (RCT) with a continuous outcome, and a twoarm superiority cluster RCT with a binary outcome. From a Bayesian perspective, for the analysis of the trial we use a (generalised) linear mixed effects model. We summarise the inference for the treatment effect for a cluster RCT based on the posterior distribution. Based on this inference we use assurance to choose the sample size. We consider and compare two different methods for the inference: Markov Chain Monte Carlo (MCMC) and Integrated Nested Laplace Approximations (INLA), and consider their implications for the assurance. We consider the Specialist Pre-hospital redirection for ischemic stroke thrombectomy (SPEEDY) trial, an RCT which has co-primary outcomes of thrombectomy rate and time to thrombectomy, as a case study for the developed Bayesian RCT designs. We demonstrate our novel approach to the sample size calculation using assurance on the SPEEDY trial, based on the results of a formal prior elicitation exercise with two clinical experts. The paper considers a range of different scenarios for cluster RCTs to evaluate INLA and MCMC, to determine when each inference scheme should be used, balancing the computational cost in terms of speed and accuracy. We make recommendations for when each should be used.

stat.ME

Recomposer: Event-roll-guided generative audio editing

Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their strong prior understanding of the data domain. We present a system for editing individual sound events within complex scenes able to delete, insert, and enhance individual sound events based on textual edit descriptions (e.g., ``enhance Door'') and a graphical representation of the event timing derived from an ``event roll'' transcription. We present an encoder-decoder transformer working on SoundStream representations, trained on synthetic (input, desired output) audio example pairs formed by adding isolated sound events to dense, real-world backgrounds. Evaluation reveals the importance of each part of the edit descriptions -- action, class, timing. Our work demonstrates ``recomposition'' is an important and practical application.

cs.SD

Comparing Methodological Variations in Seizure Onset Localisation Algorithms using intracranial EEG

During clinical treatment for epilepsy, the area of the brain thought to be responsible for pathological activity is identified. This identification is typically performed through visual assessment of EEG recordings; however, this is time consuming and prone to subjective inconsistency. Automated onset localisation algorithms provide objective identification of the onset location by highlighting changes in signal features associated with seizure onset. In this work we investigate how methodological differences in such algorithms can result in different onset locations being identified. We analysed ictal intracranial EEG (icEEG) recordings in 16 subjects (100 seizures) with drug-resistant epilepsy from the SWEZ-ETHZ public database. We identified a series of key methodological differences that must be considered when designing or selecting an onset localisation algorithm. These differences were demonstrated using three distinct algorithms that capture different, but complementary, seizure onset features: Imprint, Epileptogenicity Index, and Low Entropy Map. We assessed methodological differences (or Decision Points), and their impact on the identified onset locations. Our independent application of all three algorithms to the same ictal icEEG dataset revealed low agreement between them: 27-60% of onset channels showed minimal or no overlap. Therefore, we investigated the effect of three key differences: (i) how to define a baseline, (ii) whether low-frequency components are considered, and finally (iii) whether electrodecrement is considered. Changes at each Decision Point were found to substantially influence resultant onset channels (r>0.3). Our results demonstrate how seemingly small methodological changes can result in large differences in onset locations. We propose that key Decision Points must be considered when using or designing an onset localisation algorithm.

q-bio.NC

Incomplete resection of the icEEG seizure onset zone is not associated with post-surgical outcomes

Delineation of seizure onset regions from EEG is important for effective surgical workup. However, it is unknown if their complete resection is required for seizure freedom, or in other words, if post-surgical seizure recurrence is due to incomplete removal of the seizure onset regions. Retrospective analysis of icEEG recordings from 63 subjects (735 seizures) identified seizure onset regions through visual inspection and algorithmic delineation. We analysed resection of onset regions and correlated this with post-surgical seizure control. Most subjects had over half of onset regions resected (70.7% and 60.5% of subjects for visual and algorithmic methods, respectively). In investigating spatial extent of onset or resection, and presence of diffuse onsets, we found no substantial evidence of association with post-surgical seizure control (all AUC<0.7, p>0.05). Seizure onset regions tends to be at least partially resected, however a less complete resection is not associated with worse post-surgical outcome. We conclude that seizure recurrence after epilepsy surgery is not necessarily a result of failing to completely resect the seizure onset zone, as defined by icEEG. Other network mechanisms must be involved, which are not limited to seizure onset regions alone.

q-bio.NC

Unsupervised Multi-channel Separation and Adaptation

A key challenge in machine learning is to generalize from training data to an application domain of interest. This work generalizes the recently-proposed mixture invariant training (MixIT) algorithm to perform unsupervised learning in the multi-channel setting. We use MixIT to train a model on far-field microphone array recordings of overlapping reverberant and noisy speech from the AMI Corpus. The models are trained on both supervised and unsupervised training data, and are tested on real AMI recordings containing overlapping speech. To objectively evaluate our models, we also use a synthetic multi-channel AMI test set. Holding network architectures constant, we find that a fine-tuned semi-supervised model yields the largest improvement to SI-SNR and to human listening ratings across synthetic and real datasets, outperforming supervised models trained on well-matched synthetic data. Our results demonstrate that unsupervised learning through MixIT enables model adaptation on both single- and multi-channel real-world speech recordings.

cs.SD

Investigating Bayesian optimization for expensive-to-evaluate black box functions: Application in fluid dynamics

Bayesian optimization provides an effective method to optimize expensive-to-evaluate black box functions. It has been widely applied to problems in many fields, including notably in computer science, e.g. in machine learning to optimize hyperparameters of neural networks, and in engineering, e.g. in fluid dynamics to optimize control strategies that maximize drag reduction. This paper empirically studies and compares the performance and the robustness of common Bayesian optimization algorithms on a range of synthetic test functions to provide general guidance on the design of Bayesian optimization algorithms for specific problems. It investigates the choice of acquisition function, the effect of different numbers of training samples, the exact and Monte Carlo based calculation of acquisition functions, and both single-point and multi-point optimization. The test functions considered cover a wide selection of challenges and therefore serve as an ideal test bed to understand the performance of Bayesian optimization to specific challenges, and in general. To illustrate how these findings can be used to inform a Bayesian optimization setup tailored to a specific problem, two simulations in the area of computational fluid dynamics are optimized, giving evidence that suitable solutions can be found in a small number of evaluations of the objective function for complex, real problems. The results of our investigation can similarly be applied to other areas, such as machine learning and physical experiments, where objective functions are expensive to evaluate and their mathematical expressions are unknown.

cs.LG

Distance-Based Sound Separation

We propose the novel task of distance-based sound separation, where sounds are separated based only on their distance from a single microphone. In the context of assisted listening devices, proximity provides a simple criterion for sound selection in noisy environments that would allow the user to focus on sounds relevant to a local conversation. We demonstrate the feasibility of this approach by training a neural network to separate near sounds from far sounds in single channel synthetic reverberant mixtures, relative to a threshold distance defining the boundary between near and far. With a single nearby speaker and four distant speakers, the model improves scale-invariant signal to noise ratio by 4.4 dB for near sounds and 6.8 dB for far sounds.

cs.SD

A library of quantitative markers of seizure severity

Purpose: Understanding fluctuations of seizure severity within individuals is important for defining treatment outcomes and response to therapy, as well as developing novel treatments for epilepsy. Current methods for grading seizure severity rely on qualitative interpretations from patients and clinicians. Quantitative measures of seizure severity would complement existing approaches, for EEG monitoring, outcome monitoring, and seizure prediction. Therefore, we developed a library of quantitative electroencephalographic (EEG) markers that assess the spread and intensity of abnormal electrical activity during and after seizures. Methods: We analysed intracranial EEG (iEEG) recordings of 1056 seizures from 63 patients. For each seizure, we computed 16 markers of seizure severity that capture the signal magnitude, spread, duration, and post-ictal suppression of seizures. Results: Quantitative EEG markers of seizure severity distinguished focal vs. subclinical and focal vs. FTBTC seizures across patients. In individual patients, 71% had a moderate to large difference (ranksum r > 0.3) between focal and subclinical seizures in three or more markers. Circadian and longer-term changes in severity were found for 67% and 53% of patients, respectively. Conclusion: We demonstrate the feasibility of using quantitative iEEG markers to measure seizure severity. Our quantitative markers distinguish between seizure types and are therefore sensitive to established qualitative differences in seizure severity. Our results also suggest that seizure severity is modulated over different timescales. We envisage that our proposed seizure severity library will be expanded and updated in collaboration with the epilepsy research community to include more measures and modalities.

q-bio.NC

End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling speaker overlap and enabling straightforward handling of discriminative training, unlike traditional clustering-based diarization methods. The proposed system is designed to handle meetings with unknown numbers of speakers, using variable-number permutation-invariant cross-entropy based loss functions. We introduce several components that appear to help with diarization performance, including a local convolutional network followed by a global self-attention module, multi-task transfer learning using a speaker identification component, and a sequential approach where the model is refined with a second stage. These are trained and validated on simulated meeting data based on LibriSpeech and LibriTTS datasets; final evaluations are done using LibriCSS, which consists of simulated meetings recorded using real acoustics via loudspeaker playback. The proposed model performs better than previously proposed end-to-end diarization models on these data.

cs.SD

VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition

We introduce VoiceFilter-Lite, a single-channel source separation model that runs on the device to preserve only the speech signals from a target user, as part of a streaming speech recognition system. Delivering such a model presents numerous challenges: It should improve the performance when the input signal consists of overlapped speech, and must not hurt the speech recognition performance under all other acoustic conditions. Besides, this model must be tiny, fast, and perform inference in a streaming fashion, in order to have minimal impact on CPU, memory, battery and latency. We propose novel techniques to meet these multi-faceted requirements, including using a new asymmetric loss, and adopting adaptive runtime suppression strength. We also show that such a model can be quantized as a 8-bit integer model and run in realtime.

eess.AS

Unsupervised Sound Separation Using Mixture Invariant Training

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from synthetic mixtures created by adding up isolated ground-truth sources. Reliance on this synthetic training data is problematic because good performance depends upon the degree of match between the training data and real-world audio, especially in terms of the acoustic conditions and distribution of sources. The acoustic properties can be challenging to accurately simulate, and the distribution of sound types may be hard to replicate. In this paper, we propose a completely unsupervised method, mixture invariant training (MixIT), that requires only single-channel acoustic mixtures. In MixIT, training examples are constructed by mixing together existing mixtures, and the model separates them into a variable number of latent sources, such that the separated sources can be remixed to approximate the original mixtures. We show that MixIT can achieve competitive performance compared to supervised methods on speech separation. Using MixIT in a semi-supervised learning setting enables unsupervised domain adaptation and learning from large amounts of real world data without ground-truth source waveforms. In particular, we significantly improve reverberant speech separation performance by incorporating reverberant mixtures, train a speech enhancement system from noisy mixtures, and improve universal sound separation by incorporating a large amount of in-the-wild data.

eess.AS

Sequential Multi-Frame Neural Beamforming for Speech Separation and Enhancement

This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks for separation use an advanced convolutional architecture trained with a novel stabilized signal-to-noise ratio loss function. For beamforming, we explore multiple ways of computing time-varying covariance matrices, including factorizing the spatial covariance into a time-varying amplitude component and a time-invariant spatial component, as well as using block-based techniques. In addition, we introduce a multi-frame beamforming method which improves the results significantly by adding contextual frames to the beamforming formulations. We extensively evaluate and analyze the effects of window size, block size, and multi-frame context size for these methods. Our best method utilizes a sequence of three neural separation and multi-frame time-invariant spatial beamforming stages, and demonstrates an average improvement of 2.75 dB in scale-invariant signal-to-noise ratio and 14.2% absolute reduction in a comparative speech recognition metric across four challenging reverberant speech enhancement and separation tasks. We also use our three-speaker separation model to separate real recordings in the LibriCSS evaluation set into non-overlapping tracks, and achieve a better word error rate as compared to a baseline mask based beamformer.

cs.SD

Universal Sound Separation

Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of different types, a task we refer to as universal sound separation, and it is unknown how performance on speech tasks carries over to non-speech tasks. To study this question, we develop a dataset of mixtures containing arbitrary sounds, and use it to investigate the space of mask-based separation architectures, varying both the overall network architecture and the framewise analysis-synthesis basis for signal transformations. These network architectures include convolutional long short-term memory networks and time-dilated convolution stacks inspired by the recent success of time-domain enhancement networks like ConvTasNet. For the latter architecture, we also propose novel modifications that further improve separation performance. In terms of the framewise analysis-synthesis basis, we explore both a short-time Fourier transform (STFT) and a learnable basis, as used in ConvTasNet. For both of these bases, we also examine the effect of window size. In particular, for STFTs, we find that longer windows (25-50 ms) work best for speech/non-speech separation, while shorter windows (2.5 ms) work best for arbitrary sounds. For learnable bases, shorter windows (2.5 ms) work best on all tasks. Surprisingly, for universal sound separation, STFTs outperform learnable bases. Our best methods produce an improvement in scale-invariant signal-to-distortion ratio of over 13 dB for speech/non-speech separation and close to 10 dB for universal sound separation.

cs.SD

Differentiable Consistency Constraints for Improved Deep Speech Enhancement

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to estimate masks for complex-valued short-time Fourier transforms (STFTs) to suppress noise and preserve speech. However, current masking approaches often neglect two important constraints: STFT consistency and mixture consistency. Without STFT consistency, the system's output is not necessarily the STFT of a time-domain signal, and without mixture consistency, the sum of the estimated sources does not necessarily equal the input mixture. Furthermore, the only previous approaches that apply mixture consistency use real-valued masks; mixture consistency has been ignored for complex-valued masks. In this paper, we show that STFT consistency and mixture consistency can be jointly imposed by adding simple differentiable projection layers to the enhancement network. These layers are compatible with real or complex-valued masks. Using both of these constraints with complex-valued masks provides a 0.7 dB increase in scale-invariant signal-to-distortion ratio (SI-SDR) on a large dataset of speech corrupted by a wide variety of nonstationary noise across a range of input SNRs.

cs.SD

Exploring Tradeoffs in Models for Low-latency Speech Enhancement

We explore a variety of neural networks configurations for one- and two-channel spectrogram-mask-based speech enhancement. Our best model improves on previous state-of-the-art performance on the CHiME2 speech enhancement task by 0.4 decibels in signal-to-distortion ratio (SDR). We examine trade-offs such as non-causal look-ahead, computation, and parameter count versus enhancement performance and find that zero-look-ahead models can achieve, on average, within 0.03 dB SDR of our best bidirectional model. Further, we find that 200 milliseconds of look-ahead is sufficient to achieve equivalent performance to our best bidirectional model.

cs.SD

VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We achieve this by training two separate neural networks: (1) A speaker recognition network that produces speaker-discriminative embeddings; (2) A spectrogram masking network that takes both noisy spectrogram and speaker embedding as input, and produces a mask. Our system significantly reduces the speech recognition WER on multi-speaker signals, with minimal WER degradation on single-speaker signals.

eess.AS