arXiv ScienceSearch

arXiv subjects

Ajay Jasra

Publications and source records attributed to Ajay Jasra.

At least 19 recordsLinked to original sources

On Bridging Mixture Distributions

In this article we consider bridging between two mixture probability measures. In particular, given access to a Markov kernel between two component distributions, we provide a general mechanism to generate samples from one mixture to the other. Associated to a given reference and extended state space, we prove entropic optimality of this approach. In order to use this idea one needs to know the underlying mixtures and the Markov kernel, which is seldom available, and so we consider the case of Gaussian mixtures and Schrödinger Bridges. We prove a general $2-$Wasserstein continuity bound between the exact bridge and one that is approximated, based on $ε-$covariance inflation, and these rely on a novel continuity analysis of perturbed Riccati maps. We apply our results in the context of bridging mixtures of Gaussians, single Gaussians and empirical estimators of the Gaussian parameters and the Monge map. For mixtures of Gaussians, when the parameters are estimated using the Expectation-Maxmization algorithm, the upper-bound on the $2-$Wasserstein distance between the true and approximated bridges is, under assumptions and with probability at least $1-10N^{-1}$, $\mathcal{O}\big(\big[\big(\tfrac{d\log N}{N}\big)^{1/2}\left\{1+\big(\tfrac{d\log N}{N}\right)^{1/2}(ε^{-2}+1)\big\}+ε^2\big]\big) $ and for the other two cases, in expectation, $\mathcal{O}\left(d\left\{\tfrac{1+ε^{-2}}{1+N}+ε^2\right\}\right)$, where $d,N\in\mathbb{N}$ is the dimension of the Gaussian and the number of empirical samples respectively. We also investigate our bounds numerically.

math.ST

Sequential Markov Chain Monte Carlo for Filtering of State-Space Models with Low or Degenerate Observation Noise

We consider the discrete-time filtering problem in scenarios where the observation noise is degenerate or low. More precisely, one is given access to a discrete time observation sequence which at any time $k$ depends only on the state of an unobserved Markov chain. We specifically assume that the functional relationship between observations and hidden Markov chain has either degenerate or low noise. In this article, under suitable assumptions, we derive the filtering density and its recursions for this class of problems on a specific sequence of manifolds defined through the observation function. We then design sequential Markov chain Monte Carlo methods to approximate the filter serially in time. For a certain linear observation model, we show that using sequential Markov chain Monte Carlo for low noise will converge as the noise disappears to that of using sequential Markov chain Monte Carlo for degenerate noise. We illustrate the performance of our methodology on several challenging stochastic models arising in statistics and applied mathematics.

stat.CO

Score-Based Martingale Posteriors for Deep Neural Networks

In this paper we investigate the efficacy of the score-based martingale posteriors (SMP) (Cui & Walker, 2025; Fong et al., 2023) in the context of modern and large-scale machine learning problems and its potential for meaningful uncertainty quantification. SMPs work with a stochastic gradient ascent-type recursion on the parameter space of stochastic models and construct a martingale on the parameter space. Under simple mathematical assumptions, the recursion can be built so that the parameters form a martingale sequence which possesses a limiting, in time, random variable, the latter of which can be simulated very quickly, in contrast to Monte Carlo-based methods such as Markov chain Monte Carlo. In this expository paper we explore the SMP for inferring the parameters of deep neural networks (DNNs) and, where feasible, compare our results to the state-of-the-art Monte Carlo methods aimed at inferring conventional Bayesian posteriors.

stat.CO

Parameter Estimation for Partially Observed Time-Changed SDEs

In this paper we consider the parameter estimation problem associated to partially-observed time changed SDEs, with observations that are given at discrete times. In particular we consider both likelihood and Bayesian estimation. We develop new Markov chain Monte Carlo (MCMC) algorithms which allow an unbiased score-based stochastic approximation method to provide likelihood-type parameter estimators. We also use a variant of this MCMC algorithm to perform multilevel-based Bayesian parameter estimation. We prove that this latter method achieves a mean square error of $\mathcal{O}(ε^2)$ ($ε>0$) with a cost of $\mathcal{O}(ε^{-2}\log(ε)^2)$. Our methodologies are tested numerically on both simulated and real data.

math.NA

Unbiased Gradients for a Class of Conditional Stochastic Optimization Problems

In this paper we consider the conditional stochastic optimization (CSO) problem. This consists of optimizing a function which can be written as the expectation of a function which is itself a function of a conditional expectation, i.e.~of the type $F(ξ) := \mathbb{E}\left[f\left(Z,\mathbb{E}[g(Z,X,ξ)|Z]\right)\right]$, where precise definitions are given in the main text. We address a particular class of CSO problems where the joint law of the random variables $X,Z$ cannot be exactly sampled; this case has been addressed in Goda & Kitade (2023). We introduce a method that combines Markovian stochastic approximation with unbiased approximation methods which allows one to find the optimizer of $F(ξ)$ in the context of interest. We illustrate our methodology on two examples associated to parameter estimation with model averaging and portfolio selection associated to high-dimensional full factor multivariate stochastic volatility models.

math.OC

Martingale Posteriors for Discretely Observed Diffusions

In this paper we consider parameter estimation for discretely observed diffusion processes. In particular, we focus on data that are observed at low frequency and methodology that can estimate parameters with uncertainty quantification. Most statistical work in this domain develops advanced Markov chain Monte Carlo (MCMC) algorithms for sampling from the posterior of the parameters, a task which is often complicated by the fact that one seldom has access to the transition density of the diffusion process; one has to combine sophisticated MCMC methods which are robust to the required time discretization of the diffusion, which can yield expensive algorithms. We focus on developing the martingale posterior method for the context of interest, when one can only numerically approximate the transition density of the diffusion. Based on using types of diffusion bridges we introduce a new martingale posterior method for parameter estimation for discretely observed diffusion processes. We prove that this algorithm approximates, in some sense, the martingale posterior which has no time-discretization bias up-to $\mathcal{O}(Δ)$ if $Δ$ is the time discretization step. Our approach is illustrated on several examples, showing orders of magnitude speed up versus state-of-the-art MCMC algorithms.

stat.CO

Particle Filtering for a Class of State-Space Models with Low and Degenerate Observational Noise

We consider the discrete-time filtering problem in scenarios where the observation noise is low or degenerate. We focus on the case where the observation equation is a linear function of the state and the data involve additive noise. However, we place minimal assumptions on the hidden state process. For such a class of models we derive new particle filters (PFs) with the key property that their performance is robust to the size of the observation noise. As a consequence, the developed PFs are well-defined in the limiting case of degenerate observation noise. Indicatively, we prove (under assumptions) that the PF applied in this low noise setting inherits the properties of the PF used in the degenerate case. We extend our framework to the case where the hidden states are drawn from a diffusion process. In this scenario we develop new PFs which are robust to both low noise and fine levels of time discretization. We illustrate our algorithms numerically on several examples.

stat.CO

Unbiased Approximations for Stationary Distributions of McKean-Vlasov SDEs

We consider the development of unbiased estimators, to approximate the stationary distribution of Mckean-Vlasov stochastic differential equations (MVSDEs). These are an important class of processes, which frequently appear in applications such as mathematical finance, biology and opinion dynamics. Typically the stationary distribution is unknown and indeed one cannot simulate such processes exactly. As a result one commonly requires a time-discretization scheme which results in a discretization bias and a bias from not being able to simulate the associated stationary distribution. To overcome this bias, we present a new unbiased estimator taking motivation from the literature on unbiased Monte Carlo. We prove the unbiasedness of our estimator, under assumptions. In order to prove this we require developing ergodicity results of various discrete time processes, through an appropriate discretization scheme, towards the invariant measure. Numerous numerical experiments are provided, on a range of MVSDEs, to demonstrate the effectiveness of our unbiased estimator. Such examples include the Currie-Weiss model, a 3D neuroscience model and a parameter estimation problem.

stat.ME

New Trends in the Stability of Sinkhorn Semigroups

Entropic optimal transport problems play an increasingly important role in machine learning and generative modelling. In contrast with optimal transport maps which often have limited applicability in high dimensions, Schrodinger bridges can be solved using the celebrated Sinkhorn's algorithm, a.k.a. the iterative proportional fitting procedure. The stability properties of Sinkhorn bridges when the number of iterations tends to infinity is a very active research area in applied probability and machine learning. Traditional proofs of convergence are mainly based on nonlinear versions of Perron-Frobenius theory and related Hilbert projective metric techniques, gradient descent, Bregman divergence techniques and Hamilton-Jacobi-Bellman equations, including propagation of convexity profiles based on coupling diffusions by reflection methods. The objective of this review article is to present, in a self-contained manner, recently developed Sinkhorn/Gibbs-type semigroup analysis based upon contraction coefficients and Lyapunov-type operator-theoretic techniques. These powerful, off-the-shelf semigroup methods are based upon transportation cost inequalities (e.g. log-Sobolev, Talagrand quadratic inequality, curvature estimates), $ϕ$-divergences, Kantorovich-type criteria and Dobrushin contraction-type coefficients on weighted Banach spaces as well as Wasserstein distances. This novel semigroup analysis allows one to unify and simplify many arguments in the stability of Sinkhorn algorithm. It also yields new contraction estimates w.r.t. generalized $ϕ$-entropies, as well as weighted total variation norms, Kantorovich criteria and Wasserstein distances.

math.PR

Latent Modularity in Multi-View Data

In this article, we consider the problem of clustering multi-view data, that is, information associated to individuals that form heterogeneous data sources (the views). We adopt a Bayesian model and in the prior structure we assume that each individual belongs to a baseline cluster and conditionally allow each individual in each view to potentially belong to different clusters than the baseline. We call such a structure ''latent modularity''. Then for each cluster, in each view we have a specific statistical model with an associated prior. We derive expressions for the marginal priors on the view-specific cluster labels and the associated partitions, giving several insights into our chosen prior structure. Using simple Markov chain Monte Carlo algorithms, we consider our model in a simulation study, along with a more detailed case study that requires several modeling innovations.

stat.ME

Unbiased Parameter Estimation of Partially Observed Diffusions using Diffusion Bridges

In this article we consider the estimation of static parameters for partially observed diffusion processes with discrete-time observations over a fixed time interval. In particular, when one only has access to time-discretized solutions of the diffusions we build upon the works of \cite{ub_par,ub_grad} to devise a method that can estimate the parameters without time-discretization bias. We leverage an identity associated to the gradient of the log-likelihood associated to diffusion bridges, which has not been used before. Contrary to the afore mentioned methods, the diffusion coefficient can depend on the parameters and our approach facilitates the use of more efficient Markov chain sampling algorithms. We prove that our estimator is unbiased with finite variance and demonstrate the efficacy of our methodology in several examples.

stat.ME

Bayesian Deep Learning with Multilevel Trace-class Neural Networks

In this article we consider Bayesian inference associated to deep neural networks (DNNs) and in particular, trace-class neural network (TNN) priors which can be preferable to traditional DNNs as (a) they are identifiable and (b) they possess desirable convergence properties. TNN priors are defined on functions with infinitely many hidden units, and have strongly convergent Karhunen-Loeve-type approximations with finitely many hidden units. A practical hurdle is that the Bayesian solution is computationally demanding, requiring simulation methods, so approaches to drive down the complexity are needed. In this paper, we leverage the strong convergence of TNN in order to apply Multilevel Monte Carlo (MLMC) to these models. In particular, an MLMC method that was introduced is used to approximate posterior expectations of Bayesian TNN models with optimal computational complexity, and this is mathematically proved. The results are verified with several numerical experiments on model problems arising in machine learning, including regression, classification, and reinforcement learning.

stat.CO

Antithetic Multilevel Methods for Elliptic and Hypo-Elliptic Diffusions with Applications

We present a new antithetic multilevel Monte Carlo (MLMC) method for the estimation of expectations with respect to laws of diffusion processes that can be elliptic or hypo-elliptic. In particular, we consider the case where one has to resort to time discretization of the diffusion and numerical simulation of such schemes. Inspired by recent works, we introduce a new MLMC estimator of expectations, which does not require any Lévy area simulation and has a strong error of order 2 and a weak error of order 2. We then show how this approach can be used in the context of the filtering problem associated to partially observed diffusions with discrete time observations. We illustrate that in numerical simulations our new approaches provide efficiency gains for several problems, particularly when the diffusion process is hypo-elliptic, relative to some existing methods.

math.NA

Bayesian Parameter Estimation for Partially Observed McKean-Vlasov Diffusions Using Multilevel Markov chain Monte Carlo

In this article we consider Bayesian estimation of static parameters for a class of partially observed McKean-Vlasov diffusion processes with discrete-time observations over a fixed time interval. This problem features several obstacles to its solution, which include that the posterior density is numerically intractable in continuous-time, even if the transition probabilities are available and even when one uses a time-discretization, the posterior still cannot be used by adopting well-known computational methods such as Markov chain Monte Carlo (MCMC). In this paper we provide a solution to this problem by using new MCMC algorithms which can solve the afore-mentioned issues. This MCMC algorithm is extended to use multilevel Monte Carlo (MLMC) methods. We prove convergence bounds on our parameter estimators and show that the MLMC-based MCMC algorithm reduces the computational cost to achieve a mean square error versus ordinary MCMC by an order of magnitude. We numerically illustrate our results on two models.

stat.CO

Asymptotic Variance in the Central Limit Theorem for Multilevel Markovian Stochastic Approximation

In this note we consider the finite-dimensional parameter estimation problem associated to inverse problems. In such scenarios, one seeks to maximize the marginal likelihood associated to a Bayesian model. This latter model is connected to the solution of partial or ordinary differential equation. As such, there are two primary difficulties in maximizing the marginal likelihood (i) that the solution of differential equation is not always analytically tractable and (ii) neither is the marginal likelihood. Typically (i) is dealt with using a numerical solution of the differential equation, leading to a numerical bias and (ii) has been well studied in the literature using, for instance, Markovian stochastic approximation. It is well-known that to reduce the computational effort to obtain the maximal value of the parameter, one can use a hierarchy of solutions of the differential equation and combine with stochastic gradient methods. Several approaches do exactly this. In this paper we consider the asymptotic variance in the central limit theorem, associated to known estimates and find bounds on the asymptotic variance in terms of the precision of the solution of the differential equation. The significance of these bounds are the that they provide missing theoretical guidelines on how to set simulation parameters; that is, these appear to be the first mathematical results which help to run the methods efficiently in practice.

math.NA

Bayesian Inference for Non-Synchronously Observed Diffusions

We consider the problem of Bayesian inference for bi-variate data observed in time but with observation times which occur non-synchronously. In particular, this occurs in a wide variety of applications in finance, such as high-frequency trading or crude oil futures trading. We adopt a diffusion model for the data and formulate a Bayesian model with priors on unknown parameters along with a latent representation for the the so-called missing data. We then consider computational methodology to fit the model using Markov chain Monte Carlo (MCMC). We have to resort to time-discretization methods as the complete data likelihood is intractable and this can cause considerable issues for MCMC when the data are observed in low frequencies. In a high frequency observation frequencies we present a simple particle MCMC method based on an Euler--Maruyama time discretization, which can be enhanced using multilevel Monte Carlo (MLMC). In the low frequency observation regime we introduce a novel bridging representation of the posterior in continuous time to deal with the issues of MCMC in this case. This representation is discretized and fitted using MCMC and MLMC. We apply our methodology to real and simulated data to establish the efficacy of our methodology.

stat.ME

Unbiased Parameter Estimation for Bayesian Inverse Problems

In this paper we consider the estimation of unknown parameters in Bayesian inverse problems. In most cases of practical interest, there are several barriers to performing such estimation, This includes a numerical approximation of a solution of a differential equation and, even if exact solutions are available, an analytical intractability of the marginal likelihood and its associated gradient, which is used for parameter estimation. The focus of this article is to deliver unbiased estimates of the unknown parameters, that is, stochastic estimators that, in expectation, are equal to the maximize of the marginal likelihood, and possess no numerical approximation error. Based upon the ideas of [4] we develop a new approach for unbiased parameter estimation for Bayesian inverse problems. We prove unbiasedness and establish numerically that the associated estimation procedure is faster than the current state-of-the-art methodology for this problem. We demonstrate the performance of our methodology on a range of problems which include a PDE and ODE.

stat.ME

Multilevel Monte Carlo for a class of Partially Observed Processes in Neuroscience

In this paper we consider Bayesian parameter inference associated to a class of partially observed stochastic differential equations (SDE) driven by jump processes. Such type of models can be routinely found in applications, of which we focus upon the case of neuroscience. The data are assumed to be observed regularly in time and driven by the SDE model with unknown parameters. In practice the SDE may not have an analytically tractable solution and this leads naturally to a time-discretization. We adapt the multilevel Markov chain Monte Carlo method of [11], which works with a hierarchy of time discretizations and show empirically and theoretically that this is preferable to using one single time discretization. The improvement is in terms of the computational cost needed to obtain a pre-specified numerical error. Our approach is illustrated on models that are found in neuroscience.

q-bio.NC