arXiv ScienceSearch

arXiv subjects

Marc Hoffmann

Publications and source records attributed to Marc Hoffmann.

At least 19 recordsLinked to original sources

Rainfall is rough

We propose a new approach to model rainfall by combining heterogeneous data sources at different time scales. Continuous arrivals of rain cells are incorporated into a Hawkes process formalism that encompasses the classical Bartlett-Lewis and Neyman-Scott models, thereby providing a more flexible representation of clustering. Analysis of high frequency rainfall data (at the minute scale over several years) indicates that critical Hawkes processes with heavy-tailed power-law kernels yield a superior fit relative to classical models and alternative kernel specifications. Scaling arguments inspired by Jaisson and Rosenbaum (2016) imply that aggregated rainfall at coarse time scales converges to a rough fractional process with Hurst exponent close to zero. This prediction is supported by empirical evidence from low-frequency data (annual observations spanning centuries to millennia), where the Hurst exponent is estimated to lie between 0.01 and 0.1 based on either direct observations from weather stations or proxy reconstructions such as tree-ring records. These results establish a connection between rainfall dynamics and models developed in quantitative finance for market microstructure and volatility. They also provide a new perspective on classical scaling phenomena originally studied by Hurst and Mandelbrot.

stat.AP

Drift Estimation for Diffusion Processes Using Neural Networks Based on Discretely Observed Independent Paths

This paper addresses the nonparametric estimation of the drift function over a compact domain for a time-homogeneous diffusion process, based on high-frequency discrete observations from $N$ independent trajectories. We propose a neural network-based estimator and derive a non-asymptotic convergence rate, decomposed into a training error, an approximation error, and a diffusion-related term scaling as ${\log N}/{N}$. For compositional drift functions, we establish an explicit rate. In the numerical experiments, we consider a drift function with local fluctuations generated by a double-layer compositional structure featuring local oscillations, and show that the empirical convergence rate becomes independent of the input dimension $d$. Compared to the $B$-spline method, the neural network estimator achieves better convergence rates and more effectively captures local features, particularly in higher-dimensional settings.

stat.ML

A posteriori existence for the Keller-Segel model via a finite volume - finite element scheme

We derive two forms of conditional a posteriori error estimates for a finite volume scheme approximating the parabolic-elliptic Keller-Segel system. The estimates control the error in the $L^\infty(0,T, L^2(\Omega))$- and $L^2(0,T;H^1(\Omega))$-norm and exhibit linear convergence in the mesh size, as observed in numerical experiments. Crucially, we show that, as long as the condition of the error estimate is satisfied, a weak solution exists. This means, as long as the numerical solution has good properties, we can rigorously infer existence of an exact solution.

math.NA

Statistical estimation of a mean-field FitzHugh-Nagumo model

We consider an interacting system of particles with value in $\mathbb{R}^d \times \mathbb{R}^d$, governed by transport and diffusion on the first component, on that may serve as a representative model for kinetic models with a degenerate component. In a first part, we control the fluctuations of the empirical measure of the system around the solution of the corresponding Vlasov-Fokker-Planck equation by proving a Bernstein concentration inequality, extending a previous result of arXiv:2011.03762 in several directions. In a second part, we study the nonparametric statistical estimation of the classical solution of Vlasov-Fokker-Planck equation from the observation of the empirical measure and prove an oracle inequality using the Goldenshluger-Lepski methodology and we obtain minimax optimality. We then specialise on the FitzHugh-Nagumo model for populations of neurons. We consider a version of the model proposed in Mischler et al. arXiv:1503.00492 an optimally estimate the $6$ parameters of the model by moment estimators.

math.ST

Scaling limits for a population model with growth, division and cross-diffusion

Originally motivated by the morphogenesis of bacterial microcolonies, the aim of this article is to explore models through different scales for a spatial population of interacting, growing and dividing particles. We start from a microscopic stochastic model, write the corresponding stochastic differential equation satisfied by the empirical measure, and rigorously derive its mesoscopic (mean-field) limit. Under smoothness and symmetry assumptions for the interaction kernel, we then obtain entropy estimates, which provide us with a localization limit at the macroscopic level. Finally, we perform a thorough numerical study in order to compare the three modeling scales.

math.AP

Entropic Semi-Martingale Optimal Transport

Entropic Optimal Transport (EOT), also referred to as the Schr\"odinger problem, seeks to find a random processes with prescribed initial/final marginals and with minimal relative entropy with respect to a reference measure. The relative entropy forces the two measures to share the same support and only the drift of the controlled process can be adjusted, the diffusion being imposed by the reference measure. Therefore, at first sight, Semi-Martingale Optimal Transport (SMOT) problems (see [1]) seem out of the scope of applications of Entropic regularization techniques, which are otherwise very attractive from a computational point of view. However, when the process is observed only at discrete times, and become therefore a Markov chain, its relative entropy can remain finite even with variable diffusion coefficients, and discrete semi-martingales can be obtained as solutions of (multi-marginal) EOT problems.Given a (smooth) semi-martingale, the limit of the relative entropy of its time discretizations, scaled by the time step converges to the so-called ``specific relative entropy'', a convex functional of its variance process, similar to those used in SMOT.In this paper we use this observation to build an entropic time discretization of continuous SMOT problems. This allows to compute discrete approximations of solutions to continuous SMOT problems by a multi-marginal Sinkhorn algorithm, without the need of solving the non-linear Hamilton-Jacobi-Bellman pde's associated to the dual problem, as done for example in [1, 2]. We prove a convergence result of the time discrete entropic problem to the continuous time problem, we propose an implementation and provide numerical experiments supporting the theoretical convergence.

math.OC

Regularisation for the approximation of functions by mollified discretisation methods

Some prominent discretisation methods such as finite elements provide a way to approximate a function of $d$ variables from $n$ values it takes on the nodes $x_i$ of the corresponding mesh. The accuracy is $n^{-s_a/d}$ in $L^2$-norm, where $s_a$ is the order of the underlying method. When the data are measured or computed with systematical experimental noise, some statistical regularisation might be desirable, with a smoothing method of order $s_r$ (like the number of vanishing moments of a kernel). This idea is behind the use of some regularised discretisation methods, whose approximation properties are the subject of this paper. We decipher the interplay of $s_a$ and $s_r$ for reconstructing a smooth function on regular bounded domains from $n$ measurements with noise of order $\sigma$. We establish that for certain regimes with small noise $\sigma$ depending on $n$, when $s_a > s_r$, statistical smoothing is not necessarily the best option and {\it not regularising} is more beneficial than {\it statistical regularising}. We precisely quantify this phenomenon and show that the gain can achieve a multiplicative order $n^{(s_a-s_r)/(2s_r+d)}$. We illustrate our estimates by numerical experiments conducted in dimension $d=1$ with $\mathbb P_1$ and $\mathbb P_2$ finite elements.

math.NA

A statistical approach for simulating the density solution of a McKean-Vlasov equation

We prove optimal convergence results of a stochastic particle method for computing the classical solution of a multivariate McKean-Vlasov equation, when the measure variable is in the drift, following the classical approach of [BT97, AKH02]. Our method builds upon adaptive nonparametric results in statistics that enable us to obtain a data-driven selection of the smoothing parameter in a kernel type estimator. In particular, we generalise the Bernstein inequality of [DMH21] for mean-field McKean-Vlasov models to interacting particles Euler schemes and obtain sharp deviation inequalities for the estimated classical solution. We complete our theoretical results with a systematic numerical study, and gather empirical evidence of the benefit of using high-order kernels and data-driven smoothing parameters.

math.PR

Nonparametric Bayesian estimation in a multidimensional diffusion model with high frequency data

We consider nonparametric Bayesian inference in a multidimensional diffusion model with reflecting boundary conditions based on discrete high-frequency observations. We prove a general posterior contraction rate theorem in $L^2$-loss, which is applied to Gaussian priors. The resulting posteriors, as well as their posterior means, are shown to converge to the ground truth at the minimax optimal rate over H\"older smoothness classes in any dimension. Of independent interest and as part of our proofs, we show that certain frequentist penalized least squares estimators are also minimax optimal.

math.ST

Statistical inference for rough volatility: Minimax Theory

Rough volatility models have gained considerable interest in the quantitative finance community in recent years. In this paradigm, the volatility of the asset price is driven by a fractional Brownian motion with a small value for the Hurst parameter $H$. In this work, we provide a rigorous statistical analysis of these models. To do so, we establish minimax lower bounds for parameter estimation and design procedures based on wavelets attaining them. We notably obtain an optimal speed of convergence of $n^{-1/(4H+2)}$ for estimating $H$ based on n sampled data, extending results known only for the easier case $H>1/2$ so far. We therefore establish that the parameters of rough volatility models can be inferred with optimal accuracy in all regimes.

math.ST

Statistical inference for rough volatility: Central limit theorems

In recent years, there has been a substantive interest in rough volatility models. In this class of models, the local behavior of stochastic volatility is much more irregular than semimartingales and resembles that of a fractional Brownian motion with Hurst parameter $H < 0.5$. In this paper, we derive a consistent and asymptotically mixed normal estimator of $H$ based on high-frequency price observations. In contrast to previous works, we work in a semiparametric setting and do not assume any a priori relationship between volatility estimators and true volatility. Furthermore, our estimator attains a rate of convergence that is known to be optimal in a minimax sense in parametric rough volatility models.

math.ST

The LAN property for McKean-Vlasov models in a mean-field regime

We establish the local asymptotic normality (LAN) property for estimating a multidimensional parameter in the drift of a system of $N$ interacting particles observed over a fixed time horizon in a mean-field regime $N \rightarrow \infty$. By implementing the classical theory of Ibragimov and Hasminski, we obtain in particular sharp results for the maximum likelihood estimator that go beyond its simple asymptotic normality thanks to H\'ajek's convolution theorem and strong controls of the likelihood process that yield asymptotic minimax optimality (up to constants). Our structural results shed some light to the accompanying nonlinear McKean-Vlasov experiment, and enable us to derive simple and explicit criteria to obtain identifiability and non-degeneracy of the Fisher information matrix. These conditions are also of interest for other recent studies on the topic of parametric inference for interacting diffusions.

math.ST

A continuous multiple hypothesis testing framework for optimal exoplanet detection

When searching for exoplanets, one wants to count how many planets orbit a given star, and to determine what their characteristics are. If the estimated planet characteristics are too far from those of a planet truly present, this should be considered as a false detection. This setting is a particular instance of a general one: aiming to retrieve parametric components in a dataset corrupted by nuisance signals, with a certain accuracy on their parameters. We exhibit a detection criterion minimizing false and missed detections, either as a function of their relative cost or when the expected number of false detections is bounded. If the components can be separated in a technical sense discussed in detail, the optimal detection criterion is a posterior probability obtained as a by-product of Bayesian evidence calculations. Optimality is guaranteed within a model, and we introduce model criticism methods to ensure that the criterion is robust to model errors. We show on two simulations emulating exoplanet searches that the optimal criterion can significantly outperform other criteria. Finally, we show that our framework offers solutions for the identification of components of mixture models and Bayesian false discovery rate control when hypotheses are not discrete.

astro-ph.IM

Individual and population approaches for calibrating division rates in population dynamics: Application to the bacterial cell cycle

Modelling, analysing and inferring triggering mechanisms in population reproduction is fundamental in many biological applications. It is also an active and growing research domain in mathematical biology. In this chapter, we review the main results developed over the last decade for the estimation of the division rate in growing and dividing populations in a steady environment. These methods combine tools borrowed from PDE's and stochastic processes, with a certain view that emerges from mathematical statistics. A focus on the application to the bacterial cell division cycle provides a concrete presentation, and may help the reader to identify major new challenges in the field.

math.AP

Dispersal density estimation across scales

We consider a space structured population model generated by two point clouds: a homogeneous Poisson process $M$ with intensity $n\to\infty$ as a model for a parent generation together with a Cox point process $N$ as offspring generation, with conditional intensity given by the convolution of $M$ with a scaled dispersal density $\sigma^{-1}f(\cdot/\sigma)$. Based on a realisation of $M$ and $N$, we study the nonparametric estimation of $f$ and the estimation of the physical scale parameter $\sigma>0$ simultaneously for all regimes $\sigma=\sigma_n$. We establish that the optimal rates of convergence do not depend monotonously on the scale and we construct minimax estimators accordingly whether $\sigma$ is known or considered as a nuisance, in which case we can estimate it and achieve asymptotic minimaxity by plug-in. The statistical reconstruction exhibits a competition between a direct and a deconvolution problem. Our study reveals in particular the existence of a least favourable intermediate inference scale, a phenomenon that seems to be new.

math.ST

Nonparametric estimation for interacting particle systems : McKean-Vlasov models

We consider a system of $N$ interacting particles, governed by transport and diffusion, that converges in a mean-field limit to the solution of a McKean-Vlasov equation. From the observation of a trajectory of the system over a fixed time horizon, we investigate nonparametric estimation of the solution of the associated nonlinear Fokker-Planck equation, together with the drift term that controls the interactions, in a large population limit $N \rightarrow \infty$. We build data-driven kernel estimators and establish oracle inequalities, following Lepski's principle. Our results are based on a new Bernstein concentration inequality in McKean-Vlasov models for the empirical measure around its mean, possibly of independent interest. We obtain adaptive estimators over anisotropic H\"older smoothness classes built upon the solution map of the Fokker-Planck equation, and prove their optimality in a minimax sense. In the specific case of the Vlasov model, we derive an estimator of the interaction potential and establish its consistency.

math.ST

Estimating the reach of a manifold via its convexity defect function

The reach of a submanifold is a crucial regularity parameter for manifold learning and geometric inference from point clouds. This paper relates the reach of a submanifold to its convexity defect function. Using the stability properties of convexity defect functions, along with some new bounds and the recent submanifold estimator of Aamari and Levrard [Ann. Statist. 47 177-204 (2019)], an estimator for the reach is given. A uniform expected loss bound over a C^k model is found. Lower bounds for the minimax rate for estimating the reach over these models are also provided. The estimator almost achieves these rates in the C^3 and C^4 cases, with a gap given by a logarithmic factor.

math.ST

Testing for high frequency features in a noisy signal

Given nonstationary data, one generally wants to extract the trend from the noise by smoothing or filtering. However, it is often important to delineate a third intermediate category, that we call high frequency (HF) features: this is the case in our motivating example, which consists in experimental measurements of the time-dynamics of depolymerising protein fibrils average size. One may intuitively visualise HF features as the presence of fast, possibly nonstationary and transient oscillations, distinct from a slowly-varying trend envelope. The aim of this article is to propose an empirical definition of HF features and construct estimators and statistical tests for their presence accordingly, when the data consists of a noisy nonstationary 1-dimensional signal. We propose a parametric characterization in the Fourier domain of the HF features by defining a maximal amplitude and distance to low frequencies of significant energy. We introduce a data-driven procedure to estimate these parameters, and compute a p-value proxy based on a statistical test for the presence of HF features. The test is first conducted on simulated signals where the ratio amplitude of the HF features to the level of the noise is controlled. The test detects HF features even when the level of noise is five times larger than the amplitude of the oscillations. In a second part, the test is conducted on experimental data from Prion disease experiments and it confirms the presence of HF features in these signals with significant confidence.

eess.SP