arXiv ScienceSearch

arXiv subjects

Reda Chhaibi

Publications and source records attributed to Reda Chhaibi.

At least 19 recordsLinked to original sources

The Yang-Mills measure on surfaces via Morse theory

We introduce a Morse theoretical approach to the construction of the Yang--Mills measure on the space of connections of a compact Riemannian surface. This provides a direct continuous version of this measure which was previously obtained through lattice approximations by Chevyrev in the case of the flat torus and by one of the authors and Nohra for general compact Riemannian surfaces. The starting point is the new notion of a Morse gauge together with the resolution of random cohomological equations associated to Morse--Smale vector fields. This is achieved by improving exponential convergence to equilibrium results for Morse--Smale gradient flows that were obtained by two of the authors in the context of the study of Ruelle spectra and by Jia, Stewart and Sverak in the context of simplified models from fluid mechanics. Combining these random solutions with the data given by the Morse complex, we introduce a free Yang-Mills measure on space of connections and, using classical tools from stochastic differential equations, we show how to make sense of holonomies for random connections along a large class of curves. Finally, by setting a proper conditioning of this free measure through these random holonomies, we define the Yang--Mills measure and we compute its partition function together with the law of random holonomies with respect to this measure, recovering the formulas from the works of Migdal, Witten and L\'evy.

math.PR

A Martingale Approach To Fluctuations of Rank Estimators in Sensitivity Analysis

Given a bivariate random pair $(X,Y)$, a natural problem is to estimate, from a single sample $(X_i,Y_i)_{1\le i\le n}$, quantities such as $\mathbb{E}\left[ \mathbb{E}[ Y\mid X ]^2 \right]$. More broadly, sensitivity indices are designed to quantify the possibly nonlinear influence of an input variable $X$ on an output variable $Y$. A classical example is the Sobol' index $$ \frac{\mathrm{Var}(\mathbb{E}[Y\mid X])}{\mathrm{Var}(Y)} \in [0,1] \ . $$ Another important example is the Cram\'er--von Mises (CvM) index. Following the pioneering work of Chatterjee \cite{chatterjee2021new}, consistent rank-based estimators are now available for such quantities. In this paper, we prove sharp fluctuation results using martingale methods. Our framework yields a unified treatment of the univariate Sobol' index, a multivariate extension involving several functions of the same scalar input, and the CvM index. As a consequence, we recover, unify, and simplify results from Gamboa et al. \cite{gamboa2022global, gamboa2023erratum}, Lin--Han \cite{lin2022limit}, and Kroll \cite{kroll2024asymptotic}. In particular, we work under minimal regularity assumptions. Furthermore, while the Gaussian fluctuation phenomenon itself was already known, the novelty lies in the structure of the asymptotic variance: for the CvM index, we obtain, to the best of our knowledge, the first explicit formula, while for the Sobol' index, we derive a new expression with a more structured form.

math.ST

Faster Computation of Entropic Optimal Transport via Stable Low Frequency Modes

In this paper, we propose an accelerated version for the Sinkhorn algorithm, which is the reference method for computing the solution to Entropic Optimal Transport. Its main draw-back is the exponential slow-down of convergence as the regularization weakens $\varepsilon \rightarrow 0$. Thanks to spectral insights on the behavior of the Hessian, we propose to mitigate the problem via an original spectral warm-start strategy. This leads to faster convergence compared to the reference method, as also demonstrated in our numerical experiments.

math.NA

Feature Representation Transferring to Lightweight Models via Perception Coherence

In this paper, we propose a method for transferring feature representation to lightweight student models from larger teacher models. We mathematically define a new notion called \textit{perception coherence}. Based on this notion, we propose a loss function, which takes into account the dissimilarities between data points in feature space through their ranking. At a high level, by minimizing this loss function, the student model learns to mimic how the teacher model \textit{perceives} inputs. More precisely, our method is motivated by the fact that the representational capacity of the student model is weaker than the teacher model. Hence, we aim to develop a new method allowing for a better relaxation. This means that, the student model does not need to preserve the absolute geometry of the teacher one, while preserving global coherence through dissimilarity ranking. Importantly, while rankings are defined only on finite sets, our notion of \textit{perception coherence} extends them into a probabilistic form. This formulation depends on the input distribution and applies to general dissimilarity metrics. Our theoretical insights provide a probabilistic perspective on the process of feature representation transfer. Our experiments results show that our method outperforms or achieves on-par performance compared to strong baseline methods for representation transferring.

stat.ML

Convolutional Rectangular Attention Module

In this paper, we introduce a novel spatial attention module that can be easily integrated to any convolutional network. This module guides the model to pay attention to the most discriminative part of an image. This enables the model to attain a better performance by an end-to-end training. In conventional approaches, a spatial attention map is typically generated in a position-wise manner. Thus, it is often resulting in irregular boundaries and so can hamper generalization to new samples. In our method, the attention region is constrained to be rectangular. This rectangle is parametrized by only 5 parameters, allowing for a better stability and generalization to new samples. In our experiments, our method systematically outperforms the position-wise counterpart. So that, we provide a novel useful spatial attention mechanism for convolutional models. Besides, our module also provides the interpretability regarding the \textit{where to look} question, as it helps to know the part of the input on which the model focuses to produce the prediction.

cs.CV

Solvability of the Gaussian Kyle model with imperfect information and risk aversion

We investigate a Kyle model under Gaussian assumptions where a risk-averse informed trader has imperfect information on the fundamental price of an asset. We show that an equilibrium can be constructed by considering an optimal transport problem that is solved under a measure that renders the utility of the informed trader martingale and a filtering problem under the historical measure.

q-fin.TR

Matsumoto-Yor processes on Jordan algebras

The process $(\int_0^t e^{2b_s-b_t}\, ds\ ;\ t\ge 0)$, where $b$ is a real Brownian motion, is known as the geometric 2M-X Matsumoto--Yor process. Remarkably, it enjoys the Markov property. We provide a generalization of this process in the context of Jordan algebras, and we prove the Markov property for this generalization. Our Markov process occurs as a limit of discrete-time AX+B Markov chains on the cone of squares whose invariant probability measures classically yield a Dufresne-type identity for a perpetuity. In particular, the paper provides a generalization to any symmetric cone of the matrix--valued generalization of the Matsumoto--Yor process and Dufresne identity initially developed by Rider--Valk\'o.

math.PR

Training More Robust Classification Model via Discriminative Loss and Gaussian Noise Injection

Robustness of deep neural networks to input noise remains a critical challenge, as naive noise injection often degrades accuracy on clean (uncorrupted) data. We propose a novel training framework that addresses this trade-off through two complementary objectives. First, we introduce a loss function applied at the penultimate layer that explicitly enforces intra-class compactness and increases the margin to analytically defined decision boundaries. This enhances feature discriminativeness and class separability for clean data. Second, we propose a class-wise feature alignment mechanism that brings noisy data clusters closer to their clean counterparts. Furthermore, we provide a theoretical analysis demonstrating that improving feature stability under additive Gaussian noise implicitly reduces the curvature of the softmax loss landscape in input space, as measured by Hessian eigenvalues.This thus naturally enhances robustness without explicit curvature penalties. Conversely, we also theoretically show that lower curvatures lead to more robust models. We validate the effectiveness of our method on standard benchmarks and our custom dataset. Our approach significantly reinforces model robustness to various perturbations while maintaining high accuracy on clean data, advancing the understanding and practice of noise-robust deep learning.

stat.ML

Scattering of the Toda system and the Gaussian $\beta$-ensemble

The classical Toda flow is a well-known integrable Hamiltonian system that diagonalizes matrices. By keeping track of the distribution of entries and precise scattering asymptotics, one can exhibit matrix models for log-gases on the real line. These types of scattering asymptotics date back to fundamental work of Moser. More precisely, using the classical Toda flow acting on symmetric real tridiagonal matrices, we give a "symplectic" proof of the fact that the Dumitriu-Edelman tridiagonal model has a spectrum following the Gaussian $\beta$-ensemble.

nlin.SI

Sensitivity Analysis for Active Sampling, with Applications to the Simulation of Analog Circuits

We propose an active sampling flow, with the use-case of simulating the impact of combined variations on analog circuits. In such a context, given the large number of parameters, it is difficult to fit a surrogate model and to efficiently explore the space of design features. By combining a drastic dimension reduction using sensitivity analysis and Bayesian surrogate modeling, we obtain a flexible active sampling flow. On synthetic and real datasets, this flow outperforms the usual Monte-Carlo sampling which often forms the foundation of design space exploration.

stat.ML

Statistical Edge Detection And UDF Learning For Shape Representation

In the field of computer vision, the numerical encoding of 3D surfaces is crucial. It is classical to represent surfaces with their Signed Distance Functions (SDFs) or Unsigned Distance Functions (UDFs). For tasks like representation learning, surface classification, or surface reconstruction, this function can be learned by a neural network, called Neural Distance Function. This network, and in particular its weights, may serve as a parametric and implicit representation for the surface. The network must represent the surface as accurately as possible. In this paper, we propose a method for learning UDFs that improves the fidelity of the obtained Neural UDF to the original 3D surface. The key idea of our method is to concentrate the learning effort of the Neural UDF on surface edges. More precisely, we show that sampling more training points around surface edges allows better local accuracy of the trained Neural UDF, and thus improves the global expressiveness of the Neural UDF in terms of Hausdorff distance. To detect surface edges, we propose a new statistical method based on the calculation of a $p$-value at each point on the surface. Our method is shown to detect surface edges more accurately than a commonly used local geometric descriptor.

cs.CV

Combining Statistical Depth and Fermat Distance for Uncertainty Quantification

We measure the Out-of-domain uncertainty in the prediction of Neural Networks using a statistical notion called ``Lens Depth'' (LD) combined with Fermat Distance, which is able to capture precisely the ``depth'' of a point with respect to a distribution in feature space, without any assumption about the form of distribution. Our method has no trainable parameter. The method is applicable to any classification model as it is applied directly in feature space at test time and does not intervene in training process. As such, it does not impact the performance of the original model. The proposed method gives excellent qualitative result on toy datasets and can give competitive or better uncertainty estimation on standard deep learning datasets compared to strong baseline methods.

stat.ML

Algebra of Nonlocal Boxes and the Collapse of Communication Complexity

Communication complexity quantifies how difficult it is for two distant computers to evaluate a function f(X,Y), where the strings X and Y are distributed to the first and second computer respectively, under the constraint of exchanging as few bits as possible. Surprisingly, some nonlocal boxes, which are resources shared by the two computers, are so powerful that they allow to collapse communication complexity, in the sense that any Boolean function f can be correctly estimated with the exchange of only one bit of communication. The Popescu-Rohrlich (PR) box is an example of such a collapsing resource, but a comprehensive description of the set of collapsing nonlocal boxes remains elusive. In this work, we carry out an algebraic study of the structure of wirings connecting nonlocal boxes, thus defining the notion of the "product of boxes" $P\boxtimes Q$, and we show related associativity and commutativity results. This gives rise to the notion of the "orbit of a box", unveiling surprising geometrical properties about the alignment and parallelism of distilled boxes. The power of this new framework is that it allows us to prove previously-reported numerical observations concerning the best way to wire consecutive boxes, and to numerically and analytically recover recently-identified noisy PR boxes that collapse communication complexity for different types of noise models.

quant-ph

Estimation of large covariance matrices via free deconvolution: computational and statistical aspects

The estimation of large covariance matrices has a high dimensional bias. Correcting for this bias can be reformulated via the tool of Free Probability Theory as a free deconvolution. The goal of this work is a computational and statistical resolution of this problem. Our approach is based on complex-analytic methods methods to invert $S$-transforms. In particular, one needs a theoretical understanding of the Riemann surfaces where multivalued $S$ transforms live and an efficient computational scheme.

math.PR

To spike or not to spike: the whims of the Wonham filter in the strong noise regime

We study the celebrated Shiryaev-Wonham filter (1964) in its historical setup where the hidden Markov jump process has two states. We are interested in the weak noise regime for the observation equation. Interestingly, this becomes a strong noise regime for the filtering equations. Earlier results of the authors show the appearance of spikes in the filtered process, akin to a metastability phenomenon. This paper is aimed at understanding the smoothed optimal filter, which is relevant for any system with feedback. In particular, we exhibit a sharp phase transition between a spiking regime and a regime with perfect smoothing.

math.PR

A unified approach to informed trading via Monge-Kantorovich duality

We solve a generalized Kyle model type problem using Monge-Kantorovich duality and backward stochastic partial differential equations. First, we show that the the generalized Kyle model with dynamic information can be recast into a terminal optimization problem with distributional constraints. Therefore, the theory of optimal transport between spaces of unequal dimension comes as a natural tool. Second, the pricing rule of the market maker and an optimality criterion for the problem of the informed trader are established using the Kantorovich potentials and transport maps. Finally, we completely characterize the optimal strategies by analyzing the filtering problem from the market maker's point of view. In this context, the Kushner-Zakai filtering SPDE yields to an interesting backward stochastic partial differential equation whose measure-valued terminal condition comes from the optimal coupling of measures.

math.PR

Free Probability for predicting the performance of feed-forward fully connected neural networks

Gradient descent during the learning process of a neural network can be subject to many instabilities. The spectral density of the Jacobian is a key component for analyzing stability. Following the works of Pennington et al., such Jacobians are modeled using free multiplicative convolutions from Free Probability Theory (FPT). We present a reliable and very fast method for computing the associated spectral densities, for given architecture and initialization. This method has a controlled and proven convergence. Our technique is based on an homotopy method: it is an adaptative Newton-Raphson scheme which chains basins of attraction. In order to demonstrate the relevance of our method we show that the relevant FPT metrics computed before training are highly correlated to final test accuracies - up to 85\%. We also nuance the idea that learning happens at the edge of chaos by giving evidence that a very desirable feature for neural networks is the hyperbolicity of their Jacobian at initialization.

stat.ML

Emergence of jumps in quantum trajectories via homogeneization

In the strong noise regime, we study the homogeneization of quantum trajectories i.e. stochastic processes appearing in the context of quantum measurement. When the generator of the average semi-group can be separated into three distinct time scales, we start by describing a homogenized limiting semi-group. This result is of independent interest and is formulated outside of the scope of quantum trajectories. Going back to the quantum context, we show that, in the Meyer-Zheng topology, the time-continuous quantum trajectories converge weakly to the discontinuous trajectories of a pure jump Markov process. Notably, this convergence cannot hold in the usual Skorokhod topology.

math.PR