arXiv ScienceSearch

arXiv subjects

Jacques Ravel

Publications and source records attributed to Jacques Ravel.

5 recordsLinked to original sources

From Global to Local Correlation: Geometric Decomposition of Statistical Inference

Understanding feature-outcome associations in high-dimensional data remains challenging when relationships vary across subpopulations, yet standard methods assuming global associations miss context-dependent patterns, reducing statistical power and interpretability. We develop a geometric decomposition framework offering two strategies for partitioning inference problems into regional analyses on data-derived Riemannian graphs. Gradient flow decomposition uses path-monotonicity-validated discrete Morse theory to partition samples into gradient flow cells where outcomes exhibit monotonic behavior. Co-monotonicity decomposition utilizes vertex-level coefficients that provide context-dependent versions of the classical Pearson correlation: these coefficients measure edge-based directional concordance between outcome and features, or between feature pairs, defining embeddings of samples into association space. These embeddings induce Riemannian k-NN graphs on which biclustering identifies co-monotonicity cells (coherent regions) and feature modules. This extends naturally to multi-modal integration across multiple feature sets. Both strategies apply independently or jointly, with Bayesian posterior sampling providing credible intervals.

stat.ME

Adaptive Geometric Regression for High-Dimensional Structured Data

We present a geometric framework for regression on structured high-dimensional data that shifts the analysis from the ambient space to a geometric object capturing the data's intrinsic structure. The method addresses a fundamental challenge in analyzing datasets with high ambient dimension but low intrinsic dimension, such as microbiome compositions, where traditional approaches fail to capture the underlying geometric structure. Starting from a k-nearest neighbor covering of the feature space, the geometry evolves iteratively through heat diffusion and response-coherence modulation, concentrating mass within regions where the response varies smoothly while creating diffusion barriers where the response changes rapidly. This iterative refinement produces conditional expectation estimates that respect both the intrinsic geometry of the feature space and the structure of the response.

stat.ME

Intrinsic and Normal Mean Ricci Curvatures: A Bochner--Weitzenboeck Identity for Simple d-Vectors

We introduce two pointwise subspace averages of sectional curvature on a d-dimensional plane Pi in T_p M: (i) the intrinsic mean Ricci (the average of sectional curvatures of 2-planes contained in Pi); and (ii) the normal (mixed) mean Ricci (the average of sectional curvatures of 2-planes spanned by one vector in Pi and one in Pi^perp). Using Jacobi-field expansions, these means occur as the r^2/6 coefficients in the intrinsic (d-1)-sphere and normal (n-d-1)-sphere volume elements. A direct consequence is a Bochner--Weitzenboeck identity for simple d-vectors V (built from an orthonormal frame X_1,...,X_d with Pi = span{X_i}): the curvature term equals d(n-d) times the normal mean Ricci of Pi. This yields two immediate applications: (a) a Bochner vanishing criterion for harmonic simple d-vectors under a positive lower bound on the normal mean Ricci; and (b) a Lichnerowicz-type lower bound for the first eigenvalue of the Hodge Laplacian on simple d-eigenfields.

math.DG

The Geometry of Machine Learning Models

This paper presents a mathematical framework for analyzing machine learning models through the geometry of their induced partitions. By representing partitions as Riemannian simplicial complexes, we capture not only adjacency relationships but also geometric properties including cell volumes, volumes of faces where cells meet, and dihedral angles between adjacent cells. For neural networks, we introduce a differential forms approach that tracks geometric structure through layers via pullback operations, making computations tractable by focusing on data-containing cells. The framework enables geometric regularization that directly penalizes problematic spatial configurations and provides new tools for model refinement through extended Laplacians and simplicial splines. We also explore how data distribution induces effective geometric curvature in model partitions, developing discrete curvature measures for vertices that quantify local geometric complexity and statistical Ricci curvature for edges that captures pairwise relationships between cells. While focused on mathematical foundations, this geometric perspective offers new approaches to model interpretation, regularization, and diagnostic tools for understanding learning dynamics.

cs.LG

A New Approach to Compositional Data Analysis using \(L^{\infty}\)-normalization with Applications to Vaginal Microbiome

We introduce a novel approach to compositional data analysis based on $L^{\infty}$-normalization, addressing challenges posed by zero-rich high-throughput data. Traditional methods like Aitchison's transformations require excluding zeros, conflicting with the reality that omics datasets contain structural zeros that cannot be removed without violating inherent biological structures. Such datasets exist exclusively on the boundary of compositional space, making interior-focused approaches fundamentally misaligned. We present a family of $L^p$-normalizations, focusing on $L^{\infty}$-normalization due to its advantageous properties. This approach identifies compositional space with the $L^{\infty}$-simplex, represented as a union of top-dimensional faces called $L^{\infty}$-cells. Each cell consists of samples where one component's absolute abundance equals or exceeds all others, with a coordinate system identifying it with a d-dimensional unit cube. When applied to vaginal microbiome data, $L^{\infty}$-decomposition aligns with established Community State Types while offering advantages: each $L^{\infty}$-CST is named after its dominating component, has clear biological meaning, remains stable under sample changes, resolves cluster-based issues, and provides a coordinate system for exploring internal structure. We extend homogeneous coordinates through cube embedding, mapping data into a d-dimensional unit cube. These embeddings can be integrated via Cartesian product, providing unified representations from multiple perspectives. While demonstrated through microbiome studies, these methods apply to any compositional data.

stat.CO