arXiv ScienceSearch

SEARCH · arXiv Science

Results for “stat.AP”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

249 records · Page 3Linked to original sources

Spatial symmetry invariance of solution of Kolmogorov flow

We prove a mathematical theorem that solution for all $t > 0$ of the two-dimensional (2D) Kolmogorov flow governed by Navier-Stokes (NS) equations with periodic boundary condition keeps the same spatial symmetry as its smooth initial condition. The proof of a similar theorem for the three-dimensional NS equations is given in the appendix. These mathematical theorems can be used to check the correctness and reliability of numerical simulations of NS turbulence. For example, they support the corresponding CNS (clean numerical simulation) results of the 2D and 3D turbulent Kolmogorov flows [1-3] that remain the same spatial symmetry in the whole time interval of simulation, but do not support the corresponding DNS (direct numerical simulation) results that lose the spatial symmetry quickly. In other words, these DNS results violate these mathematical theorems. Thus, these mathematical theorems rigorously confirm that the spatiotemporal trajectories of NS turbulence given by DNS are indeed quickly polluted by numerical noises badly. All of these indicate that CNS can indeed provide helpful enlightenments to deepen our understanding about turbulence and besides approach some mathematical truths about NS equations.

physics.flu-dyn

Hybrid Event Frame Sensors: Modeling, Calibration, and Simulation

Hybrid event-frame sensors integrate an Event Vision Sensor (EVS) and an Active Pixel Sensor (APS) within a single chip, combining the high dynamic range and low latency of the EVS with the rich spatial intensity information from the APS. While this tight integration offers compact and temporally precise imaging, the complex circuit architecture introduces nontrivial noise patterns that remain poorly understood and unmodeled. In this work, we present the first unified statistics-based imaging noise model that jointly describes the noise behavior of APS and EVS pixels. Our formulation explicitly incorporates photon shot noise, dark current noise, fixed-pattern noise, and quantization noise, and links EVS noise to illumination level and dark current. Based on this formulation, we further develop a calibration pipeline to estimate noise parameters from real data and provide a detailed analysis of both APS and EVS noise behaviors. Finally, we propose H-ESIM, a statistically grounded simulator that generates RAW frames and events under realistic jointly calibrated noise statistics. Experiments on two hybrid sensors validate our model across multiple imaging tasks, including video frame interpolation and deblurring, demonstrating strong transfer from simulation to real data.

cs.CV

Inverse Source Problem for a Time-Fractional Diffusion-Wave Equation with a Singular Inverse-Square Potential

This paper investigates an inverse source problem for a time-fractional diffusion-wave equation with a singular inverse-square potential. The source term is assumed to consist of a known temporal factor and an unknown spatial component, which is to be recovered from terminal-state measurements. The well-posedness and regularity of the forward problem are established within an appropriate energy framework by exploiting Hardy-type inequalities and the spectral properties of the associated singular elliptic operator. The terminal observation operator is then shown to be compact, and uniqueness of the spatial source is established under a suitable nondegeneracy condition on the temporal factor. To stabilize the resulting ill-posed inverse problem, a Tikhonov regularization approach is introduced. The gradient of the regularized functional is derived through an adjoint problem involving a right-sided fractional derivative, leading to an adjoint-based conjugate gradient method with an exact line search for the numerical reconstruction of the unknown source. Numerical experiments are conducted on both one-and two-dimensional spatial domains, using both exact and noisy terminal data, to demonstrate the effectiveness and stability of the proposed source reconstruction method.

math.NA

Yield Trajectory Tracking for Hyperbolic Age-Structured Population Systems

For population systems modeled by age-structured hyperbolic partial differential equations (PDEs) that are bilinear in the input and evolve with a positive-valued infinite-dimensional state, global stabilization of constant yield set points was achieved in prior work. Seasonal demands in biotechnological production processes give rise to time-varying yield references. For the proposed control objective aiming at a global attractivity of desired yield trajectories, multiple non-standard features have to be considered: a non-local boundary condition, a PDE state restricted to the positive orthant of the function space and arbitrary restrictive but physically meaningful input constraints. Moreover, we provide Control Lyapunov Functionals ensuring an exponentially fast attraction of adequate reference trajectories. To achieve this goal, we make use of the relation between first-order hyperbolic PDEs and integral delay equations leading to a decoupling of the input-dependent dynamics and the infinite-dimensional internal one. Furthermore, the dynamic control structure does not necessitate exact knowledge of the model parameters or online measurements of the age-profile. With a Galerkin-based numerical simulation scheme using the key ideas of the Karhunen-Loève-decomposition, we demonstrate the controller's performance.

math.OC

Computational Oncology of Chemotaxis-Driven Tumour--Immune Spatial Patterning and Stability

We develop a reaction--diffusion--chemotaxis model for spatial tumour--immune--chemokine dynamics that couples logistic tumour growth, immune-mediated killing, chemokine-dependent immune recruitment, chemotactic migration, and signal production. For the nondimensional system, we establish local classical solvability, nonnegativity, a uniform tumour-density bound, and global mass estimates for the immune and chemokine components. The tumour-free equilibrium is stable precisely when the baseline immune-control index satisfies \(σ_0/δ>1\), whereas positive homogeneous coexistence is characterized by a scalar nonlinear equation. Linearization in the Neumann Laplacian eigenbasis yields a mode-dependent cubic dispersion relation, showing that chemotaxis does not alter the tumour-invasion threshold but can destabilize homogeneous coexistence through a finite-wavelength oscillatory instability above a critical sensitivity \(ξ_c\). A conservative finite-volume discretization with upwind chemotactic fluxes and implicit backward differentiation formula time integration is used to test these predictions. Numerical experiments recover the analytical equilibria and growth rates, identify the dominant unstable mode, reproduce the transition to spatial heterogeneity, and quantify the effects of immune recruitment, decay, and diffusion on the stability boundary. Grid-refinement, mass-balance, residual, and nonnegativity diagnostics support the computational reliability of the results.

math.AP

The $α$-Limit Problem: Convergence of a Linear Degenerate Interface Transmission Problem

We study the singular limit of a family of linear degenerate interface transmission problems arising from a regularization procedure in the newly proposed Two-Parameter Diffuse Domain Method (DDM2p). For $α>0$, the regularized problem admits a strictly convex variational formulation on $H^{1}(Ω)$. In the limit $α\to0$, the problem degenerates to a weakly coupled interface system with a nonstandard energy structure. To characterize the limit, we introduce a closed Hilbert subspace $\mathcal{H}\subset H^{1}(Ω)$, defined through an auxiliary Helmholtz problem on an annular subdomain $Ω_2\subset Ω$, and identify the limiting energy functional $\mathcal{E}_{0}$ on $\mathcal{H}$. We prove that the regularized energies $\mathcal{E}_α$ $Γ$-converge to $\mathcal{E}_{0}$ in the strong $L^{2}(Ω)$ topology, using the standard framework. Consequently, minimizers of $\mathcal{E}_α$ converge to the unique minimizer of $\mathcal{E}_{0}$, which is shown to be equivalent to the solution of the limiting interface problem. We further prove strong convergence $u_α\to u_{0}$ in $H^{1}(Ω)$ and establish an $O(α)$ convergence rate. Numerical experiments in one spatial dimension confirm the predicted first-order convergence rate and suggest that this rate is sharp.

math.AP

L2RDaS: Synthesizing 4D Radar Tensors for Model Generalization via Dataset Expansion

4-dimensional (4D) radar is increasingly adopted in autonomous driving for perception tasks, owing to its robustness under adverse weather conditions. To better utilize the spatial information inherent in 4D radar data, recent deep learning methods have transitioned from using sparse point cloud to 4D radar tensors. However, the scarcity of publicly available 4D radar tensor datasets limits model generalization across diverse driving scenarios. Previous methods addressed this by synthesizing radar data, but the outputs did not fully exploit the spatial information characteristic of 4D radar. To overcome these limitations, we propose LiDAR-to-4D radar data synthesis (L2RDaS), a framework that synthesizes spatially informative 4D radar tensors from LiDAR data available in existing autonomous driving datasets. L2RDaS integrates a modified U-Net architecture to effectively capture spatial information and an object information supplement (OBIS) module to enhance reflection fidelity. This framework enables the synthesis of radar tensors across diverse driving scenarios without additional sensor deployment or data collection. L2RDaS improves model generalization by expanding real datasets with synthetic radar tensors, achieving an average increase of 4.25\% in ${{AP}_{BEV}}$ and 2.87\% in ${{AP}_{3D}}$ across three detection models. Additionally, L2RDaS supports ground-truth augmentation (GT-Aug) by embedding annotated objects into LiDAR data and synthesizing them into radar tensors, resulting in further average increases of 3.75\% in ${{AP}_{BEV}}$ and 4.03\% in ${{AP}_{3D}}$. The implementation will be available at https://github.com/kaist-avelab/K-Radar.

cs.CV

A Tensor Neural Network Method for High-Order Homogenization of Locally Periodic Elliptic Problems

We develop a high-order tensor neural network (TNN) method for locally periodic elliptic multiscale problems of the form $-\nabla\cdot(A(x,x/\varepsilon)\nabla u_\varepsilon)=f$. Because the coefficient depends on both the slow variable $x$ and the fast periodic variable $y=x/\varepsilon$, the high-order cell problems and macroscopic corrector equations are more involved than in the classical case $A=A(y)$, and the correctors depend parametrically on $x$. We derive a computable high-order two-scale expansion and prove an $H^1$ convergence estimate for the partial expansion in boundary-layer-free settings, including periodic domains and ideal boundary-matching configurations. The proof uses the recursive compatibility structure of the corrector hierarchy and a zero-mean oscillation estimate in $H^{-1}$. We then construct a TNN framework for the high-dimensional corrector problems. Its tensor-product structure permits deterministic one-dimensional quadrature for the cell problems, homogenized coefficients, macroscopic source terms, and loss functions, avoiding Monte Carlo integration error. The numerical realization assumes that the coefficient entries and assembled data admit finite or controlled tensor-product representations; this computational assumption is separate from the general matrix-valued coefficient class used in the analysis. Experiments with scalar locally periodic coefficients show accurate high-order correctors. The $H^1$ semi-norm errors are consistent with the proved estimate, while point-normalized $L^2$ errors display the nominal high-order behavior predicted by the formal expansion.

math.NA

Shape Holomorphy and Sparse Approximation of the Maxwell Electric Field Integral Operator

Uncertainty quantification for time-harmonic Maxwell scattering by obstacles of uncertain shape needs more than holomorphic dependence of the scattered field: for a boundary element method it is the boundary integral operator family itself that must depend holomorphically on the shape parameters. Two obstructions stand in the way. The natural energy space of the electric field integral equation, $\boldsymbol H^{-1/2}_{\mathrm{div}_Γ}(Γ)$, depends on the geometry, and the available operator-valued shape-holomorphy theory for weakly singular kernels is set in $L^2$, which does not reach it. We remove both. A surface contravariant Piola transformation identifies the geometry-dependent Maxwell trace spaces with a fixed reference space, and in the pulled-back variational formulation the surface Jacobians cancel exactly. The principal analytical ingredient is then a uniform fractional mapping theorem $H^{-1/2}\to H^{1/2}$ for the complex-deformed scalar single-layer family on uniformly $C^{1,1}$ surfaces, obtained by realizing the Laplace principal part as the trace of a complex-coefficient Newton problem on a fixed ambient space. The pulled-back operators are consequently $(\bm b,p,\eps)$-holomorphic for $\bm b\in\ell^p(\N)$, $0<p<1$, and pointwise exclusion of interior electric resonances over the compact real parameter set yields uniform invertibility. Legendre coefficients are therefore $\ell^p$ summable, so the operator family, the surface current and the far field all admit sparse polynomial approximations at dimension-independent best $N$-term rates. These statements are for the operator family itself in its energy-space operator norm, not only for individual solutions.

math.NA

Projection-based low-rank assembly in IgA

Isogeometric Analysis (IgA) uses the same spline functions to represent the computational domain and to approximate the solution. This allows exact geometry descriptions, but the resulting mass and stiffness matrices are expensive to assemble and to store, especially in three dimensions. We present a projection-based low-rank approach for assembling the mass and stiffness tensors of orientation-preserving tensor-product B-spline geometries. For the mass tensor, we exploit the polynomial structure of the determinant of the Jacobian of the geometry map and represent it in reduced spline product spaces by univariate coefficient transfer operators; with exact quadrature and without truncation, the resulting low-rank tensor is an exact reformulation of the standard Galerkin mass tensor. For the stiffness tensor, we split the rational weight function into a polynomial numerator, again represented in reduced spline product spaces, and the reciprocal determinant, which is in general not a spline function and is therefore approximated by an $L^2$-projection onto a tensor-product spline space. Both constructions are carried out entirely in the tensor-train (TT) format, with the projection system solved by the alternating minimal energy (AMEn) method, so that full high-order coefficient tensors are never formed and the multidimensional integrals reduce to univariate integrals and contracted products. The method is implemented in MATLAB using GeoPDEs and the TT-Toolbox. Numerical experiments show that it is competitive with full assembly and with the interpolation-based low-rank method, and that it applies in two situations in which interpolation is problematic: a singular interpolation system and nearly singular geometries. The construction is restricted to orientation-preserving tensor-product B-spline geometries and does not cover NURBS.

math.NA

Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood. Modern PE methods such as RoPE still struggle on tasks such as long-context understanding or retrieval \cite{chen-etal-2025-hope}. Hence, a better understanding of the internal positional mechanism could help design better PE. Building on evidence that positional and semantic signals occupy nearly orthogonal subspaces in trained Transformers, we modify an encoder Transformer to process three explicitly disentangled streams: semantic, absolute positional (AP) and relative positional (RP), and confine the masked-language-modeling (MLM) objective to the semantic stream. This decoupling enables a clean mechanistic study and yields three take-aways. (1) The isolated AP subspace spontaneously collapses into a low-frequency two-dimensional manifold that captures the structure of the document; (2) Attention heads specialize into structure and semantic-oriented groups, with RP exclusively supporting the latter; (3) Standard positional encodings do not robustly retain macroscopic structure: RoPE and RP only weakly encode it, and entangled AP loses it in the final layers under MLM pressure. The disentangled approach preserves positional encoding, which improves linguistic representation on 49 of the 65 linguistic phenomena of the Flash-Holmes probing benchmark.

cs.CL

A simple derivation of the Kalman filter

In this lecture note, we present a concise and self-contained derivation of the discrete-time Kalman filter equations that requires only a basic understanding of least squares estimation. The treatment is designed to minimize mathematical overhead while preserving both rigor and generality.

math.OC

Clustering Three-Way Data with Outliers

Matrix-variate distributions are a relatively recent addition to the model-based clustering literature, thereby making it possible to analyze data in matrix form with complex structure such as images and time series. Due to its recent appearance, there is limited literature on matrix-variate data, with even less on dealing with outliers in these models. An approach for clustering matrix-variate normal data with outliers is discussed. The approach, which uses the distribution of subset log-likelihoods, extends the OCLUST algorithm to matrix-variate normal data and uses an iterative approach to detect and trim outliers.

stat.ML

Probabilistic Symbolic Regression for Equation Discovery via Operator-induced and Regularized Symbolic Forests

Symbolic regression has emerged as a powerful tool for artificial intelligence-driven scientific discovery by learning interpretable analytical expressions that reveal governing relationships directly from data. Existing methods, however, often rely on heuristic search, struggle to balance predictive accuracy with expression complexity in noisy settings, and offer limited characterization of symbolic uncertainty. Probabilistic approaches that address these challenges in a unified manner remain underexplored. We introduce a probabilistic symbolic regression framework that represents mathematical expressions as ensembles of symbolic trees. A regularizing prior over tree topology controls expression complexity, while an Occam's window-based posterior summary captures uncertainty across multiple plausible symbolic models. Given the limited existing theoretical treatment of symbolic regression, we develop posterior concentration guarantees when symbolic expressions approximate the underlying relationship arbitrarily well, with a near-parametric rate when an exact finite formula exists. Additionally, we establish a sharp oracle concentration result under symbolic misspecification. Comparisons of our proposed framework with state-of-the-art competitors demonstrate superior predictive accuracy, optimal symbolic complexity, and stable structural recovery when learning benchmark scientific equations, together with the identification of scientifically interpretable descriptor formulas in a challenging materials discovery application.

stat.ME

Symmetry-driven embedding of networks in hyperbolic space

Hyperbolic models are known to produce networks with properties observed empirically in most network datasets, including heavy-tailed degree distribution, high clustering, and hierarchical structures. As a result, several embeddings algorithms have been proposed to invert these models and assign hyperbolic coordinates to network data. Current algorithms for finding these coordinates, however, do not quantify uncertainty in the inferred coordinates. We present BIGUE, a Markov chain Monte Carlo (MCMC) algorithm that samples the posterior distribution of a Bayesian hyperbolic random graph model. We show that the samples are consistent with current algorithms while providing added credible intervals for the coordinates and all network properties. We also show that some networks admit two or more plausible embeddings, a feature that an optimization algorithm can easily overlook.

stat.CO

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.

stat.ME

Diffusion Models in Simulation-Based Inference: A Tutorial Review

Diffusion models have recently emerged as powerful learners for simulation-based inference (SBI), enabling fast and accurate estimation of latent parameters from simulated and real data. Their score-based formulation offers a flexible way to learn conditional or joint distributions over parameters and observations, thereby providing a versatile solution to various modeling problems. In this tutorial review, we synthesize recent developments on diffusion models for SBI, covering design choices for training, inference, and evaluation. We highlight opportunities created by various concepts such as guidance, score composition, flow matching, consistency models, and joint modeling. Furthermore, we discuss how efficiency and statistical accuracy are affected by noise schedules, parameterizations, and samplers. Finally, we illustrate these concepts with case studies across parameter dimensionalities, simulation budgets, and model types, and outline open questions for future research.

stat.ML

Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for large language model (LLM) outputs, where we test whether an output $X$ generated from a source text $Z$ carries information about an attribute $Y$ beyond $Z$ itself. For this purpose, we propose embedded CITs (eCITs), which embed $X$ and $Z$ and apply an existing CIT to the resulting representations and to $Y$. We show that, provided the embedding of $Z$ is sufficient, i.e. retains the information $Z$ carries about either $Y$ or the representation of $X$, the null hypothesis transfers from $X$ and $Z$ to their representations, so that a CIT valid for the embedded hypothesis is valid for the original one. We further give conditions for equivalence of the two hypotheses, and show that sufficiency weakens to mean sufficiency when the embedded test targets conditional mean independence. We propose a semi-synthetic simulation design to assess type I error (T1E) control and power of the eCITs for given embedding maps on a specific dataset and task, and use it to evaluate them on our application. Applying the eCITs to German Parliament speeches, we find for all combinations of embedding maps considered that the summaries of two LLMs contain information about the speaker's faction and gender beyond the speech they were generated from.

stat.ML