arXiv ScienceSearch

arXiv subjects

Long Feng

Publications and source records attributed to Long Feng.

At least 19 recordsLinked to original sources

Learning CNN Filters via Generalized Stein's Method

Convolutional Neural Networks (CNNs) have undoubtedly revolutionized image data analysis and the field of computer vision. As the cornerstone of CNNs, the convolution operation enables the networks to extract abstract features and uncover hidden relationships in the image data. This paper considers the problem of estimating convolution filters from a statistical perspective using a classical tool --- Stein's formula. We first formulate CNNs into a general index model with matrix-valued input, where convolution filters can be viewed as index vectors. Furthermore, we propose a novel singular value decomposition (SVD) based approach to accurately learn the convolution filters based on a generalized version of the first-order Stein's formula. Theoretical analysis suggests that our estimation achieves an optimal convergence rate, comparable to that of generalized linear models where the link function is known. Extensive simulation studies and real data analyses demonstrate that our approach outperforms popular deep learning algorithms, such as Adam. Notably, our method extends beyond filter estimation and can be applied to nonlinear dimension reduction, providing a viable pathway for representation learning.

stat.AP

Spatial-sign-based multilinear principal component analysis for tensor data

Multilinear principal component analysis (MPCA) reduces the dimension of tensor-valued data while preserving their mode-specific structure, but its quadratic scatter criterion can be unstable under heavy-tailed distributions and contamination. We propose spatial-sign-based multilinear principal component analysis (SMPCA), a robust dimension-reduction method that centers the observations by their spatial median, removes radial magnitude through spatial-sign normalization, and estimates the mode-wise loading spaces by alternating eigendecompositions. Under a separable tensor elliptical model, we show that the target mode-wise loading spaces uniquely maximize the population criterion and that one complete sweep of exact population block updates recovers them from any initialization. We also characterize exactly when their tensor-product subspace coincides with a leading unrestricted subspace of vectorized spatial-sign PCA and, when finite second moments exist, ordinary vectorized PCA. At the sample level, we derive explicit statistical rates for the mode-wise subspaces and the joint multilinear projector, obtain corresponding reconstruction guarantees, establish consistency of the cumulative-contribution dimension selector, and prove that the objective values generated by exact cyclic updates are nondecreasing and convergent. Simulations and an empirical application show that SMPCA is more accurate and stable than competitors under heavy-tailed distributions and outlier contamination, while retaining competitive performance under light-tailed settings.

stat.ME

Factor-Adjusted Location Tests for High-Dimensional Time Series

We study high-dimensional one-sample mean testing for time series with strong common serial dependence driven by latent dynamic factors. After estimating the dynamic factor loading space from lagged autocovariance, we project the data onto its orthogonal complement and construct three factor-adjusted tests: a max test for sparse alternatives, a quadratic test for dense alternatives, and a Cauchy combination test for unknown sparsity. The idiosyncratic component is allowed to be non-Gaussian sub-Gaussian vector white noise. We establish the Gumbel limit of the max statistic, the normal limit and local power function of the quadratic statistic, their asymptotic independence, and the validity of the Cauchy combination. In the strong-factor case, the refined projection expansion shows that the quadratic statistic remains valid for dimensions as large as $p=o(n^2)$. A random-loading residual bootstrap is developed for finite-sample calibration. Simulation studies and a real data application demonstrate reliable size control and competitive power for high-dimensional observations with strong dependence.

stat.ME

Elliptical Regularized Hotelling Tests for High-Dimensional Change-Point Detection

We propose an elliptical regularized Hotelling (ERHT) procedure for detecting location changes in high-dimensional sequences with heavy-tailed, cross-sectionally dependent observations. ERHT contrasts spatial medians on adjacent segments using a ridge-regularized inverse of the pooled centered spatial-sign covariance matrix, thereby combining robustness to radial variation with dependence-aware weighting. We establish Gaussian-process limits for the single- and multiple-change scans and joint convergence over a finite set of regularization parameters. These results provide asymptotically exact calibration of a Cauchy-aggregated adaptive test through the joint Gaussian limit, together with guarantees for local power and single-change localization. We further embed the ERHT score in wild binary segmentation and prove consistency for estimating the number and locations of multiple changes. Simulations show that ERHT is generally well calibrated and delivers competitive power under heavy-tailed distributions, particularly when cross-sectional dependence is substantial. An analysis of the Fama--French 49 industry portfolios reveals persistent evidence of location instability and identifies four structural breaks.

stat.ME

High-Dimensional Change Point Analysis for Temporally Dependent Data

This paper develops adaptive procedures for detecting and locating mean changes in high-dimensional time series. Quadratic CUSUM statistics target dense changes, whereas coordinatewise maximum statistics target sparse changes. Two weighting schemes are considered to accommodate both interior and boundary changes. Under general non-Gaussian vector dependence, we establish the limiting distributions, validate the required centering and scaling estimators, and prove asymptotic independence between matched quadratic and maximum statistics. These results justify Cauchy combination tests. We further establish single-change localization and consistent multiple-change recovery using wild binary segmentation. Numerical results illustrate the effectiveness of the proposed methods.

stat.ME

Adaptive Ridge-Regularized Hotelling Change-Point Tests for Functional Data

We propose a unified ridge-regularized Hotelling framework for detecting and locating mean changes in functional time series. A growing basis expansion converts the functional observations into high-dimensional score vectors. Their long-run covariance is estimated by an edge-corrected difference-based procedure. Ridge regularization stabilizes inference under spectral decay. An explicit local-power formula shows that the power-maximizing ridge depends on the unknown spectral orientation of the change. We therefore combine a family of ridge CUSUM statistics by a Cauchy transform and calibrate the aggregate directly from their joint weighted-bridge limit. For multiple changes, we embed local maximum-ridge statistics in a wild binary segmentation procedure, followed by local refinement. Under mild conditions, we establish the validity, local power, consistency, and localization properties of the proposed tests. In the multiple-change setting, the procedure consistently recovers the number of changes and uniformly estimates their locations. The framework accommodates weak dependence and non-Gaussian functional errors. Simulations and two empirical applications demonstrate the favorable finite-sample performance of the proposed methods.

stat.ME

Elliptical Regularized Hotelling Testing for High Dimensional Data

We consider one-sample testing of a high-dimensional location parameter under elliptically symmetric distributions with heavy tails and pervasive cross-sectional dependence. We propose an elliptical regularized Hotelling test with Cauchy combination (ERHT--CC), based on the sample spatial median and the spatial-sign covariance matrix centered at that median. We derive its null asymptotic normality, consistent estimators of the centering and variance, and an explicit local power function. Since the power-optimal ridge parameter depends on the unknown alternative, we aggregate fixed-ridge $p$-values over a deterministic grid using the Cauchy rule. We establish a finite-grid joint Gaussian limit, justify the analytic combined $p$-value without estimating cross-ridge correlations, and characterize its local power. Simulation studies and an empirical analysis demonstrate the favorable finite-sample performance of ERHT--CC under heavy tails and pervasive dependence.

stat.ME

Cauchy Aggregation of Ridge-Regularized Hotelling Tests for High-Dimensional Change-Point Detection

Ridge-regularized Hotelling-type (RHT) change-point tests depend on a ridge parameter $\lambda$, but the power-optimal value is determined by the unknown covariance structure and the unknown mean shift. We avoid selecting a single ridge value by computing fixed-ridge p-values on a finite deterministic grid and aggregating them with the Cauchy combination rule. Under the standard random-matrix conditions for fixed-ridge RHT statistics, we establish finite-grid joint weak convergence of the ridge processes. This leads to fixed-level validity under joint-limit calibration and small-tail validity for the analytic Cauchy p-value. Monte Carlo experiments show that deterministic-grid Cauchy aggregation has stable size behavior and achieves power close to the best stable fixed ridge choice across a range of covariance and signal configurations.

stat.ME

Rank-Based Tests for Mutual Independence of High-Dimensional Random Vectors via $L_q$ Norm

We consider the problem of testing mutual independence among the components of a high-dimensional random vector. Building on the rank-based max-sum framework, we introduce fixed finite-$L_q$ power-sum statistics under three general classes of rank-based correlations: simple linear rank statistics, non-degenerate rank-based U-statistics and degenerate rank-based U-statistics. The proposed statistics interpolate between the dense-alternative sensitivity of the $L_2$ statistic and the sparse-alternative sensitivity of the $L_\infty$ statistic. We establish the asymptotic independence between any fixed finite-$L_q$ block and the corresponding $L_\infty$ statistic, and combine $L_2,L_4,L_6$ and $L_\infty$ p-values through a Cauchy rule. Numerical studies show that the resulting $L_{2,4,6,\infty}$ procedure is highly robust to the sparsity of the alternative and has strong empirical power across the considered designs.

stat.ME

Adaptive Test for Jump

We develop an adaptive jump test for discretely observed high-frequency semimartingales by combining the A"it-Sahalia--Jacod ratio statistic (A"it-Sahalia and Jacod, 2009) and the Lee--Mykland extreme-return statistic (Lee and Mykland, 2008) with the Cauchy combination rule. Allowing stochastic It^o drift, volatility, and leverage, we show asymptotic independence under the continuous-path null and dense local alternatives, yielding an analytically calibrated test with closed-form power; under finite-activity jumps, the test is consistent. We also extend the method to additive microstructure noise. Simulations show that the combined procedure performs well under both dense and sparse alternatives and is typically best overall.

stat.ME

Federated LoRA Fine-Tuning for LLMs via Collaborative Alignment

Low-rank adaptation (LoRA) has emerged as a powerful tool for parameter-efficient fine-tuning of large language models (LLMs). This paper studies LoRA under a federated learning setting, enabling collaborative fine-tuning across clients while preserving parameter efficiency. We focus on a highly heterogeneous regime in which clients share only partial structure and a substantial subset may be contaminated. We propose Collaborative Low-rank Alignment and Identifiable Recovery (CLAIR), a contamination-aware framework that relies only on preliminary local estimators. Its formulation applies broadly, from linear regression to neural network and LLM modules, whenever local adaptation can be represented by matrix-valued updates. CLAIR recovers the shared LoRA subspace and detects contaminated clients via a structured low-rank plus block-sparse decomposition. We prove exact recovery of the shared LoRA subspace in the noiseless case, stable recovery under preliminary estimation error, and consistent collaborative-set recovery under mild separation conditions. We further quantify the gain from CLAIR refinement: it reduces off-subspace estimation error through cross-client averaging while preserving client-specific variation within the shared LoRA subspace, thus improves over local fine-tuning whenever this oracle gain outweighs the costs of subspace estimation and benign-client heterogeneity. Empirically, we demonstrate the benefits of CLAIR by fine-tuning a Transformer architecture on a text-copying task. The results show accurate contamination detection and improved benign-client performance compared with local fine-tuning and non-robust federated averaging.

stat.ML

Semiparametric Elliptical Mixture Clustering for High-Dimensional Data

Clustering high-dimensional data is especially challenging when cluster distributions are heavy tailed and only approximately elliptical. Existing high-dimensional methods are largely built for Gaussian or other light-tailed models, whereas classical robust elliptical procedures are mostly low dimensional or rely on fully parametric radial families. We propose a semiparametric elliptical mixture clustering framework with cluster-specific centers, an unknown common radial generator, and a common sparse precision-shape matrix, together with a data-driven rule for selecting the number of clusters. A generalized expectation-maximization (GEM) algorithm is developed by combining transformed-radius estimation of the radial generator, radial-score center updates, and a Tyler-POET-GLASSO update for the common precision-shape matrix. The method avoids specifying a parametric radial family and remains computationally feasible in high dimensions. We establish high-dimensional consistency for the estimated model components and the excess misclustering error. Simulation studies and a handwritten-digit application demonstrate the competitive performance and robustness of the proposed method, particularly in heavy-tailed elliptical settings.

stat.ME

FedFrozen: Two-Stage Federated Optimization via Attention Kernel Freezing

Federated learning with heterogeneous clients remains a significant challenge for deep learning, primarily due to client drift arising from inconsistent local updates. Existing federated optimization methods typically address this issue through objective-level regularization or update-correction mechanisms. Recent studies, however, suggest that Transformer-based architectures may be inherently more robust than conventional models under heterogeneous federated training. Motivated by this observation, we investigate how different parameter components within the attention mechanism influence federated optimization. Specifically, we decompose the attention module into a query/key block, which determines the attention kernel, and a value block, which performs semantic transformation under the induced kernel. Based on this perspective, we propose FedFrozen, a two-stage federated optimization framework that first performs full-model warm-up training and then freezes the query/key block while continuing to optimize the value block. Under a linear-attention formulation, we show that the warm-up stage can be interpreted as an inexact descent procedure on a regularized kernel-profile objective, while the frozen stage reduces to a restricted value-block optimization problem under a fixed attention kernel. Our analysis further reveals an explicit trade-off that governs the choice of warm-up length. Simulations validate the predicted bias-drift behavior, and real-data experiments demonstrate that FedFrozen improves both the stability and effectiveness of Transformer models in heterogeneous federated learning.

cs.LG

High-Dimensional Two-Sample Test for Elliptical Symmetry Distribution

We study the high-dimensional two-sample location problem under elliptical symmetry with arbitrary dependence in the scatter matrix. Existing spatial-sign procedures are attractive for heavy-tailed data, but their null calibration is tied to weakly dependent scatter matrices and their diagonal standardization does not, in general, recover the diagonal shape under strong dependence. We propose a new spatial-sign test based on coordinatewise pairwise-difference quantile scales. The new diagonal standardizer is location free, requires no positive moment condition on the radial variable, and estimates the diagonal of the elliptical shape up to a scalar specific to the sample, which disappears after spatial normalization. For the resulting full-sample statistic, we derive an explicit-rate stochastic expansion, establish a general weighted chi-square null distribution under arbitrary correlation structure, justify an empirical diagonal-deletion correction, and show that a Rademacher wild bootstrap consistently estimates the null law. The usual normal approximation appears only as a special case when no eigenvalue dominates.

stat.ME

High-Dimensional Tests for Elliptical Models via Radial--Directional Dependence

We develop high-dimensional goodness-of-fit tests for elliptical models by testing radial--directional independence after affine standardization. The method forms coordinatewise correlations between the log-radius and directional components, using a sum statistic for dense departures, a max statistic for sparse departures, and a Cauchy combination for adaptation. We derive oracle null limits, prove asymptotic independence of the sum and max components under both the null and a balanced local alternative, and establish validity of high-dimensional Hettmansperger--Randles plug-in standardization under explicit perturbation rates. Simulations and data analyses show stable size control, dense--sparse power complementarity, and interpretable coordinate-level diagnostics.

stat.ME

Sparse $K$-spatial-median clustering for high-dimensional data

We propose a robust clustering framework for high-dimensional data with heavy tails and a large fraction of irrelevant variables. The method replaces the mean updates of Lloyd's $K$-means with \emph{spatial medians} to enhance robustness. For the assignment step, it admits either a Euclidean rule for computational simplicity or a robust Mahalanobis-type metric constructed from the spatial sign covariance matrix to account for heterogeneous scales and feature dependence. To handle the $p \gg n$ regime, we further introduce a simple \emph{hard feature-exclusion} mechanism that removes weakly separating dimensions based on across-center dispersion, with the exclusion threshold selected automatically via a permutation-based Gap criterion. Simulation studies under correlated Gaussian and multivariate $t$ models demonstrate that the proposed approach provides competitive clustering accuracy and improved stability relative to $K$-means and sparse $K$-means baselines.

stat.ME

Testing Alpha in High-Dimensional Conditional Time-Varying Factor Models with Dependent Observations

This paper studies alpha testing in a high-dimensional conditional time-varying factor model with temporally dependent observations. Both factor loadings and alpha processes are allowed to vary smoothly over time, and the cross-sectional dimension may be comparable to or larger than the sample size. Using a B-spline sieve method, we develop a sum-type test for dense alternatives, a max-type test for sparse alternatives, and a Cauchy combination test for adaptive inference. On the theoretical side, we derive explicit stochastic expansions for the estimated average alphas, establish asymptotic normality of the sum statistic, and develop the extreme-value limit theory for the max statistic by showing its Gumbel convergence under temporal dependence together with the validity of block-bootstrap calibration. We further prove asymptotic independence between the sum and max statistics and thereby justify the Cauchy combination test. Simulation results demonstrate that the proposed procedures achieve satisfactory size control and competitive power across a wide range of dense and sparse alternatives. An empirical application further illustrates the usefulness of the proposed methods in testing asset-pricing models with time-varying structure.

stat.ME

High-Dimensional Data Analysis for Elliptically Symmetric Distributions

High-dimensional data arise routinely in modern statistics, econometrics, finance, genomics, and machine learning. While a large body of existing methodology is developed under Gaussian or light-tailed assumptions, many real data sets exhibit heavy tails, heterogeneity, and departures from classical covariance-based models. This book provides a systematic treatment of high-dimensional data analysis under elliptically symmetric distributions, with an emphasis on robust inference based on spatial signs, spatial ranks, multivariate Kendall's tau matrices, and related shape-based methods.The book covers the basic theory of elliptical symmetry, high-dimensional location inference, estimation and testing for covariance and precision matrices, sphericity and proportionality testing, high-dimensional alpha testing in factor pricing models, change-point analysis, white-noise and independence testing, high-dimensional discriminant analysis, and dimension reduction through principal component analysis and factor models. Throughout, we review classical low-dimensional and high-dimensional benchmark methods and then develop robust alternatives tailored to elliptical models. Particular attention is paid to the interplay between sum-type, max-type, and adaptive procedures, as well as to the role of scatter, shape, and rank-based dependence measures in heavy-tailed settings. This book is intended as a unified overview of robust high-dimensional methods under elliptical symmetry and as a synthesis of the author's recent research contributions in this area. It is written for researchers and graduate students in statistics, econometrics, and related fields who are interested in modern high-dimensional inference beyond the Gaussian paradigm.

stat.ME