arXiv ScienceSearch

arXiv subjects

Supratik Basu

Publications and source records attributed to Supratik Basu.

5 recordsLinked to original sources

Extended rank regression for all ordinal data

The accuracy of inference from a regression model depends largely on how well the model represents the relationship between the mean and variance of the outcomes. As this relationship is rarely of direct interest, it is natural to treat it as a nuisance parameter, rather than attempt to estimate it. We take this approach in the context of a monotonically transformed linear regression model using a pseudo-likelihood based on an extended notion of ranks. This approach can accommodate a wide range of mean-variance relationships and any ordinal data type, including continuous and discrete ordered data, and requires no estimation or prior specification of the transformation, or decision to treat an outcome as continuous or discrete. We show that the extended rank likelihood incurs no asymptotic information loss at the two extremes of continuous and binary data, and that rank-based prediction intervals can obtain approximate coverage control conditional on the features. Bayesian parameter estimates and prediction intervals are available via a simple Gibbs sampling algorithm. For settings where the model is in doubt, conformal calibration of the Bayesian predictive distribution provides intervals with guaranteed marginal frequentist coverage.

stat.ME

Universally Optimal Robustness-Efficiency Tradeoffs for a General Class of Minimum Divergence Estimators

Balancing the efficiency of an estimator under ideal conditions against its robustness under contamination remains a central challenge in robust statistics. While minimum divergence methods offer a flexible alternative to traditional M-estimation, choosing the appropriate discrepancy measure has historically relied on heuristic or empirical justifications. This manuscript introduces a rigorous optimality criterion for this selection process. By investigating the comprehensive Generalized Alpha-Beta Divergence (GABD) family, we explicitly characterize the Pareto frontier dictating the lowest possible asymptotic variance for any strictly enforced asymptotic breakdown point. Our main theoretical results establish that the estimator achieving this mathematical optimum invariably falls within the extended $(\phi, \gamma)$-divergence class. Crucially, the derived optimal tuning parameter, $\phi^*$, given other parameters, depends solely on the desired breakdown threshold and is entirely invariant to both the assumed parametric model and the exact nature of the data contamination. Supported by comprehensive derivations of asymptotic normality, influence functions, and breakdown thresholds for both continuous and discrete settings, this work offers a unified, theoretical resolution to the long-standing problem of optimal divergence selection in robust inference.

math.ST

Characterization of Generalized Alpha-Beta Divergence and Associated Entropy Measures

Minimum divergence estimators provide a natural framework for robust (parametric) statistical inference. Useful properties of several such divergence measures, including, the Hellinger distance, the power divergence, the density power divergence, the logarithmic density power divergence, etc., have been established in the literature; many of them lead to estimators with high statistical efficiency, sometimes even full asymptotic efficiency. The notable success of these divergences as tools of parametric inference motivates us to explore possible extensions of the alpha-beta divergence family, leading to a superfamily of divergence measures called the ``generalized alpha-beta (GAB) divergences''. This family contains all the aforementioned popular divergence measures as special cases, and additionally provides opportunities to discover new and novel classes of divergences that generate estimators having strong robustness properties without allowing a significant drop in statistical efficiency in various applications. In this paper, we provide the necessary and sufficient conditions for the validity of these generalized divergence measures that enable us to employ them for improved statistical inference. We also show various characterizing properties like duality, inversion, semi-continuity, etc., for the general class of GAB divergences. A discussion on the entropy measure derived from this general family and its properties are also presented along with the associated maximum entropy principle. The class of GAB divergences provide a delicate balance between local and global robustness, and this is illustrated by two examples of robust parameter estimation under the Geometric and the normal scale models.

math.ST

Maximal Inequalities for Independent Random Vectors

Maximal inequalities refer to bounds on expected values of the supremum of averages of random variables over a collection. They play a crucial role in the study of non-parametric and high-dimensional estimators, and especially in the study of empirical risk minimizers. Although the expected supremum over an infinite collection appears more often in these applications, the expected supremum over a finite collection is a basic building block. This follows from the generic chaining argument. For the case of finite maximum, most existing bounds stem from the Bonferroni inequality (or the union bound). The optimality of such bounds is not obvious, especially in the context of heavy-tailed random vectors. In this article, we consider the problem of finding sharp upper and lower bounds for the expected $L_{\infty}$ norm of the mean of finite-dimensional random vectors under marginal variance bounds and an integrable envelope condition.

math.PR

Dirichlet Process-based Robust Clustering using the Median-of-Means Estimator

Clustering stands as one of the most prominent challenges in unsupervised machine learning. Among centroid-based methods, the classic $k$-means algorithm, based on Lloyd's heuristic, is widely used. Nonetheless, it is a well-known fact that $k$-means and its variants face several challenges, including heavy reliance on initial cluster centroids, susceptibility to converging into local minima of the objective function, and sensitivity to outliers and noise in the data. When data contains noise or outliers, the Median-of-Means (MoM) estimator offers a robust alternative for stabilizing centroid-based methods. On a different note, another limitation in many commonly used clustering methods is the need to specify the number of clusters beforehand. Model-based approaches, such as Bayesian nonparametric models, address this issue by incorporating infinite mixture models, which eliminate the requirement for predefined cluster counts. Motivated by these facts, in this article, we propose an efficient and automatic clustering technique by integrating the strengths of model-based and centroid-based methodologies. Our method mitigates the effect of noise on the quality of clustering; while at the same time, estimates the number of clusters. Statistical guarantees on an upper bound of clustering error, and rigorous assessment through simulated and real datasets, suggest the advantages of our proposed method over existing state-of-the-art clustering algorithms.

stat.ML