arXiv ScienceSearch

arXiv subjects

Samuel Orso

Publications and source records attributed to Samuel Orso.

10 recordsLinked to original sources

An accurate percentile method for parametric inference based on asymptotically biased estimators

Inference methods for computing confidence intervals in parametric settings usually rely on consistent estimators of the parameter of interest. However, it may be computationally and/or analytically burdensome to obtain such estimators in various parametric settings, for example when the data exhibit certain features such as censoring, misclassification errors or outliers. To address these challenges, we propose a simulation-based inferential method, called the implicit bootstrap, that remains valid regardless of the potential asymptotic bias of the estimator on which the method is based. We demonstrate that this method allows for the construction of asymptotically valid percentile confidence intervals of the parameter of interest. Additionally, we show that these confidence intervals can also achieve second-order accuracy. We also show that the method is exact in three instances where the standard bootstrap fails. Using simulation studies, we illustrate the coverage accuracy of the method in three examples where standard parametric bootstrap procedures are computationally intensive and less accurate in finite samples.

stat.ME

Just Identified Indirect Inference Estimator: Accurate Inference through Bias Correction

An important challenge in statistical analysis lies in controlling the estimation bias when handling the ever-increasing data size and model complexity of modern data settings. In this paper, we propose a reliable estimation and inference approach for parametric models based on the Just Identified iNdirect Inference estimator (JINI). The key advantage of our approach is that it allows to construct a consistent estimator in a simple manner, while providing strong bias correction guarantees that lead to accurate inference. Our approach is particularly useful for complex parametric models, as it allows to bypass the analytical and computational difficulties (e.g., due to intractable estimating equation) typically encountered in standard procedures. The properties of JINI (including consistency, asymptotic normality, and its bias correction property) are also studied when the parameter dimension is allowed to diverge, which provide the theoretical foundation to explain the advantageous performance of JINI in increasing dimensional covariates settings. Our simulations and an alcohol consumption data analysis highlight the practical usefulness and excellent performance of JINI when data present features (e.g., misclassification, rounding) as well as in robust estimation.

stat.ME

A General Approach for Simulation-based Bias Correction in High Dimensional Settings

An important challenge in statistical analysis lies in controlling the bias of estimators due to the ever-increasing data size and model complexity. Approximate numerical methods and data features like censoring and misclassification often result in analytical and/or computational challenges when implementing standard estimators. As a consequence, consistent estimators may be difficult to obtain, especially in complex and/or high dimensional settings. In this paper, we study the properties of a general simulation-based estimation framework that allows to construct bias corrected consistent estimators. We show that the considered approach leads, under more general conditions, to stronger bias correction properties compared to alternative methods. Besides its bias correction advantages, the considered method can be used as a simple strategy to construct consistent estimators in settings where alternative methods may be challenging to apply. Moreover, the considered framework can be easily implemented and is computationally efficient. These theoretical results are highlighted with simulation studies of various commonly used models, including the negative binomial regression (with and without censoring) and the logistic regression (with and without misclassification errors). Additional numerical illustrations are provided in the supplementary materials.

math.ST

SWAG: A Wrapper Method for Sparse Learning

The majority of machine learning methods and algorithms give high priority to prediction performance which may not always correspond to the priority of the users. In many cases, practitioners and researchers in different fields, going from engineering to genetics, require interpretability and replicability of the results especially in settings where, for example, not all attributes may be available to them. As a consequence, there is the need to make the outputs of machine learning algorithms more interpretable and to deliver a library of "equivalent" learners (in terms of prediction performance) that users can select based on attribute availability in order to test and/or make use of these learners for predictive/diagnostic purposes. To address these needs, we propose to study a procedure that combines screening and wrapper approaches which, based on a user-specified learning method, greedily explores the attribute space to find a library of sparse learners with consequent low data collection and storage costs. This new method (i) delivers a low-dimensional network of attributes that can be easily interpreted and (ii) increases the potential replicability of results based on the diversity of attribute combinations defining strong learners with equivalent predictive power. We call this algorithm "Sparse Wrapper AlGorithm" (SWAG).

stat.ML

Asymptotically Optimal Bias Reduction for Parametric Models

An important challenge in statistical analysis concerns the control of the finite sample bias of estimators. This problem is magnified in high-dimensional settings where the number of variables $p$ diverges with the sample size $n$, as well as for nonlinear models and/or models with discrete data. For these complex settings, we propose to use a general simulation-based approach and show that the resulting estimator has a bias of order $\mathcal{O}(0)$, hence providing an asymptotically optimal bias reduction. It is based on an initial estimator that can be slightly asymptotically biased, making the approach very generally applicable. This is particularly relevant when classical estimators, such as the maximum likelihood estimator, can only be (numerically) approximated. We show that the iterative bootstrap of Kuk (1995) provides a computationally efficient approach to compute this bias reduced estimator. We illustrate our theoretical results in simulation studies for which we develop new bias reduced estimators for the logistic regression, with and without random effects. These estimators enjoy additional properties such as robustness to data contamination and to the problem of separability.

math.ST

Wavelet-Based Moment-Matching Techniques for Inertial Sensor Calibration

The task of inertial sensor calibration has required the development of various techniques to take into account the sources of measurement error coming from such devices. The calibration of the stochastic errors of these sensors has been the focus of increasing amount of research in which the method of reference has been the so-called "Allan variance slope method" which, in addition to not having appropriate statistical properties, requires a subjective input which makes it prone to mistakes. To overcome this, recent research has started proposing "automatic" approaches where the parameters of the probabilistic models underlying the error signals are estimated by matching functions of the Allan variance or Wavelet Variance with their model-implied counterparts. However, given the increased use of such techniques, there has been no study or clear direction for practitioners on which approach is optimal for the purpose of sensor calibration. This paper formally defines the class of estimators based on this technique and puts forward theoretical and applied results that, comparing with estimators in this class, suggest the use of the Generalized Method of Wavelet Moments as an optimal choice.

stat.ME

Phase Transition Unbiased Estimation in High Dimensional Settings

An important challenge in statistical analysis concerns the control of the finite sample bias of estimators. For example, the maximum likelihood estimator has a bias that can result in a significant inferential loss. This problem is typically magnified in high-dimensional settings where the number of variables $p$ is allowed to diverge with the sample size $n$. However, it is generally difficult to establish whether an estimator is unbiased and therefore its asymptotic order is a common approach used (in low-dimensional settings) to quantify the magnitude of the bias. As an alternative, we introduce a new and stronger property, possibly for high-dimensional settings, called phase transition unbiasedness. An estimator satisfying this property is unbiased for all $n$ greater than a finite sample size $n^\ast$. Moreover, we propose a phase transition unbiased estimator built upon the idea of matching an initial estimator computed on the sample and on simulated data. It is not required for this initial estimator to be consistent and thus it can be chosen for its computational efficiency and/or for other desirable properties such as robustness. This estimator can be computed using a suitable simulation based algorithm, namely the iterative bootstrap, which is shown to converge exponentially fast. In addition, we demonstrate the consistency and the limiting distribution of this estimator in high-dimensional settings. Finally, as an illustration, we use our approach to develop new estimators for the logistic regression model, with and without random effects, that also enjoy other properties such as robustness to data contamination and are also not affected by the problem of separability. In a simulation exercise, the theoretical results are confirmed in settings where the sample size is relatively small compared to the model dimension.

math.ST

A simple recipe for making accurate parametric inference in finite sample

Constructing tests or confidence regions that control over the error rates in the long-run is probably one of the most important problem in statistics. Yet, the theoretical justification for most methods in statistics is asymptotic. The bootstrap for example, despite its simplicity and its widespread usage, is an asymptotic method. There are in general no claim about the exactness of inferential procedures in finite sample. In this paper, we propose an alternative to the parametric bootstrap. We setup general conditions to demonstrate theoretically that accurate inference can be claimed in finite sample.

stat.ME

On the Properties of Simulation-based Estimators in High Dimensions

Considering the increasing size of available data, the need for statistical methods that control the finite sample bias is growing. This is mainly due to the frequent settings where the number of variables is large and allowed to increase with the sample size bringing standard inferential procedures to incur significant loss in terms of performance. Moreover, the complexity of statistical models is also increasing thereby entailing important computational challenges in constructing new estimators or in implementing classical ones. A trade-off between numerical complexity and statistical properties is often accepted. However, numerically efficient estimators that are altogether unbiased, consistent and asymptotically normal in high dimensional problems would generally be ideal. In this paper, we set a general framework from which such estimators can easily be derived for wide classes of models. This framework is based on the concepts that underlie simulation-based estimation methods such as indirect inference. The approach allows various extensions compared to previous results as it is adapted to possibly inconsistent estimators and is applicable to discrete models and/or models with a large number of parameters. We consider an algorithm, namely the Iterative Bootstrap (IB), to efficiently compute simulation-based estimators by showing its convergence properties. Within this framework we also prove the properties of simulation-based estimators, more specifically the unbiasedness, consistency and asymptotic normality when the number of parameters is allowed to increase with the sample size. Therefore, an important implication of the proposed approach is that it allows to obtain unbiased estimators in finite samples. Finally, we study this approach when applied to three common models, namely logistic regression, negative binomial regression and lasso regression.

math.ST

A Paradigmatic Regression Algorithm for Gene Selection Problems

Motivation: Gene selection has become a common task in most gene expression studies. The objective of such research is often to identify the smallest possible set of genes that can still achieve good predictive performance. The problem of assigning tumours to a known class is a particularly important example that has received considerable attention in the last ten years. Many of the classification methods proposed recently require some form of dimension-reduction of the problem. These methods provide a single model as an output and, in most cases, rely on the likelihood function in order to achieve variable selection. Results: We propose a prediction-based objective function that can be tailored to the requirements of practitioners and can be used to assess and interpret a given problem. The direct optimization of such a function can be very difficult because the problem is potentially discontinuous and nonconvex. We therefore propose a general procedure for variable selection that resembles importance sampling to explore the feature space. Our proposal compares favorably with competing alternatives when applied to two cancer data sets in that smaller models are obtained for better or at least comparable classification errors. Furthermore by providing a set of selected models instead of a single one, we construct a network of possible models for a target prediction accuracy level.

stat.ME