arXiv ScienceSearch

arXiv subjects

Lu Lin

Publications and source records attributed to Lu Lin.

At least 73 records · Page 4Linked to original sources

Unified Rules of Renewable Weighted Sums for Various Online Updating Estimations

This paper establishes unified frameworks of renewable weighted sums (RWS) for various online updating estimations in the models with streaming data sets. The newly defined RWS lays the foundation of online updating likelihood, online updating loss function, online updating estimating equation and so on. The idea of RWS is intuitive and heuristic, and the algorithm is computationally simple. This paper chooses nonparametric model as an exemplary setting. The RWS applies to various types of nonparametric estimators, which include but are not limited to nonparametric likelihood, quasi-likelihood and least squares. Furthermore, the method and the theory can be extended into the models with both parameter and nonparametric function. The estimation consistency and asymptotic normality of the proposed renewable estimator are established, and the oracle property is obtained. Moreover, these properties are always satisfied, without any constraint on the number of data batches, which means that the new method is adaptive to the situation where streaming data sets arrive perpetually. The behavior of the method is further illustrated by various numerical examples from simulation experiments and real data analysis.

stat.ME

Optimally estimating the sample standard deviation from the five-number summary

When reporting the results of clinical studies, some researchers may choose the five-number summary (including the sample median, the first and third quartiles, and the minimum and maximum values) rather than the sample mean and standard deviation, particularly for skewed data. For these studies, when included in a meta-analysis, it is often desired to convert the five-number summary back to the sample mean and standard deviation. For this purpose, several methods have been proposed in the recent literature and they are increasingly used nowadays. In this paper, we propose to further advance the literature by developing a smoothly weighted estimator for the sample standard deviation that fully utilizes the sample size information. For ease of implementation, we also derive an approximation formula for the optimal weight, as well as a shortcut formula for the sample standard deviation. Numerical results show that our new estimator provides a more accurate estimate for normal data and also performs favorably for non-normal data. Together with the optimal sample mean estimator in Luo et al., our new methods have dramatically improved the existing methods for data transformation, and they are capable to serve as "rules of thumb" in meta-analysis for studies reported with the five-number summary. Finally for practical use, an Excel spreadsheet and an online calculator are also provided for implementing our optimal estimators.

stat.ME

JNET: Learning User Representations via Joint Network Embedding and Topic Embedding

User representation learning is vital to capture diverse user preferences, while it is also challenging as user intents are latent and scattered among complex and different modalities of user-generated data, thus, not directly measurable. Inspired by the concept of user schema in social psychology, we take a new perspective to perform user representation learning by constructing a shared latent space to capture the dependency among different modalities of user-generated data. Both users and topics are embedded to the same space to encode users' social connections and text content, to facilitate joint modeling of different modalities, via a probabilistic generative framework. We evaluated the proposed solution on large collections of Yelp reviews and StackOverflow discussion posts, with their associated network structures. The proposed model outperformed several state-of-the-art topic modeling based user models with better predictive power in unseen documents, and state-of-the-art network embedding based user models with improved link prediction quality in unseen nodes. The learnt user representations are also proved to be useful in content recommendation, e.g., expert finding in StackOverflow.

cs.SI

A race-DC in Big Data

The strategy of divide-and-combine (DC) has been widely used in the area of big data. Bias-correction is crucial in the DC procedure for validly aggregating the locally biased estimators, especial for the case when the number of batches of data is large. This paper establishes a race-DC through a residual-adjustment composition estimate (race). The race-DC applies to various types of biased estimators, which include but are not limited to Lasso estimator, Ridge estimator and principal component estimator in linear regression, and least squares estimator in nonlinear regression. The resulting global estimator is strictly unbiased under linear model, and is acceleratingly bias-reduced in nonlinear model, and can achieve the theoretical optimality, for the case when the number of batches of data is large. Moreover, the race-DC is computationally simple because it is a least squares estimator in a pro forma linear regression. Detailed simulation studies demonstrate that the resulting global estimator is significantly bias-corrected, and the behavior is comparable with the oracle estimation and is much better than the competitors.

stat.ME

A Global Bias-Correction DC Method for Biased Estimation under Memory Constraint

This paper establishes a global bias-correction divide-and-conquer (GBC-DC) rule for biased estimation under the case of memory constraint. In order to introduce the new estimation, a closed representation of the local estimators obtained by the data in each batch is adopted, aiming to formulate a pro forma linear regression between the local estimators and the true parameter of interest. Least square method is then used within this framework to composite a global estimator of the parameter. Thus, the main advantage over the classical DC method is that the new GBC-DC method can absorb the information hidden in the statistical structure and the variables in each batch of data. Consequently, the resulting global estimator is strictly unbiased even if the local estimator has a non-negligible bias. Moreover, the global estimator is consistent, and even can achieve root-$n$ consistency, without the constraint on the number of batches. Another attractive feature of the new method is computationally simple and efficient, without use of any iterative algorithm and local bias-correction. Specifically, the proposed GBC-DC method applies to various biased estimations such as shrinkage-type estimation and nonparametric regression estimation. Detailed simulation studies demonstrate that the proposed GBC-DC approach is significantly bias-corrected, and the behavior is comparable with the full data estimation and is much better than the competitors.

stat.ME

How to estimate the sample mean and standard deviation from the five number summary?

In some clinical studies, researchers may report the five number summary (including the sample median, the first and third quartiles, and the minimum and maximum values) rather than the sample mean and standard deviation. To conduct meta-analysis for pooling studies, one needs to first estimate the sample mean and standard deviation from the five number summary. A number of studies have been proposed in the recent literature to solve this problem. However, none of the existing estimators for the standard deviation is satisfactory for practical use. After a brief review of the existing literature, we point out that Wan et al.'s method (BMC Med Res Methodol 14:135, 2014) has a serious limitation in estimating the standard deviation from the five number summary. To improve it, we propose a smoothly weighted estimator by incorporating the sample size information and derive the optimal weight for the new estimator. For ease of implementation, we also provide an approximation formula of the optimal weight and a shortcut formula for estimating the standard deviation from the five number summary. The performance of the proposed estimator is evaluated through two simulation studies. In comparison with Wan et al.'s estimator, our new estimator provides a more accurate estimate for normal data and performs favorably for non-normal data. In real data analysis, our new method is also able to provide a more accurate estimate of the true sample standard deviation than the existing method. In this paper, we propose an optimal estimator of the standard deviation from the five number summary. Together with the optimal mean estimator in Luo et al. (Stat Methods Med Res, in press, 2017), our new methods have improved the existing literature and will make a solid contribution to meta-analysis and evidence-based medicine.

stat.ME

A simple and efficient profile likelihood for semiparametric exponential family

Semiparametric exponential family proposed by Ning et al. (2017) is an extension of the parametric exponential family to the case with a nonparametric base measure function. Such a distribution family has potential application in some areas such as high dimensional data analysis. However, the methodology for achieving the semiparametric efficiency has not been proposed in the existing literature. In this paper, we propose a profile likelihood to efficiently estimate both parameter and nonparametric function. Due to the use of the least favorable curve in the procedure of profile likelihood, the semiparametric efficiency is achieved successfully and the estimation bias is reduced significantly. Moreover, by making the most of the structure information of the semiparametric exponential family, the estimator of the least favorable curve has an explicit expression. It ensures that the newly proposed profile likelihood can be implemented and is computationally simple. Simulation studies can illustrate that our proposal is much better than the existing methodology for most cases under study, and is robust to the different model conditions.

stat.ME

Group-Average and Convex Clustering for Partially Heterogeneous Linear Regression

In this paper, a subgroup least squares and a convex clustering are introduced for inferring a partially heterogenous linear regression that has potential application in the areas of precision marketing and precision medicine. The homogenous parameter and the subgroup-average of the heterogenous parameters can be consistently estimated by the subgroup least squares, without need of the sparsity assumption on the heterogenous parameters. The heterogenous parameters can be consistently clustered via the convex clustering. Unlike the existing methods for regression clustering, our clustering procedure is a standard mean clustering, although the model under study is a type of regression, and the corresponding algorithm only involves low dimensional parameters. Thus, it is simple and stable even if the sample size is large. The advantage of the method is further illustrated via simulation studies and the analysis of car sales data.

stat.ME

Upper expectation parametric regression

Every observation may follow a distribution that is randomly selected in a class of distributions. It is called the distribution uncertainty. This is a fact acknowledged in some research fields such as financial risk measure. Thus, the classical expectation is not identifiable in general.In this paper, a distribution uncertainty is defined, and then an upper expectation regression is proposed, which can describe the relationship between extreme events and relevant covariates under the framework of distribution uncertainty. As there are no classical methods available to estimate the parameters in the upper expectation regression, a two-step penalized maximum least squares procedure is proposed to estimate the mean function and the upper expectation of the error. The resulting estimators are consistent and asymptotically normal in a certain sense.Simulation studies and a real data example are conducted to show that the classical least squares estimation does not work and the penalized maximum least squares performs well.

stat.ME

Inference for biased models: a quasi-instrumental variable approach

For linear regression models who are not exactly sparse in the sense that the coefficients of the insignificant variables are not exactly zero, the working models obtained by a variable selection are often biased. Even in sparse cases, after a variable selection, when some significant variables are missing, the working models are biased as well. Thus, under such situations, root-n consistent estimation and accurate prediction could not be expected. In this paper, a novel remodelling method is proposed to produce an unbiased model when quasi-instrumental variables are introduced. The root-n estimation consistency and the asymptotic normality can be achieved, and the prediction accuracy can be promoted as well. The performance of the new method is examined through simulation studies.

stat.ME

Asymptotic Composite Estimation

Composition methodologies in the current literature are mainly to promote estimation efficiency via direct composition, either, of initial estimators or of objective functions. In this paper, composite estimation is investigated for both estimation efficiency and bias reduction. To this end, a novel method is proposed by utilizing a regression relationship between initial estimators and values of model-independent parameter in an asymptotic sense. The resulting estimators could have smaller limiting variances than those of initial estimators, and for nonparametric regression estimation, could also have faster convergence rate than the classical optimal rate that the corresponding initial estimators can achieve. The simulations are carried out to examine its performance in finite sample situations.

stat.ME

Sublinear expectation linear regression

Nonlinear expectation, including sublinear expectation as its special case, is a new and original framework of probability theory and has potential applications in some scientific fields, especially in finance risk measure and management. Under the nonlinear expectation framework, however, the related statistical models and statistical inferences have not yet been well established. The goal of this paper is to construct the sublinear expectation regression and investigate its statistical inference. First, a sublinear expectation linear regression is defined and its identifiability is given. Then, based on the representation theorem of sublinear expectation and the newly defined model, several parameter estimations and model predictions are suggested, the asymptotic normality of estimations and the mini-max property of predictions are obtained. Furthermore, new methods are developed to realize variable selection for high-dimensional model. Finally, simulation studies and a real-life example are carried out to illustrate the new models and methodologies. All notions and methodologies developed are essentially different from classical ones and can be thought of as a foundation for general nonlinear expectation statistics.

math.ST

Estimation and inference for high-dimensional non-sparse models

To successfully work on variable selection, sparse model structure has become a basic assumption for all existing methods. However, this assumption is questionable as it is hard to hold in most of cases and none of existing methods may provide consistent estimation and accurate model prediction in nons-parse scenarios. In this paper, we propose semiparametric re-modeling and inference when the linear regression model under study is possibly non-sparse. After an initial working model is selected by a method such as the Dantzig selector adopted in this paper, we re-construct a globally unbiased semiparametric model by use of suitable instrumental variables and nonparametric adjustment. The newly defined model is identifiable, and the estimator of parameter vector is asymptotically normal. The consistency, together with the re-built model, promotes model prediction. This method naturally works when the model is indeed sparse and thus is of robustness against non-sparseness in certain sense. Simulation studies show that the new approach has, particularly when $p$ is much larger than $n$, significant improvement of estimation and prediction accuracies over the Gaussian Dantzig selector and other classical methods. Even when the model under study is sparse, our method is also comparable to the existing methods designed for sparse models.

stat.ME

Adaptive post-Dantzig estimation and prediction for non-sparse "large $p$ and small $n$" models

For consistency (even oracle properties) of estimation and model prediction, almost all existing methods of variable/feature selection critically depend on sparsity of models. However, for ``large $p$ and small $n$" models sparsity assumption is hard to check and particularly, when this assumption is violated, the consistency of all existing estimations is usually impossible because working models selected by existing methods such as the LASSO and the Dantzig selector are usually biased. To attack this problem, we in this paper propose adaptive post-Dantzig estimation and model prediction. Here the adaptability means that the consistency based on the newly proposed method is adaptive to non-sparsity of model, choice of shrinkage tuning parameter and dimension of predictor vector. The idea is that after a sub-model as a working model is determined by the Dantzig selector, we construct a globally unbiased sub-model by choosing suitable instrumental variables and nonparametric adjustment. The new estimation of the parameters in the sub-model can be of the asymptotic normality. The consistent estimator, together with the selected sub-model and adjusted model, improves model predictions. Simulation studies show that the new approach has the significant improvement of estimation and prediction accuracies over the Gaussian Dantzig selector and other classical methods have.

stat.ME

Covariate-adjusted nonlinear regression

In this paper, we propose a covariate-adjusted nonlinear regression model. In this model, both the response and predictors can only be observed after being distorted by some multiplicative factors. Because of nonlinearity, existing methods for the linear setting cannot be directly employed. To attack this problem, we propose estimating the distorting functions by nonparametrically regressing the predictors and response on the distorting covariate; then, nonlinear least squares estimators for the parameters are obtained using the estimated response and predictors. Root $n$-consistency and asymptotic normality are established. However, the limiting variance has a very complex structure with several unknown components, and confidence regions based on normal approximation are not efficient. Empirical likelihood-based confidence regions are proposed, and their accuracy is also verified due to its self-scale invariance. Furthermore, unlike the common results derived from the profile methods, even when plug-in estimates are used for the infinite-dimensional nuisance parameters (distorting functions), the limit of empirical likelihood ratio is still chi-squared distributed. This property eases the construction of the empirical likelihood-based confidence regions. A simulation study is carried out to assess the finite sample performance of the proposed estimators and confidence regions. We apply our method to study the relationship between glomerular filtration rate and serum creatinine.

math.ST

Extension of Dirac theory and the classification of elementary particles

The Dirac theory implies the existence of an internal vector space, in addition to spin space. Using Dirac's coupling of variables in internal space to those in physical space, we construct a new configuration structure for particles in the combined physical plus internal spaces. The importance of this is that the internal degrees of freedom implicit in Dirac's theory allow a new classification of elementary particles. An important consequence is the prediction of a new type of quark. As expected, our theory groups fermions into doublets, which are then divided into color singlets (leptons) and color triplets (quarks), and which are then further divided into generation singlets and generation triplets. If the Pauli exclusion principle for fermions is also valid within a particle's internal space, then our theory makes two important predictions. First, we can explain why the widely studied quarks (up, down, charm, strange, top, bottom) cannot be observed as free states in nature. Second, we predict the existence of a new quark which can indeed be observed as a free state in nature, and whose wave function is antisymmetric in internal space. W. M. Fairbank, a Guggenheim Fellow, has published experimental data which supports our second prediction.

physics.gen-ph

Complex space-time and the classification of elementary particles

It is shown that the Dirac theory implies complex space-time and complex space-time can lead to the Dirac equation. It is suggested that fermions are grouped into doublets, those doublets are then divided into color singlets (leptons) and color triplets (quarks), then they are further divided into generation singlets and generation triplets. If the exclusion principle for fermions works in the internal space, then the non-observation of free quarks can be explained. However, it is suggested that a possible new free quark may exist with a color triplet and generation singlet state which is antisymmetric in the internal space.

physics.gen-ph