arXiv ScienceSearch

arXiv subjects

Kazuharu Harada

Publications and source records attributed to Kazuharu Harada.

7 recordsLinked to original sources

Efficient estimation of weighted treatment effects under two-phase sampling

Two-phase sampling offers a practical way to collect costly confounders only in a subsample while retaining inexpensive information for a larger cohort. In observational causal studies, however, phase-2 selection can distort estimation of population causal effects if the sampling mechanism is ignored, and available phase-1 information may also be exploited to improve efficiency. Yet efficiency theory for causal estimands under such designs remains limited, particularly beyond the average treatment effect. In this paper, we derive the semiparametric efficiency bound for a class of propensity-score-weighted average treatment effects, which includes the average treatment effect, effects among treated and untreated populations, and the overlap effect, under two-phase sampling. In addition to straightforward weighting estimators based on the known sampling probabilities, we propose an enriched doubly robust estimator that attains the efficiency bound when all nuisance functions are consistently estimated. In particular, under outcome-dependent sampling, substantial efficiency gains can arise in some settings by appropriately incorporating phase-1 information. We further conduct extensive simulation studies, varying the choice of phase-1 variables and sampling schemes, to characterize when and to what extent leveraging phase-1 information leads to efficiency gains.

stat.ME

Prognosis-equivalent mapping of clinical measurements via survival analysis

A continuous clinical measurement recorded in fixed physical units may have different prognostic meaning across patients when its effect depends on a patient-level modifier. For example, the same tumor diameter may imply markedly different prognosis in an infant and an adult, because patient size can modify its prognostic effect. We formalize this problem through prognosis-equivalent mapping for time-to-event outcomes. The resulting estimand maps a measurement value observed at one modifier level to the value under a reference modifier level that yields the same conditional prognostic quantity. The basic operation is to equate a monotone conditional prognostic score and invert the reference-side curve. In observational data, however, treatment may be selected according to the modifier and the measurement, so a direct equate-and-invert procedure may reflect differences in treatment assignment as well as the prognostic meaning of the measurement itself. We therefore define, in addition to a direct approach, a policy-standardized mapping that standardizes treatment under a common policy using the g-formula. We also introduce an origin-referenced mapping that compares changes in prognostic score from a common anchor, thereby separating modifier-specific prognostic levels from modifier dependence in the measurement--prognosis relationship. Using Cox models, we develop linear and flexible inversion estimators with analytic and bootstrap confidence bands. Simulations evaluate finite-sample performance and illustrate how treatment standardization and origin referencing clarify what is being mapped.

stat.ME

Simultaneous Modeling of Disease Screening and Severity Prediction: A Multi-task and Sparse Regularization Approach

Identifying clinically relevant biomarkers and developing predictive models are central challenges in biomedical research. Biomarkers are commonly used for disease screening, and some provide information not only on the presence or absence of a disease but also on its severity. Such biomarkers can contribute to treatment prioritization and support clinical decision-making. To address both disease screening and severity prediction, this paper focuses on regression modeling for ordinal outcomes with a hierarchical structure. When the response variable is a combination of the presence of disease and severity, such as {healthy, mild, intermediate, severe}, a straightforward approach is to apply the conventional ordinal regression model. However, such models may lack the flexibility needed to capture heterogeneity in how predictors relate to response levels, particularly when the response levels have a heterogeneous association structure with predictors. Therefore, this paper proposes a model that treats screening and severity prediction as separate tasks, along with an estimation method based on structural sparse regularization. This method is designed to leverage a shared structure between the tasks. In numerical experiments, the proposed method demonstrated stable performance across many scenarios compared to existing ordinal regression methods.

stat.ME

False Discovery Rate Control for Confounder Selection Using Mirror Statistics

While data-driven confounder selection requires careful consideration, it is frequently employed in observational studies. Widely recognized criteria for confounder selection include the minimal-set approach, which involves selecting variables relevant to both treatment and outcome, and the union-set approach, which involves selecting variables associated with either treatment or outcome. These approaches are often implemented using heuristics and off-the-shelf statistical methods, where the degree of uncertainty may not be clear. In this paper, we focus on the false discovery rate (FDR) to measure uncertainty in confounder selection. We define the FDR specific to confounder selection and propose methods based on the mirror statistic, a recently developed approach for FDR control that does not rely on p-values. The proposed methods are p-value-free and require only the assumption of some symmetry in the distribution of the mirror statistic. It can be combined with sparse estimation and other methods that involve difficulties in deriving p-values. The properties of the proposed methods are investigated through exhaustive numerical experiments. Particularly in high-dimensional data scenarios, the proposed methods effectively control FDR and perform better than the p-value-based methods.

stat.ME

An interpretable neural network-based non-proportional odds model for ordinal regression

This study proposes an interpretable neural network-based non-proportional odds model (N$^3$POM) for ordinal regression. N$^3$POM is different from conventional approaches to ordinal regression with non-proportional models in several ways: (1) N$^3$POM is defined for both continuous and discrete responses, whereas standard methods typically treat the ordered continuous variables as if they are discrete, (2) instead of estimating response-dependent finite-dimensional coefficients of linear models from discrete responses as is done in conventional approaches, we train a non-linear neural network to serve as a coefficient function. Thanks to the neural network, N$^3$POM offers flexibility while preserving the interpretability of conventional ordinal regression. We establish a sufficient condition under which the predicted conditional cumulative probability locally satisfies the monotonicity constraint over a user-specified region in the covariate space. Additionally, we provide a monotonicity-preserving stochastic (MPS) algorithm for effectively training the neural network. We apply N$^3$POM to several real-world datasets.

stat.ME

Outlier-Resistant Estimators for Average Treatment Effect in Causal Inference

The inverse probability (IPW) and doubly robust (DR) estimators are often used to estimate the average causal effect (ATE), but are vulnerable to outliers. The IPW/DR median can be used for outlier-resistant estimation of the ATE, but the outlier resistance of the median is limited and it is not resistant enough for heavy contamination. We propose extensions of the IPW/DR estimators with density power weighting, which can eliminate the influence of outliers almost completely. The outlier resistance of the proposed estimators is evaluated through the unbiasedness of the estimating equations. Unlike the median-based methods, our estimators are resistant to outliers even under heavy contamination. Interestingly, the naive extension of the DR estimator requires bias correction to keep the double robustness even under the most tractable form of contamination. In addition, the proposed estimators are found to be highly resistant to outliers in more difficult settings where the contamination ratio depends on the covariates. The outlier resistance of our estimators from the viewpoint of the influence function is also favorable. Our theoretical results are verified via Monte Carlo simulations and real data analysis. The proposed methods were found to have more outlier resistance than the median-based methods and estimated the potential mean with a smaller error than the median-based methods.

stat.ME

Estimation of Structural Causal Model via Sparsely Mixing Independent Component Analysis

We consider the problem of inferring the causal structure from observational data, especially when the structure is sparse. This type of problem is usually formulated as an inference of a directed acyclic graph (DAG) model. The linear non-Gaussian acyclic model (LiNGAM) is one of the most successful DAG models, and various estimation methods have been developed. However, existing methods are not efficient for some reasons: (i) the sparse structure is not always incorporated in causal order estimation, and (ii) the whole information of the data is not used in parameter estimation. To address {these issues}, we propose a new estimation method for a linear DAG model with non-Gaussian noises. The proposed method is based on the log-likelihood of independent component analysis (ICA) with two penalty terms related to the sparsity and the consistency condition. The proposed method enables us to estimate the causal order and the parameters simultaneously. For stable and efficient optimization, we propose some devices, such as a modified natural gradient. Numerical experiments show that the proposed method outperforms existing methods, including LiNGAM and NOTEARS.

stat.ML