arXiv ScienceSearch

arXiv subjects

Shonosuke Sugasawa

Publications and source records attributed to Shonosuke Sugasawa.

At least 19 recordsLinked to original sources

Spatially Dependent Indian Buffet Processes

We develop a new stochastic process called spatially dependent Indian buffet processes (sIBP) for binary feature matrices of unbounded columns with spatial correlations between subjects, and propose general spatial factor models for various multivariate response variables. We introduce spatial dependency through the stick-breaking representation of the original Indian buffet process (IBP; Griffiths and Ghahramani, 2005, 2011) and latent Gaussian process for the logit-transformed breaking proportions to capture underlying spatial correlation. We show that sIBP retains the sparsity and finite-feature behavior of the original IBP, while its joint feature allocation probabilities are affected by spatial correlation. Using binomial expansion and Polya-gamma data augmentation, we provide an efficient Gibbs sampler for posterior computation. The usefulness of our sIBP is demonstrated through simulation studies and two applications for large-dimensional multinomial data of areal dialects and geographical distribution of multiple tree species.

stat.ME

Beyond Tweedie's Formula: Conditional Score Modeling for Empirical Bayes Inference

We propose conditional f-modeling (Cf-modeling), a framework for empirical Bayes inference with covariates. A central identity shows that the conditional marginal score function determines not only the posterior mean through Tweedie's formula, but also the posterior moment-generating function, providing a basis for recovering posterior quantities without explicit prior modeling. Motivated by this observation, we treat the conditional marginal score as the primary object of inference and estimate it directly using an energy-based representation and Hyvärinen score matching, thereby avoiding potentially intractable covariate-dependent normalizing constants. The resulting framework flexibly accommodates covariate effects and heteroscedasticity and provides a practical approach to posterior moment estimation and uncertainty quantification. We demonstrate the effectiveness of the proposed method through simulations and an RNA-seq application.

stat.ME

Contrastive Bayesian Inference for Unnormalized Models

Unnormalized (or energy-based) models provide a flexible framework for capturing the characteristics of data with complex dependency structures. However, the application of standard Bayesian inference methods has been severely limited because the parameter-dependent normalizing constant is either analytically intractable or computationally prohibitive to evaluate. A promising approach is score-based generalized Bayesian inference, which avoids evaluating the normalizing constant by replacing the likelihood with a scoring rule. However, this approach still requires careful tuning of the likelihood information, and it may fail to yield valid inference without appropriate control. To overcome this difficulty, we propose a fully Bayesian framework for inference on unnormalized models that does not require such tuning. We build on noise-contrastive estimation, which recasts inference as a binary classification problem between observed and noise samples, and treat the normalizing constant as an additional unknown parameter within the resulting likelihood. For exponential families, the classification likelihood becomes conditionally Gaussian via Pólya-Gamma data augmentation, leading to a simple Gibbs sampler. We further establish posterior concentration and a Bernstein-von Mises theorem for our proposed method, providing theoretical justification for its uncertainty quantification. We demonstrate the proposed approach through two models: time-varying density models of temporal point processes and sparse torus graph models of multivariate circular data. Through simulation studies and real-data analyses, our proposed method provides accurate point estimates and principled uncertainty quantification.

stat.ME

Finite Mixtures of Generalized Estimating Equations for Clustering Multivariate Correlated Outcomes

Multivariate correlated outcomes occur across disciplines, including ecology, social sciences, and psychometrics. This paper focuses on clustering these outcomes across observational units, specifically, finding groups of units with the same ``outcome profile". Our motivation comes from bioregionalization in ecology, which aims to cluster sites into bioregions with the same species profiles, where site membership can depend on environmental or habitat covariates. To accomplish this, we propose finite mixtures of generalized estimating equations (MixGEE). Unlike existing approaches to model-based bioregionalization, MixGEE partitions sites into regions while accounting for between-species correlations through a region-specific working correlation structure. Thus, each region is characterized by a marginal species mean vector and a between-species correlation matrix. Unlike likelihood-based finite mixture models, MixGEE does not require a full joint distribution of the multivariate outcomes. Instead, we construct a pseudo-posterior probability for region membership motivated by the large-sample distribution of the estimating equation. This leads to an iterative algorithm alternating between updating these probabilities and solving weighted estimating equations. We determine the number of regions using cross-validation based on predictive performance for held-out sites and species components, and use a clustered Dirichlet random-weight bootstrap for uncertainty quantification. Simulations demonstrate reliable estimation and inference under various correlation structures and more stable selection of the number of groups than methods that ignore dependence. Applying MixGEE to presence--absence records of fish species around the Kerguelen Plateau reveals three distinct fish assemblage profiles with heterogeneous occurrence and within-site correlation patterns.

stat.ME

Log-regularly varying scale mixture of asymmetric Laplaces for robust Bayesian quantile regression

Bayesian quantile regression based on the asymmetric Laplace (AL) distribution can be sensitive to extreme observations because of its exponentially decaying tails. We propose a robust error distribution constructed as a finite mixture of the AL distribution and a log-Pareto scale mixture of asymmetric Laplace distributions (LPAL). Unlike a direct log-Pareto extension of the normal location-scale representation of the AL distribution, the proposed AL-LPAL mixture preserves the prescribed quantile and exhibits log-regularly varying behavior in both tails. The LPAL component also has an unbounded density at the target quantile, yielding a distribution that combines sharp central concentration with super-heavy tails. We establish posterior robustness under arbitrarily extreme contamination and provide sufficient conditions for the existence of posterior moments of the regression coefficients and scale parameter. For posterior computation, we develop a Gibbs sampler using latent-variable augmentations and a computationally efficient mean-field variational Bayes approximation. Simulation studies show that the proposed method is competitive under moderate contamination and maintains stable point estimation with comparatively concentrated posterior intervals, particularly when severe contamination affects the quantile of interest. Applications to carbon dioxide and Boston housing data, using the same preprocessing as existing robust Bayesian quantile regression analyses, show favorable predictive performance across nearly all quantile levels and loss criteria considered.

stat.ME

Small Area Estimation under Spatial Regimes: Spatially Clustered Fay-Herriot Models for Agricultural Indicators

Area-level small area estimation (SAE) models, such as the Fay--Herriot (FH) model, borrow strength across domains through covariates and random effects, but they can struggle when the relationship between the covariates and the outcome is spatially heterogeneous, that is, when it changes across the spatial domain of interest. We propose a spatially-clustered FH (SC-FH) framework that simultaneously (i) estimates cluster-specific regression coefficients and random effects variances and (ii) generates spatially coherent partitions of the geographical domain. Estimation maximizes a penalized likelihood that augments the FH likelihood with a Potts-type spatial cohesion term over the areal adjacency graph, through an efficient strategy that alternates between sequential label updates and closed-form FH updates within clusters. In simulation experiments run on the real geography of the application, the method recovers the latent regimes almost exactly whenever they are separated in the covariate--response space and improves prediction accuracy over the standard FH benchmark, with the spatial penalty acting as a stabilizer of both classification and estimation. An empirical application to the average standard output of farms in the Po Valley (Northern Italy) identifies two spatially compact production regimes with significantly different cluster-wise coefficients, and shows that the clusterwise predictor improves on the direct estimates while avoiding the over-shrinkage of the pooled model.

stat.AP

Scalable Estimation of Crossed Random Effects Models via Multi-way Discretization

Cross-classified data frequently arise in scientific fields such as education, healthcare, and social sciences. A common modeling strategy is to introduce crossed random effects within a regression framework. However, this approach often encounters serious computational bottlenecks, particularly for non-Gaussian outcomes. In this paper, we propose a scalable and flexible method that approximates the distribution of each random effect by a discrete distribution, effectively partitioning the random effects into a finite number of representative values. This approximation allows us to express the model as a multi-way discrete structure, which can be efficiently estimated using a simple and fast iterative algorithm. The proposed method accommodates a wide range of outcome models and remains applicable even in settings with more than two-way cross-classification. We theoretically establish the consistency and asymptotic normality of the estimator under general settings of classification levels. Through simulation studies and real data applications, we demonstrate the practical performance of the proposed method in logistic, Poisson, and ordered probit regression models involving cross-classified structures.

stat.ME

Spatio-temporal smoothing, interpolation and prediction of income distributions based on grouped data

The Housing and Land Survey (HLS) of Japan provides municipality-level grouped data on household incomes. Although such data can be invaluable for effective local policymaking, their analysis is often hindered by several challenges, including limited information inherent in the grouped format, the presence of missing areas, and the low frequency of survey implementation. To address these issues, we propose a novel grouped-data-based spatio-temporal finite mixture model for estimating income distributions across multiple spatial units and time points. A unique feature of the proposed method is that all areas share common latent distributions, while the mixing proportions, incorporating spatial and temporal effects, capture the potential area-wise heterogeneity. Consequently, the inclusion of these effects enables smoothing quantities of interest over space and time, imputing missing values, and predicting future trends. By applying the proposed method to the HLS data, we generate complete maps of income and inequality measures at any given time, thereby facilitating rapid and efficient policymaking with fine granularity.

stat.AP

Semiparametric Copula Estimation for Spatially Correlated Multivariate Mixed Outcomes: Analyzing Visual Sightings of Fin Whales from a Line Transect Survey

For marine biologists, ascertaining the dependence structures between marine species and marine environments, such as sea surface temperature and ocean depth, is imperative for defining ecosystem functioning and providing insights into the dynamics of marine ecosystems. However, obtained data include not only continuous but also discrete data, such as binaries and counts (referred to as mixed outcomes), as well as spatial correlations, both of which make conventional multivariate analysis tools impractical. To solve this issue, we propose semiparametric Bayesian inference and develop an efficient algorithm for computing the posterior of the dependence structure based on the rank likelihood under a latent multivariate spatial Gaussian process using the Markov chain Monte Carlo method. To alleviate the computational intractability caused by the Gaussian process, we also provide a scalable implementation that leverages the nearest-neighbor Gaussian process. Extensive numerical experiments reveal that the proposed method reliably infers the dependence structures of spatially correlated mixed outcomes. Finally, we apply the proposed method to a dataset collected during an international synoptic krill survey in the Scotia Sea of the Antarctic Peninsula to infer the dependence structure between fin whales (Balaenoptera physalus), krill biomass, and relevant oceanographic data.

stat.ME

The Covariate-Assisted Bayesian Intransitive Bradley-Terry Model via Combinatorial Hodge Theory

Pairwise comparison data are widely used to recover latent rankings, yet the models in dominant use assume stochastic transitivity. When preferences are in fact intransitive, a single scalar strength conflates genuine hierarchy with cycle-induced structure, biasing both the recovered ranking and any covariate effects attributed to it. To address this limitation, we propose the Covariate-Assisted Bayesian Intransitive Bradley-Terry (CA-BIBT) model, which uses a combinatorial Hodge decomposition to resolve the latent match-up into identifiable and mutually orthogonal flows, attributing the component lying in the covariate-induced subspace to observed covariates and assigning the remaining components to the residuals. A global-local shrinkage prior on the residual cycle-induced flow adapts the model from transitive to intransitive regimes without prespecifying the regime, and a Gibbs sampler yields, as posterior byproducts, calibrated uncertainty for each flow, the posterior probability of the level at which the entities are rankable, and two complementary decision summaries that remain well-defined under intransitivity. In simulations, the CA-BIBT model recovers all flow components accurately with near-nominal coverage, and in applications to two animal dominance datasets, it distinguishes covariate-induced from residual cyclic dominance while quantifying posterior uncertainty, and demonstrates the practical utility of two complementary decision summaries.

stat.ME

Efficient Prior Sensitivity and Tipping-point Analysis for Medical Research: Revisiting Sampling Importance Resampling

Bayesian methods have received increasing attention in medical research, where sensitivity analysis of prior distributions is essential. Such analyses typically require the evaluation of the posterior distribution of a parameter under multiple alternative prior settings. When the posterior distribution of the parameter of interest cannot be derived analytically, the standard approach is to re-fit the model using Markov chain Monte Carlo (MCMC) for each setting, which incurs substantial computational costs. This issue is particularly relevant in tipping-point analysis, in which the posterior must be evaluated across gradually changing degrees of borrowing. Sampling-importance resampling (SIR) provides an efficient alternative by approximating posterior samples under new settings without MCMC re-fitting. Despite its potential computational advantages, the practical performance of SIR in repeated prior-sensitivity analyses, including tipping-point analysis, and in complex Bayesian models used in medical research has not been sufficiently illustrated. In this study, we illustrate the practical utility of SIR through two case studies: one involving tipping-point analysis under external data borrowing and another involving sensitivity analysis for a Bayesian nonparametric model in meta-analysis. In both examples, SIR substantially reduced computation time and produced posterior summaries similar to those obtained by MCMC re-fitting.

stat.ME

Tree-Embedded Bayesian Factor Models for Multidimensional Categorical Distributions

Analyzing data collected from multiple observational units to estimate common and heterogeneous structures through a hierarchical model is a central task in Bayesian inference, and to this end, Bayesian factor models are one of the most widely used tools for this purpose. In this paper, we propose a novel Bayesian latent factor model for categorical distributions from grouped data, providing a parsimonious model for describing many observed distributions through lower-dimensional structures. Grouped data arise in a wide range of applications in social science, for example, distributions of age composition and income observed across locations. In these contexts, standard mixture models can be inefficient because the distributions do not necessarily exhibit clear clustering structures, and the distributions can be more accurately approximated as a combination of lower-dimensional characteristics. To analyze distribution-valued data with the Bayesian factor analysis, we adopt a tree-based transformation that embeds distributions into a Euclidean space and construct a Bayesian latent factor model in the transformed space. We develop the hierarchical model by incorporating the infinite factor model, which can adaptively estimate the number of effective factors. In addition, we propose its generalization by incorporating a spatial dependence by introducing a prior based on a SAR model. The proposed model provides smooth estimates of multivariate distributional structures, because once a tree-based transformation is applied, both univariate and multivariate distributions are essentially treated as the same Euclidean vectors. Through numerical experiments using real population data, we demonstrate that the proposed model outperforms existing parametric and Bayesian nonparametric models in various scenarios involving smooth spatial variations, especially under small sample sizes.

stat.ME

Information Gap and Feasibility-Aware Inference in Binomial Logistic Mixtures

This paper studies the information gap between mixture detection and label recovery in binomial logistic mixtures. Standard likelihood-based criteria such as the Bayesian information criterion (BIC) can detect the presence of two components, but this does not guarantee that the corresponding labels are recoverable. We show that this gap is intrinsic to binomial logistic mixtures with a fixed number of trials: observed-data evidence for mixture structure and per-observation information for label recovery have different local orders in the component separation, and only the former accumulates with the sample size. As a result, there exists a detectable-but-unrecoverable regime in which BIC selects two components while the posterior labels remain essentially uninformative. To address this issue, we propose two feasibility-aware inference procedures: a recoverability-aware BIC with a posterior-entropy penalty and an entropy-regularized estimator that mitigates the tendency of the maximum likelihood estimator to produce overly separated components and overly concentrated posterior responsibilities. Numerical experiments confirm the predicted gap and demonstrate that the proposed methods avoid misleading component selections and improve the calibration of posterior label probabilities.

stat.ML

Propensity Patchwork Kriging for Scalable Inference on Heterogeneous Treatment Effects

Gaussian process-based models are attractive for estimating heterogeneous treatment effects (HTE), but their computational cost limits scalability in causal inference settings. In this work, we address this challenge by extending Patchwork Kriging into the causal inference framework. Our proposed method partitions the data according to the estimated propensity score and applies Patchwork Kriging to enforce continuity of HTE estimates across adjacent regions. By imposing continuity constraints only along the propensity score dimension, rather than the full covariate space, the proposed approach substantially reduces computational cost while avoiding discontinuities inherent in simple local approximations. The resulting method can be interpreted as a smoothing extension of stratification and provides an efficient approach to HTE estimation. The proposed method is demonstrated through simulation studies and a real data application.

stat.ME

Causal Small Area Estimation with Survey-only Covariates

Area-specific causal inference is important in many policy and survey applications, where the goal is to evaluate treatment effects for small geographic or demographic domains. Existing causal small area estimation methods, however, typically rely on a strong data requirement that treatment status is observed for all units in the population. This assumption is often unrealistic in practical survey settings, where both treatment and outcome variables are observed only for sampled units, while auxiliary covariates are available for the full population. To address this limitation, we develop a new identification strategy for area-specific treatment effects under this more realistic data structure by combining survey-only covariates with population-level auxiliary information. Based on this result, we propose a doubly robust estimator that remains consistent when either the outcome regression model or the treatment and area assignment models are correctly specified. We further derive the semiparametric efficiency bound for the target parameter and show that the proposed estimator attains this bound under regularity conditions. Simulation studies demonstrate favorable finite-sample performance, particularly in settings with small sample sizes within areas, and an empirical application illustrates the practical relevance of the proposed framework.

math.ST

Efficient Bayesian Inference in the Cox Model via Rank-Ordered Likelihood

In Bayesian inference for the Cox proportional hazards model, modeling the baseline hazard function is challenging. Recently, direct Bayesian inference using the partial likelihood is considered in the framework of general Bayesian inference. In terms of posterior computation, several studies have examined sampling algorithms under the Cox model. In this study, we propose two Gibbs sampling algorithms for Bayesian inference in the Cox proportional hazards model, motivated by a rank-ordered data representation and based on the Plackett--Luce and generalized Plackett--Luce models with P'{o}lya--Gamma data augmentation, referred to as PL-Cox and GPL-Cox, respectively. The two proposed methods offer practical advantages, as they do not require correction of posterior samples, naturally handle tied event times, and are readily extensible to shared frailty models. In simulation study, we considered multiple survival model settings, including continuous and discrete survival time models, as well as scenarios with varying degrees of ties, and found that the PL-Cox model exhibited relatively stable performance. In analyses of a large real dataset, the proposed methods remained computationally feasible, and the GPL-Cox model showed more favorable computational scalability than the PL-Cox model. In analyses of real data incorporating shared frailty, both methods demonstrated good computational efficiency.

stat.ME

Dynamic Bayesian regression quantile synthesis for forecasting outlook-at-risk

This paper proposes dynamic Bayesian regression quantile synthesis (DRQS), a novel method for quantile forecasting within the Bayesian predictive synthesis (BPS) framework designed to combine quantile-specific information from multiple agent models. While existing BPS approaches primarily focus on mean forecasting, our method directly targets the conditional quantiles of the response variable by utilizing the asymmetric Laplace distribution for the synthesis function. The resulting framework can be interpreted as a dynamic quantile linear model with latent predictors. We extend the univariate DRQS to a multivariate setting-factor DRQS (FDRQS)-by introducing a time-varying latent factor structure for the synthesis weights. This allows the model to leverage cross-sectional dependencies and shared information across multiple time series simultaneously. We develop an efficient Markov chain Monte Carlo (MCMC) algorithm for posterior inference, utilizing data augmentation and forward-filtering backward-sampling. Empirical applications to US inflation and global GDP growth demonstrate the improved performance of the proposed methods for quantile forecasting. In particular, FDRQS exhibits superior resilience during periods of extreme economic stress, such as the COVID-19 pandemic, by adaptively rebalancing agent contributions and capturing emergent global dependencies.

stat.ME

Direct Bayesian Additive Regression Trees for Conditional Average Treatment Effects in Regression Discontinuity Designs

Regression discontinuity designs (RDD) are widely used for causal inference. In many empirical applications, treatment effects vary substantially with covariates, and ignoring such heterogeneity can lead to misleading conclusions, which motivates flexible modeling of heterogeneous treatment effects in RDD. To this end, we propose a Bayesian nonparametric approach to estimating heterogeneous treatment effects based on Bayesian Additive Regression Trees (BART). The key feature of our method lies in adopting a general Bayesian framework using a pseudo-model defined through a loss function for fitting local linear models around the cutoff, which gives direct modeling of heterogeneous treatment effects by BART. Optimal selection of the bandwidth parameter for the local model is implemented using the Hyvärinen score. Through numerical experiments, we demonstrate that the proposed approach flexibly captures complicated structures of heterogeneous treatment effects as a function of covariates.

stat.ME