arXiv ScienceSearch

arXiv · 2509.07031

Scalable Sample-to-Population Estimation of Hyperbolic Space Models for Hypergraphs

Abstract

Hypergraphs are useful mathematical representations of overlapping and nested subsets of interacting units, including groups of genes or brain regions, economic cartels, political or military coalitions, and groups of products that are purchased together. Despite the vast range of applications, the statistical analysis of hypergraphs is challenging: There are many hyperedges of small and large sizes, and hyperedges can overlap or be nested. Existing approaches to hypergraphs are either not scalable or achieve scalability at the expense of model realism. We develop a statistical framework that enables scalable estimation, simulation, and model assessment of hypergraph models, which is supported by non-asymptotic and asymptotic theoretical guarantees. First, we introduce a novel model of hypergraphs capturing core-periphery structure in addition to proximity, by embedding units in an unobserved hyperbolic space. Second, we achieve scalability by developing manifold optimization algorithms for learning hyperbolic space models based on samples from a population hypergraph. Third, we provide non-asymptotic and asymptotic theoretical guarantees for learning hyperbolic space models based on samples from a population hypergraph. We use the proposed statistical framework to detect core-periphery structure along with proximity among U.S.\ politicians based on historical media reports.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Cornelius Fritz, Yubai Yuan, Michael Schweinberger. 2026-02-01. Scalable Sample-to-Population Estimation of Hyperbolic Space Models for Hypergraphs. https://arxiv.org/abs/2509.07031

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Robust Estimation and Inference with Categorical Data

Categorical data pose a distinctive robustness challenge: contingency-table cells need not have a meaningful magnitude, ordering, or metric, so departures must instead be assessed through discrepancies between observed cell frequencies and model-implied probabilities. We develop $C$-estimation, a unifying framework for robust estimation in structured categorical models. $C$-estimation limits the influence of large frequency discrepancies and can be applied to unconditional, composite, and regression models of categorical data. Building on minimum-disparity estimation, the framework also accommodates clipped nonsmooth loss functions. A common asymptotic theory establishes Fisher consistency, consistency for the population target and asymptotic normality under contamination, and sandwich covariance estimation. At a correctly specified model, regular $C$-estimators retain the first-order efficiency of maximum likelihood. To quantify global robustness, we derive computable lower and upper envelopes for maximum-bias curves and, for a Huber-like loss, connect its clipping constants to a global robustness bound. Simulations support the theory and illustrate the estimators' robustness. An application to questionnaire data illustrates robust estimation of a latent factor model and identifies misfitting response strings that may reflect careless responding. A software implementation is provided.

stat.ME

Multi-Attribute Preferences: A Transfer Learning Approach

We introduce a transfer-learning method based on the Bradley--Terry model for multi-attribute pairwise-comparison data. The aim is to estimate the log-worth parameters of one primary attribute while using information from related secondary attributes. The method first pools the primary data with data from informative secondary attributes. It then corrects this pooled estimate using the primary likelihood, with a ridge penalty controlling the size of the correction. When the informative set is unknown, we use held-out primary data to select secondary attributes. For a known informative set and under a pooled Bradley--Terry compatibility condition, we derive high-probability $\ell_\infty$ and $\ell_2$ error bounds. Under additional conditions, these bounds can have a smaller asymptotic order than the corresponding primary-only Bradley--Terry bounds. We also establish asymptotic normality for a one-step estimator. A simulation study evaluates the method under more general settings, and an application to consumer preferences for eba, a cassava-derived food product, illustrates its use and interpretation. An R package implementing the method is available at https://CRAN.R-project.org/package=BTTL.

stat.ME

Spatially Dependent Indian Buffet Processes

We develop a new stochastic process called spatially dependent Indian buffet processes (sIBP) for binary feature matrices of unbounded columns with spatial correlations between subjects, and propose general spatial factor models for various multivariate response variables. We introduce spatial dependency through the stick-breaking representation of the original Indian buffet process (IBP; Griffiths and Ghahramani, 2005, 2011) and latent Gaussian process for the logit-transformed breaking proportions to capture underlying spatial correlation. We show that sIBP retains the sparsity and finite-feature behavior of the original IBP, while its joint feature allocation probabilities are affected by spatial correlation. Using binomial expansion and Polya-gamma data augmentation, we provide an efficient Gibbs sampler for posterior computation. The usefulness of our sIBP is demonstrated through simulation studies and two applications for large-dimensional multinomial data of areal dialects and geographical distribution of multiple tree species.

stat.ME