arXiv Science⌕ Search

arXiv · 2610.02682

Characterizing Identifiability in Boolean Factor Models

Abstract

Boolean factor models, including prominent subfamilies such as Boolean matrix decompositions and cognitive diagnosis models, find broad applications ranging from social sciences to engineering. Despite their flexibility, a key challenge lies in establishing the identifiability of their graphical structures, which specify how latent variables influence observed variables. Existing identifiability conditions typically rely on the strong assumption of pure nodes, which may be unrealistic in many applications. We develop a novel approach leveraging the Hasse diagram to represent the distribution of observed variables and transform identifiability into a graph isomorphism challenge. Based on this, we establish {\it sufficient and necessary} graphical identifiability conditions that do not require pure nodes. We further derive equivalent algebraic conditions and develop an efficient Boolean satisfiability (SAT)-based verification algorithm. We extend the analysis to probabilistic Boolean factor models, establishing conditions for jointly identifying the graphical structure and additional model parameters without requiring pure nodes. Our results substantially broaden the class of identifiable and interpretable Boolean factor models by removing the pure-node requirement, yielding new theoretical insights, while also providing practitioners with a concrete and easily implementable tool to assess model identifiability.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mengqi Lin, Gongjun Xu. 2026-10-02. Characterizing Identifiability in Boolean Factor Models. https://arxiv.org/abs/2610.02682

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

KL Convergence Guarantees for Score diffusion models under minimal data assumptions

Diffusion models are a new class of generative models that revolve around the estimation of the score function associated with a stochastic differential equation. Subsequent to its acquisition, the approximated score function is then harnessed to simulate the corresponding time-reversal process, ultimately enabling the generation of approximate data samples. Despite their evident practical significance these models carry, a notable challenge persists in the form of a lack of comprehensive quantitative results, especially in scenarios involving non-regular scores and estimators. In almost all reported bounds in Kullback Leibler (KL) divergence, it is assumed that either the score function or its approximation is Lipschitz uniformly in time. However, this condition is very restrictive in practice or appears to be difficult to establish. To circumvent this issue, previous works mainly focused on establishing convergence bounds in KL for an early stopped version of the diffusion model and a smoothed version of the data distribution, or assuming that the data distribution is supported on a compact manifold. These explorations have led to interesting bounds in either Wasserstein or Fortet-Mourier metrics. However, the question remains about the relevance of such early-stopping procedure or compactness conditions. In particular, if there exist a natural and mild condition ensuring explicit and sharp convergence bounds in KL. In this article, we tackle the aforementioned limitations by focusing on score diffusion models with fixed step size stemming from the Ornstein-Uhlenbeck semigroup and its kinetic counterpart. Our study provides a rigorous analysis, yielding simple, improved and sharp convergence bounds in KL applicable to any data distribution with finite Fisher information with respect to the standard Gaussian distribution.

math.ST↗

Estimation of conditional inequality curves and measures via estimating the conditional quantile function

In the paper conditional inequality curves and measures are proposed which allow us to describe the inequality/concentration of the conditional distribution of the feature we are interested in with respect to certain continuous variables. Moreover, for a graphical illustration of the change in values of the proposed conditional indices, a curve of conditional inequality measures is introduced. To estimate the curves and measures, a new method is proposed to estimate the conditional quantile function. This method uses quantile regression estimates for a given set of quantile orders, followed by isotonic regression on the estimated regression coefficients to ensure that the estimated conditional quantile function is nondecreasing. The consistency of the proposed estimators is proved while their finite sample performance is evaluated through simulation studies and compared with existing approaches. Finally, practical application of conditional curves and measures is demonstrated by determining estimated curves of conditional salary inequalities with respect to years of experience in different employee tenure groups, based on some real data. The code used to prepare the simulation results presented in this paper is available in a dedicated GitHub repository.

math.ST↗

Tracy-Widom, Gaussian, and Bootstrap: Approximations for Leading Eigenvalues in High-Dimensional PCA

Under certain conditions, the largest eigenvalue of a sample covariance matrix undergoes a well-known phase transition when the sample size $n$ and data dimension $p$ diverge proportionally. In the subcritical regime, this eigenvalue has fluctuations of order $n^{-2/3}$ that can be approximated by a Tracy-Widom distribution, while in the supercritical regime, it has fluctuations of order $n^{-1/2}$ that can be approximated with a Gaussian distribution. However, the statistical problem of determining which regime underlies a given dataset is far from resolved. We develop a new testing framework and procedure to address this problem. In particular, we demonstrate that the procedure has an asymptotically controlled level, and that it is power consistent for certain alternatives. Also, this testing procedure enables the design a new bootstrap method for approximating the distributions of functionals of the leading sample eigenvalues within the subcritical regime -- which is the first such method that is supported by theoretical guarantees.

math.ST↗