arXiv ScienceSearch

arXiv subjects

Samaneh Nazari

Publications and source records attributed to Samaneh Nazari.

4 recordsLinked to original sources

Scalable Bayesian structure learning of directed acyclic graphs via Laplace approximation, with an application to breast cancer gene expression networks

Structure learning of directed acyclic graphs (DAGs) from observational data is a foundational task in causal discovery and is widely used to infer regulatory networks from medical and genomic measurements. The Bayesian formulation quantifies model uncertainty and admits prior biological knowledge, but its practical use has been hampered by the super-exponential growth of the DAG space and by the intractability of the node-marginal likelihood under flexible, non-conjugate priors. Existing closed-form solutions are largely confined to the conjugate Normal--Inverse-Gamma prior. We develop a Laplace-approximated Bayesian scoring function for the non-conjugate Normal--Gamma prior on the modified Cholesky parameterisation of the precision matrix, embed it in a Metropolis--Hastings sampler over DAGs, and couple the latent Gaussian network to a binary clinical outcome through a probit link. We show that the node-marginal integral is of generalised inverse-Gaussian form, so that its exact value is a modified Bessel function of the second kind and the proposed scoring function is its leading large-argument asymptotic; the posterior of each conditional variance is likewise generalised inverse-Gaussian and is sampled exactly. In simulation, the proposed prior improves on the conjugate baseline and on the PC, greedy-equivalence-search, NOTEARS, and DAGMA benchmarks at sample sizes typical of clinical cohorts. On two real datasets, the Sachs protein-signalling network, scored against its validated consensus graph, and the Wisconsin Diagnostic Breast Cancer data, the method recovers known structure and, through the DAG-probit extension, predicts malignancy from nuclear morphometry with a cross-validated ROC-AUC of $0.94$ using a sparse, interpretable set of direct predictors.

stat.ME

Semiparametric Bayesian structure learning of nonparanormal directed acyclic graphs with local--global shrinkage

Bayesian structure learning of directed acyclic graphs (DAGs) is central to high-dimensional causal discovery, yet existing methods mostly assume multivariate Gaussian data, which is routinely violated by measurements displaying heavy tails, skewness, or bounded support. We address this by introducing NPN-DAG-HS, a fully Bayesian semiparametric DAG model coupling the nonparanormal family with horseshoe shrinkage on Cholesky off-diagonals via an extended-rank likelihood. This handles arbitrary unknown monotone marginal transformations without estimating them, preserving exact-zero shrinkage and conditionally conjugate posteriors for practical inference in large dimensions. Our method uses a partially collapsed Metropolis-within-Gibbs sampler that augments rank-likelihood Gaussian copies and alternates score-based DAG moves with horseshoe Gibbs updates. Theoretically, in the high-dimensional regime ($p \to \infty$ with $\log p / n \to 0$), we establish posterior contraction at the rate $\sqrt{(s_0+p)\log p/n}$, strong skeleton selection consistency, and a parametric Bernstein-von Mises theorem for smooth total causal-effect functionals, yielding asymptotic frequentist calibration of credible intervals. Simulations confirm our method outperforms Gaussian baselines and standard frequentist learners. Applied to an acute myeloid leukaemia dataset, it recovers a sparse, interpretable network, pruning unsupported edges declared by Gaussian analyses and illustrating the value of our semiparametric relaxation.

stat.ME

Nonparanormal Bayesian Learning of Directed Acyclic Graphs under Gamma and Inverse-Gamma Innovation Priors: Closed-Form Scores and Informed Sampling

Bayesian structure learning for directed acyclic graphs (DAGs) is a central tool for reconstructing biological networks, yet it often assumes the data are jointly Gaussian. In motivating proteomic applications, this assumption is routinely violated by skewed and heavy-tailed data, causing Gaussian DAGs to recover spurious or misdirected edges. We develop a fully Bayesian framework for DAG learning in the nonparanormal family, replacing Gaussianity with the weaker requirement that unknown strictly increasing marginal transformations are jointly Gaussian. Working on the modified Cholesky parameterization of the latent precision matrix, we introduce two innovation-variance priors: a non-conjugate Normal-Gamma prior, which decouples coefficient shrinkage from variance regularization, and a conjugate Normal-Inverse-Gamma prior. For both, we obtain the node-wise marginal likelihood in closed form -- through a modified Bessel function of the third kind for the Gamma prior and a Student-$t$ form for the Inverse-Gamma prior -- allowing MCMC sampler moves to be scored without numerical integration. Exploiting these, we build a locally-balanced informed sampler and a Bessel-free score that scales the sampler to hundreds of nodes. On simulated data, our nonparanormal samplers match Gaussian methods when data are Gaussian and dominate them sharply when margins are skewed. On human T-cell protein-signalling data, they recover well-established interactions at high posterior probability, clearly outperforming constraint-based competitors and performing comparably to a Gaussian Bayesian model.

stat.ME

Bayesian DAG Structure Learning with Simultaneous Shrinkage Covariance Estimation under Scale-Mixture Error Distributions in the Proportional High-Dimensional Regime

We propose a unified Bayesian framework namely robust DAG-Cholesky horseshoe (R-DACH) for joint directed acyclic graph (DAG) structure learning and precision matrix estimation in the high-dimensional proportional asymptotic regime $p/n \to c \in (0,\infty)$, under the scale mixture of normal errors. The construction places a global-local horseshoe-type prior directly on the strictly lower-triangular entries of the modified Cholesky factor of the DAG-Markov precision matrix, so that sparsity in the Cholesky parameters induces a coherent parent-set selection consistent with a topological ordering of the variables. A per-observation inverse-gamma scale mixture yields automatic robustness to heavy-tailed and contaminated observations and admits Student-$t$, Laplace, and slash distributions as special cases. We design a partially-collapsed blocked Gibbs sampler that traverses the joint space of orderings, sparsity patterns and continuous parameters. Simulations across $(n,p)$ configurations with $p$ up to several hundreds confirm the theoretical rates and demonstrate substantial gains over graphical-horseshoe, DAG-Wishart, and PC-based competitors under contamination. An application to RNA-seq gene-expression data from \emph{The Cancer Genome Atlas} reveals biologically interpretable regulatory structure that competing methods fail to recover.

stat.ME