arXiv ScienceSearch

arXiv · 2407.04465

A Compounded Burr Probability Distribution for Fitting Heavy-Tailed Data with Applications to Biological Networks

Abstract

Complex biological networks, encompassing metabolic pathways, gene regulatory systems, and protein-protein interaction networks, often exhibit scale-free structures characterized by heavy-tailed degree distributions. However, empirical studies reveal significant deviations from ideal power law behavior, underscoring the need for more flexible and accurate probabilistic models. In this work, we propose the Compounded Burr (CBurr) distribution, a novel four parameter family derived by compounding the Burr distribution with a discrete mixing process. This model is specifically designed to capture both the body and tail behavior of real-world network degree distributions with applications to biological networks. We rigorously derive its statistical properties, including moments, hazard and risk functions, and tail behavior, and develop an efficient maximum likelihood estimation framework. The CBurr model demonstrates broad applicability to networks with complex connectivity patterns, particularly in biological, social, and technological domains. Extensive experiments on large-scale biological network datasets show that CBurr consistently outperforms classical power-law, log-normal, and other heavy-tailed models across the full degree spectrum. By providing a statistically grounded and interpretable framework, the CBurr model enhances our ability to characterize the structural heterogeneity of biological networks.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tanujit Chakraborty, Swarup Chattopadhyay, Suchismita Das, Shraddha M. Naik, Chittaranjan Hens. 2025-04-26. A Compounded Burr Probability Distribution for Fitting Heavy-Tailed Data with Applications to Biological Networks. https://doi.org/10.1063/5.0270403

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Quantifying Portfolio Demutualization: A Benchmark-Relative Pooling--Profiling Scale

Insurance pricing combines pooling with differentiation: a tariff may leave benchmark differences in expected loss partly mutualized or translate them into policy-level premium differences. We propose a benchmark-relative pooling--profiling scale with two complementary coordinates. The coupled $L^p$ coordinate measures policy-level alignment between an evaluated tariff and a stated benchmark pure premium, whereas the marginal Wasserstein coordinate compares their exposure-weighted premium distributions. The difference between their residual $p$-costs defines an allocation mismatch. Under portfolio balance, the coupled $L^1$ coordinate has an exact actuarial interpretation: it is the fraction of the transfer volume induced by full pooling that the tariff removes. Synthetic and motor-insurance applications show that broad classes, proxies, shrinkage and tail caps can affect marginal differentiation, policy-level allocation and transfers differently. A barycentric group-parity intervention further shows that conditional premium disparities can fall mainly through reallocation and restored benchmark-relative transfers, with little change in marginal differentiation.

stat.AP

Constructing Reliable Social Networks from Conversational Data: An Ensemble Prompt Engineering Approach with Uncertainty Quantification

Conversational data are central to the study of interaction dynamics and social structures across psychological research. However, constructing structured social networks from unstructured conversational data remains a major methodological challenge. This study presents a pipeline for network construction using prompt engineering. We employ an ensemble of five Large Language Models (LLMs) with majority voting to automate utterance classification, reducing dependence on manual coding without task-specific parameter fine-tuning. Classification uncertainty is assessed through an uncertainty quantification framework based on Shannon entropy, which can be used to prioritize ambiguous cases for review. The classified utterances are used to construct directed interaction networks for subsequent analysis. Reliability and accuracy are established relative to the criterion-referenced and human comparisons reported here rather than as unconditional guarantees across settings. We demonstrate the utility of this approach through two illustrative applications to classroom interaction data: network centrality analysis to characterize participant roles, and network mediation analysis using the additive and multiplicative effects network (AMEN) model to examine how interaction structures mediate the relationship between gender and mathematics performance. This pipeline provides a scalable foundation for automated network construction from conversational data across diverse research contexts.

stat.AP

Issue-Specific Polarization and Cohesion in a Multi-Party Legislature: Integrating the Latent Space Item Response Model with Topic-Based Regression

We develop a two-stage, cut-posterior Bayesian framework for quantifying issue-specific legislative alignment in multi-party systems. The approach integrates a Latent Space Item Response Model (LSIRM), embedding legislators and bills in a shared Euclidean space, with Bayesian beta regression using text-derived topic proportions as bill-level covariates. The resulting legislator- and issue-specific coefficients allow polarization and cohesion to be compared across policy domains. The beta-regression layer must not feed back into the estimated latent geometry, so that the two components target a cut posterior. We estimate it with a two-stage Multiple-Imputation procedure that propagates uncertainty in the latent positions into every downstream quantity. In the 17th Korean National Assembly, fiscal domains such as Taxation and Grants and Local Government Budget show sharp polarization with tight within-party clustering, whereas Armed Services, Patriots, and Veterans exhibits weak party structuring and greater intra-party variability. The Democratic Labor Party forms a distinct cluster on several issues even where the two major parties are not strongly polarized, showing that legislative conflict escapes a single left--right ordering. The framework supports analysis of issue-structured voting in legislatures where one-dimensional ideal point models are unreliable.

stat.AP