arXiv ScienceSearch

arXiv · 2101.07619

Cancer driver gene detection in transcriptional regulatory networks using the structure analysis of weighted regulatory interactions

Abstract

Identification of genes that initiate cell anomalies and cause cancer in humans is among the important fields in the oncology researches. The mutation and development of anomalies in these genes are then transferred to other genes in the cell and therefore disrupt the normal functionality of the cell. These genes are known as cancer driver genes (CDGs). Various methods have been proposed for predicting CDGs, most of which based on genomic data and based on computational methods. Therefore, some researchers have developed novel bioinformatics approaches. In this study, we propose an algorithm, which is able to calculate the effectiveness and strength of each gene and rank them by using the gene regulatory networks and the stochastic analysis of regulatory linking structures between genes. To do so, firstly we constructed the regulatory network using gene expression data and the list of regulatory interactions. Then, using biological and topological features of the network, we weighted the regulatory interactions. After that, the obtained regulatory interactions weight was used in interaction structure analysis process. Interaction analysis was achieved using two separate Markov chains on the bipartite graph obtained from the main graph of the gene network. To do so, the stochastic approach for link-structure analysis has been implemented. The proposed algorithm categorizes higher-ranked genes as driver genes. The efficiency of the proposed algorithm, regarding the F-measure value and number of identified driver genes, was compared with 23 other computational and network-based methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mostafa Akhavan Safar, Babak Teimourpour, Abbas Nozari-Dalini. 2021-01-19. Cancer driver gene detection in transcriptional regulatory networks using the structure analysis of weighted regulatory interactions. https://doi.org/10.2174/1574893617666220127094224

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Thermodynamic and Statistical Signatures of Modality Changes in Concentration Distributions Driven by Stochastic Switching Between Two Activity States

Stochastic switching between gene expression states, coupled with production and degradation dynamics, governs the accumulation of mRNA and proteins in cells. The concentrations of these accumulated entities dictate the phenotypic distribution of genetically identical cells. The underlying accumulation dynamics are well-captured by a two-state promoter switching model, with statistical and thermodynamic properties quantified via the Fano factor and entropy production rates. However, how these measures correlate with concentration distributions and their shifts under varying kinetic parameters remains largely unexplored. To this end, we use chemical master equations to study a generalized model of mRNA accumulation dynamics in the presence of stochastic switching between two activity states and state-dependent production and degradation rates. We derive exact expressions for the steady-state probability distribution and analytically compute the mean concentration, Fano factor, and entropy production rate (EPR). Simplifying these expressions, we identify contributions arising from stochastic switching rates and relaxation dynamics toward equilibrium in each activity state. Next, using our theoretical results, we characterize the variation in the Fano factor and EPR as a function of mean expression during modality changes of the distributions mediated by the variation of switching rates. We also identify the conditions in kinetic parameters that achieve the highest Fano factor and entropy production rates. Our findings establish a generalized framework for examining stochastic accumulation dynamics, clarifying how kinetic parameters dictate molecular distributions, noise, and dissipation. These insights extend readily to broader contexts coupling stochastic switching with accumulation, including protein burst dynamics, phenotype-switching-mediated drug intake, and queuing theory.

q-bio.MN

Are You Learning Biological Signal or Shortcuts? Auditing and Mitigating Bias in Protein-Protein Interaction Datasets

Protein-protein interaction (PPI) databases do not faithfully reflect biological realities. Instead, they are influenced by study and technical biases that distort certain protein and interaction attributes. Machine learning models can exploit these as learning shortcuts if the negative dataset is not constructed with care. So far, the shortcuts introduced during PPI dataset construction have only been examined in isolation. Here, we systematically characterize both reported and, to our knowledge, previously unreported biases in PPI datasets that lead machine learning models to learn shortcuts instead of biological signal. We analyze HIPPIE, IntAct, and STRING, dedicated PPI databases, as well as two datasets derived from 3D-structural information in the Protein Data Bank (PDB). We show that random data splitting introduces strong topological shortcuts. When train-test protein overlap is removed, the resulting datasets still retain usable shortcuts stemming from self-interactions, taxonomic identity, and functional relatedness, whose prevalence interestingly depends on the data source. We further show that sampling negatives from a set of high-confidence non-interactors, an intuitively appealing choice, can amplify the shortcut stemming from functional relatedness. To detect and mitigate these biases, we provide an open Nextflow pipeline that combines similarity-aware, data-loss-minimizing dataset splitting with bias-minimizing negative sampling, both formulated as integer linear programs. Its key concept of quantifying biases to minimize them through optimization-based negative sampling can, in principle, be extended to any machine learning problem where the pool of negative candidates is much larger than the positives and is thus of interest also beyond PPI prediction.

q-bio.MN

Obstruction of Absolute Concentration Robustness by Conservation Laws in Non-Redundant Zero-One Networks

Absolute concentration robustness (ACR) is a structural property of biochemical reaction networks in which a species attains the same steady-state concentration at every positive steady state, independently of initial conditions and rate constants. Existing detection methods rely on algebraic elimination and typically scale exponentially with network size. We develop a topology-based alternative for non-redundant zero-one networks of stoichiometric dimension at most two, a class that already captures enzyme catalysis, carbon-nanotube transitions, and other elementary biochemical mechanisms. Organizing our analysis around a structural index $s^*$, the number of distinct rows in the stoichiometric matrix, we obtain a complete classification of all such networks admitting non-vacuous ACR for generic rate constants. In dimension one, ACR occurs only for the elementary inflow and outflow module. In dimension two, ACR is possible if and only if $s^*\leq 3$; for $s^*=3$, the admissible networks are precisely those obtained as species refinements of consistent subnetworks of five canonical biochemical prototypes. For $s^*\geq 4$, non-vacuous ACR is impossible for any generic rate assignment.

q-bio.MN