arXiv ScienceSearch

arXiv · 2510.22174

Identification of Shared Genetic Biomarkers to Discover Candidate Drugs for Cervical and Endometrial Cancer by Using the Integrated Bioinformatics Approaches

Abstract

Cervical (CC) and endometrial cancers (EC) are two common types of gynecological tumors that threaten the health of females worldwide. Since their underlying mechanisms and associations remain unclear, computational bioinformatics analysis is required. In the present study, bioinformatics methods were used to screen for key candidate genes, their functions and pathways, and drug agents associated with CC and EC, aiming to reveal the possible molecular-level mechanisms. Four publicly available microarray datasets of CC and EC from the Gene Expression Omnibus database were downloaded, and 72 differentially expressed genes (DEGs) were selected through integrated analysis. Then, we performed the protein-protein interaction (PPI) analysis and identified 9 shared genetic biomarkers (SGBs). The GO functional and KEGG pathway enrichment analyses of these SGBs revealed some important functions and signaling pathways significantly associated with CC and EC. The interaction network analysis identified four transcription factors (TFs) and two miRNAs as key transcriptional and post-transcriptional regulators of SGBs. The expression of the AURKA, TOP2A, and UBE2C genes was higher in CC and EC tissues than in normal samples, and this gene expression was linked to disease progression. Furthermore, we performed docking analysis between 9 SGBs-based proteins and 145 meta-drugs, and identified the top-ranked 10 drugs as candidate drugs. Finally, we investigated the binding stability of the top-ranked three drugs (Sorafenib, Paclitaxel, Sunitinib) using 100 ns MD-based MM-PBSA simulations with UBE2C, AURKA, and TOP2A proteins, and observed their stable performance. Therefore, the proposed drugs might play a vital role in the treatment against CC and EC.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Md. Selim Reza, Mst. Ayesha Siddika, Md. Tofazzal Hossain, Md. Ashad Alam, Md. Nurul Haque Mollah. 2025-10-25. Identification of Shared Genetic Biomarkers to Discover Candidate Drugs for Cervical and Endometrial Cancer by Using the Integrated Bioinformatics Approaches. https://arxiv.org/abs/2510.22174

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

EPI-KAN: A Method For Estimating and Forecasting Time-Dependent COVID-19 Parameters

We introduce EPI-KAN, a novel method for estimating COVID-19 time-varying parameters. EPI-KAN uses historical epidemiological data, Physics-Informed Neural Network (PINN), and the novel Kolmogorov-Arnold Network (KAN). The method harnesses the novel Kolmogorov-Arnold Network (KAN), which is a type of artificial neural network. For the KAN in this paper, we learn activation functions that are represented using Fourier series, hence we abbreviate as KAN-F. In this study, we estimate parameters in the context of an SIRD compartmental differential equations. The time-dependent parameters are the transmission rate $β(t)$, recovery rate $γ(t)$, and mortality rate $μ(t)$. We define three KAN-F functions $\widehatβ$, $\widehatγ$, $\widehatμ$ that model the true parameters $β(t)$, $γ(t)$, $μ(t)$, respectively. We test two model architectures for the KAN-F: the first has 8 input variables consisting of $S$, $I$, $R$, $D$, and their numerical gradients at any time $t$, while the second has 4 input variables excluding the numerical gradients. The objective loss function that has to be minimized is subject to Physics-Informed Neural Network (PINN). Using historical data of COVID-19 from three South-East Asian countries: Indonesia, Singapore, and Malaysia, we are able to estimate $β(t)$, $γ(t)$, and $μ(t)$ on each country with decent accuracy and efficiency. The time period of choice coincides with the period where SARS-CoV-2 Delta variant (B.1.617.2) was dominant. In addition to estimating the rates during the training period, we also predict transmission rates over 30 days during forecast period. We found that the output of KAN-F over the forecast period can give good predictions if we scale the output by a factor of 17\% for Indonesia and 30\% for Singapore and Malaysia.

q-bio.OT

Making Models That Matter: How to Build Trustworthy and Useful Systems Biology Models

Computational models supporting mechanistic understanding of (complex) biological systems, systems behaviour prediction, and experimental design are becoming more and more embedded in research on complex biological systems. Reuse and refinement of models, rather than continuous reinvention, is becoming increasingly important as models' demands on computational infrastructure increase. However published models - despite the variety of efforts taken so far - are frequently difficult to reproduce or reuse, substantially limiting their scientific value. Here we address the requirements for model reusability in the light of the field-specific CURE framework (Credible, Understandable, Reproducible, Extensible) and the more general FAIR principles (Findable, Accessible, Interoperable, Reusable). Considering published guidance we identify broad agreement on requirements for findability, accessibility, and interoperability, but continued lack of clarity and consensus around reusability. Focusing on the scientific quality and usability of computational models we discuss six key practices underpinning model sharing and re-use. Mapping the FAIR and CURE principles onto the model lifecycle we propose ten recommendations for building and sharing systems biology models that are both FAIR- and CURE-compliant.

q-bio.OT

SenSASP: A Unified, Multi-Layer Database of Senescence and SASP Genes

Research on cellular senescence and the senescence-associated secretory phenotype (SASP) draws on independently curated gene resources that differ in scope, identifiers, and update cycles, making cross-resource integration error-prone. We unified four widely used resources, CellAge, GenAge, the SenMayo signature, and the Reactome Cellular Senescence pathway, onto a single canonical identifier (the Ensembl gene ID) and enriched every gene with three annotation layers absent from all four inputs: cross-species conservation, tissue and cell-type expression, and high-confidence protein-protein interactions. Unification collapsed 1,460 summed source entries into 1,250 unique genes (210 redundant entries removed, 14.4%) while preserving full source provenance: 173 genes are corroborated by two or more resources and two (IL6, JUN) by all four. The three annotation layers reach 95.8%, 97.9%, and 93.0% of genes, with 89.4% annotated across all three. A 500-gene random sample of identifier mappings was validated against HGNC and Ensembl (98.0% exact match). The result, SenSASP, is a single, machine-readable, provenance-tracked database of harmonized identifiers and net-new functional context, illustrated here with a gene-prioritization score and a tissue-expression atlas. SenSASP is freely available at https://xuan13hao.github.io/sensasp/

q-bio.OT