arXiv Science⌕ Search

arXiv · 2609.37735

CancerZigZag: Iterative Seed-Anchored Diffusion for Generative Modeling of Single-Cell State Transitions

Abstract

Single-cell cancer datasets are predominantly cross-sectional and rarely provide paired or longitudinal observations linking individual healthy-like cells to tumor-associated states. We introduce CancerZigZag, a seed-initialized diffusion-based framework for exploratory generation of tumor-associated single-cell candidate clouds from unpaired epithelial cell populations. For each cancer context, a diffusion model is trained exclusively on tumor-derived epithelial cells and applied through repeated partial latent-space perturbation and reverse diffusion to held-out healthy-like seeds, generating stochastic candidate clouds without paired measurements or classifier guidance. We applied CancerZigZag to colorectal, breast, lung, and renal cell carcinoma contexts and explored parameter landscapes defined by perturbation depth and the number of ZigZag cycles. Across the reported operating configurations, candidate clouds contained outputs classified toward held-out tumor-derived reference populations for each evaluated seed. Residual seed-dependent organization varied across contexts, with the clearest structure in colorectal cancer, more modest organization in lung cancer, and limited cloud-level structure in breast and renal cell carcinoma. Representative candidates also showed directional concordance with transcriptional shifts observed between held-out healthy-like and tumor-derived reference populations. CancerZigZag is not interpreted as a model of deterministic healthy-to-tumor transformation or cellular progression. Instead, it provides a reference-informed framework for exploring tumor-associated candidate distributions from unpaired healthy-like seeds and quantifying the context-dependent relationship between tumor-associated displacement and residual seed dependence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Johannes Schlüter, Alexander Schönhuth. 2026-09-29. CancerZigZag: Iterative Seed-Anchored Diffusion for Generative Modeling of Single-Cell State Transitions. https://arxiv.org/abs/2609.37735

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

CellMSA: Context Modeling for Single-Cell Representation Learning

Single-cell transcriptomics enables profiling of cellular states at unprecedented resolution, but its high dimensionality, sparsity, and technical batch effects pose significant challenges for representation learning. Existing single-cell foundation models typically encode each cell independently or only model cells from the same batch for denoising, thereby underutilizing the rich relational information across batches and cell types to model gene expression patterns. We argue that single-cell models can benefit from more informative cell-context modeling. By comparing consistency and variation across cells, models can capture fine-grained gene-gene dependencies associated with cell states, which are essential for learning high-quality representations. Inspired by the use of multiple sequence alignment (MSA) context in protein modeling, we propose CellMSA, a single-cell representation learning framework that introduces an MSA-inspired inductive bias into transcriptomic modeling. For each target cell, CellMSA retrieves relevant cells from different batches and biologically related cell types as context, and summarizes cross-cell patterns into a context-dependent gene-pair representation. This representation is then injected into a pair-aware target-cell encoder for fine-grained representation learning. We pretrain CellMSA on a large-scale human single-cell corpus of approximately 109 million cell observations, including 65.6 million primary observations. Experiments show that our framework consistently outperforms existing methods across multiple benchmarks. Code is available at the following repository: https://github.com/PharMolix/CellMSA.

q-bio.GN↗

Fragment Path Complex Networks for Viral Genome Classification

Accurate classification of viral families is essential to understanding genomic diversity, but rapid evolution and heterogeneous genome structure complicate reliable assignment. Newly discovered viruses, particularly those from underrepresented ecosystems, expand the known diversity of viral genomes. This growing diversity poses challenges for classification methods that rely on similarity to existing references. We present Fragment Path Complex Networks (FPCN), a framework that represents viral genomes as sequences of ordered fragments at various scales. FPCN captures contextual sequence information within each fragment and uses path complexes to model complex relationships between neighboring fragments. We evaluated FPCN on well-established NCBI benchmarks, and FPCN achieved higher classification accuracy than competing methods. Further evaluations on newly released datasets revealed FPCN's ability to generalize to genomes that were not presented in the training. Together, these results position FPCN as an effective tool for classifying viral genomes within known families across benchmark datasets and newly released genomes.

q-bio.GN↗

Learning Interpretable Tumor Microenvironment Representations by Fitting Pan-Cancer Cell State-Niche Correlation

In the tumor microenvironment, a cell's state is influenced by cell-cell interactions (CCIs) with neighboring cells in its niches. Identifying dysregulated CCIs that are associated with pathogenic processes pinpoints targets for drug discovery. Imaging-based spatial transcriptomics and single-cell RNA sequencing provide, respectively, single-cell spatial information and transcriptome-wide measurements needed to study CCIs, but neither modality provides both. Existing spatial transcriptomics foundation models also cannot effectively learn from spatially resolved single-cell data with full-transcriptome coverage, explicitly infer the CCI mechanisms driving cell state-niche associations, or be interpretable enough to support direct biological interpretations. Here, we present GITIII-scale, a hierarchical, interpretable pan-cancer spatial transcriptomics foundation model for TME representation learning that investigates cell state-niche associations and their underlying ligand-receptor (LR) signaling pathways. GITIII-scale uses transformers to model interactions between pairs of cells at defined spatial distances, an interpretable single-layer graph transformer without a feed-forward network to decompose how each gene in a receiver cell is influenced by each neighboring sender cell, and a graph transformer to generate cellular-neighborhood embeddings. Trained on our assembled pan-cancer database of specimen-matched scRNA-seq and imaging-based spatial transcriptomics datasets, GITIII-scale generated TME embeddings that recovered niche-associated state changes more accurately than existing spatial transcriptomics foundation models in cancer types unseen during training. A case study of an unseen breast cancer dataset further demonstrated the model's interpretability by identifying potentially drug-targetable LR pathways associated with endothelial overgrowth and tumorigenesis.

q-bio.GN↗