arXiv · 2610.03029
SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation
Abstract
Gene set analysis is a cornerstone of functional genomics, yet it remains labor-intensive and heavily dependent on manual curation and expert biological interpretation. While Large Language Models (LLMs) have emerged as powerful tools for genomic reasoning and annotation, most existing approaches rely on symbolic gene names and fail to capture domain-specific biological structure, particularly protein sequence information that governs molecular activity, interactions, and downstream gene function. In this work, we propose SoftGene, a novel framework for LLM-based gene set annotation that leverages the hierarchical structure of gene sets. First, we use a hierarchical attention-based encoder built on ESM, a protein language model, to represent each gene set using protein-level amino acid sequence information. Second, we construct a hybrid prompting scheme that combines soft prompts derived from gene set embeddings with hard prompts containing auxiliary context generated by an LLM, and feed the resulting prompt into a local LLM for annotation. We evaluate our framework on two benchmark datasets: Gene Ontology (GO) and the Molecular Signatures Database (MSigDB). Our results show that integrating protein-sequence representations with textual context improves gene set annotation overall, while per-domain analyses reveal that the contribution of protein embeddings varies across biological domains.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Drew Ross, Arya Hadizadeh Moghaddam, Dongjie Wang, Xiaoyu Zhang, Zijun Yao. 2026-10-02. SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation. https://arxiv.org/abs/2610.03029
Cite the original work for its findings. Save a collection to share your selection of sources.