SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing gene set annotation methods, which rely heavily on manual curation and where large language models (LLMs) often overlook protein sequence-structure information. To overcome these challenges, this work proposes a hierarchical automated annotation framework that integrates protein language model representations with hybrid soft-hard prompting. Specifically, the framework employs ESM to encode protein sequences, aggregates features via a hierarchical attention encoder, and leverages a hybrid prompting mechanism to drive a local LLM for annotation generation. Evaluations on Gene Ontology (GO) and MSigDB benchmarks demonstrate significant improvements in annotation accuracy. Furthermore, this research reveals distinct contributions of cross-domain features, establishing an efficient new paradigm for gene set functional interpretation.
📝 Abstract
Gene set analysis is a cornerstone of functional genomics, yet it remains labor-intensive and heavily dependent on manual curation and expert biological interpretation. While Large Language Models (LLMs) have emerged as powerful tools for genomic reasoning and annotation, most existing approaches rely on symbolic gene names and fail to capture domain-specific biological structure, particularly protein sequence information that governs molecular activity, interactions, and downstream gene function. In this work, we propose SoftGene, a novel framework for LLM-based gene set annotation that leverages the hierarchical structure of gene sets. First, we use a hierarchical attention-based encoder built on ESM, a protein language model, to represent each gene set using protein-level amino acid sequence information. Second, we construct a hybrid prompting scheme that combines soft prompts derived from gene set embeddings with hard prompts containing auxiliary context generated by an LLM, and feed the resulting prompt into a local LLM for annotation. We evaluate our framework on two benchmark datasets: Gene Ontology (GO) and the Molecular Signatures Database (MSigDB). Our results show that integrating protein-sequence representations with textual context improves gene set annotation overall, while per-domain analyses reveal that the contribution of protein embeddings varies across biological domains.
Problem

Research questions and friction points this paper is trying to address.

Gene set annotation
Large Language Models
Protein sequence information
Functional genomics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Protein Language Model
Soft Prompting
Gene Set Annotation
Hierarchical Attention
Hybrid Prompting
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
D
Drew Ross
Electrical Engineering and Computer Science, University of Kansas, USA
A
Arya Hadizadeh Moghaddam
Electrical Engineering and Computer Science, University of Kansas, USA
Dongjie Wang
Dongjie Wang
Assistant Professor, University of Kansas
Machine LearningData-Centric AIRoot Cause AnalysisAutomated Urban Planning
X
Xiaoyu Zhang
Cell Biology and Physiology, University of Kansas Medical Center, USA
Zijun Yao
Zijun Yao
University of Kansas
AI/MLData MiningHealth Informatics