🤖 AI Summary
This study addresses the limitations of existing gene set annotation methods, which rely heavily on manual curation and where large language models (LLMs) often overlook protein sequence-structure information. To overcome these challenges, this work proposes a hierarchical automated annotation framework that integrates protein language model representations with hybrid soft-hard prompting. Specifically, the framework employs ESM to encode protein sequences, aggregates features via a hierarchical attention encoder, and leverages a hybrid prompting mechanism to drive a local LLM for annotation generation. Evaluations on Gene Ontology (GO) and MSigDB benchmarks demonstrate significant improvements in annotation accuracy. Furthermore, this research reveals distinct contributions of cross-domain features, establishing an efficient new paradigm for gene set functional interpretation.
📝 Abstract
Gene set analysis is a cornerstone of functional genomics, yet it remains labor-intensive and heavily dependent on manual curation and expert biological interpretation. While Large Language Models (LLMs) have emerged as powerful tools for genomic reasoning and annotation, most existing approaches rely on symbolic gene names and fail to capture domain-specific biological structure, particularly protein sequence information that governs molecular activity, interactions, and downstream gene function. In this work, we propose SoftGene, a novel framework for LLM-based gene set annotation that leverages the hierarchical structure of gene sets. First, we use a hierarchical attention-based encoder built on ESM, a protein language model, to represent each gene set using protein-level amino acid sequence information. Second, we construct a hybrid prompting scheme that combines soft prompts derived from gene set embeddings with hard prompts containing auxiliary context generated by an LLM, and feed the resulting prompt into a local LLM for annotation. We evaluate our framework on two benchmark datasets: Gene Ontology (GO) and the Molecular Signatures Database (MSigDB). Our results show that integrating protein-sequence representations with textual context improves gene set annotation overall, while per-domain analyses reveal that the contribution of protein embeddings varies across biological domains.