Score
Designs and implements computational and statistical workflows to identify and validate gene-expression biomarkers from transcriptomic and related genomics data, including differential expression testing, feature selection, and biomarker signature construction. Produces reproducible analyses and interpretable outputs such as validated biomarker sets, performance summaries, and reviewer-ready interpretation reports.
Current evaluations of bioinformatics agents overemphasize answer correctness while neglecting workflow auditability and scientific credibility. This work proposes a Function–Evidence–Validation (FEV) tri-dimensional evaluation framework centered on inspectable workflow trajectories, shifting the primary focus to workflow correctness for the first time. Through systematic literature review, trajectory analysis, and cross-domain benchmark mapping, the study comprehensively analyzes 109 agent systems and 28 evaluation resources across subfields including genomics, single-cell and spatial omics, and protein science. The findings reveal that while agents perform adequately in planning and execution, they exhibit significant deficiencies in reproducibility, traceability, external validation, and prospective experimental design. This research provides both theoretical grounding and practical guidance for developing transparent, auditable next-generation bioinformatics agents.
To address inflated false discovery rates (FDR) in high-dimensional RNA-seq differential expression analysis arising from multiple testing, this study develops a robust statistical inference framework. Methodologically, we systematically compare three FDR control procedures—Benjamini–Hochberg (BH), Benjamini–Yekutieli (BY), and Storey’s q-value—and introduce an adaptive q-value approach to enhance statistical power under sparsity and high dimensionality. Visualization and performance evaluation integrate PCA, volcano plots, MA plots, and confusion matrices. Our key contribution is the first systematic quantification—within transcriptomic contexts—of how inter-gene correlation and batch effects compromise FDR control; we demonstrate that Storey’s method maintains strict FDR control while substantially improving detection sensitivity and cross-dataset reproducibility of differentially expressed genes. This workflow provides a standardized, statistically rigorous, and practically applicable analytical pipeline for large-scale transcriptomic studies.
This study addresses the challenge of predicting radiotherapy sensitivity in non-small cell lung cancer (NSCLC) by establishing, for the first time, an integrative transcriptomic (RNA-seq) and proteomic (DIA-MS) analytical framework using SF2—the surviving fraction after 2 Gy irradiation—as the phenotypic endpoint. Leveraging Lasso-based feature selection coupled with support vector regression (SVR), the model was optimized via ten repetitions of five-fold cross-validation to enhance robustness. The integrative model achieved stable predictive performance across both omics layers (R² = 0.461–0.604), outperforming unimodal models. It identified 20 consistently dysregulated cross-omics biomarker genes enriched in DNA damage repair and cellular stress response pathways. This work not only validates the complementary value of multi-omics integration for mechanistic insight and clinical translation but also establishes a generalizable paradigm for radiobiological sensitivity prediction.
Identifying disease-associated genes from gene expression data remains heavily reliant on manual curation and lacks scalability. Method: We propose GenoTEX, the first automated evaluation benchmark for this task—covering data selection, preprocessing, and statistical analysis—and provide expert-annotated code and results. We formalize domain-expert practices as quantifiable LLM-agent tasks and introduce GenoAgent, a self-correcting multi-agent framework integrating LLMs, workflow orchestration, differential expression analysis, GO/KEGG enrichment, and expert-knowledge alignment. Contribution/Results: On GenoTEX, GenoAgent achieves end-to-end automated analysis with significantly reduced human intervention. Error analysis identifies semantic understanding and domain-logic modeling as primary bottlenecks. This work establishes a reproducible, evaluable benchmark and methodological paradigm for biomedical AI agents.
Traditional gene expression analysis relies heavily on manual curation, resulting in low efficiency and poor reproducibility. To address this, we propose the Team of AI Scientists (TAIS), a novel multi-agent framework that pioneers a role-based collaboration paradigm grounded in large language models (LLMs). TAIS decomposes the disease-predictive gene identification pipeline into three modular, schedulable, and verifiable AI agents—Project Manager, Data Engineer, and Domain Expert—enabling end-to-end automation. Our method integrates a multi-agent architecture, a custom gene expression benchmark, prompt-driven task decomposition, and rigorous result verification. Evaluated on our curated benchmark, TAIS achieves human-expert-level accuracy in gene identification while automating 92% of the workflow. This significantly enhances research efficiency, transparency, and reproducibility, overcoming fundamental limitations of monolithic LLM-based analysis.
This work addresses the critical disconnect between tool references in scientific narratives and their implementations in executable bioinformatics workflows, which severely hinders reproducibility and reuse. To bridge this gap, we propose CoPaLink, the first end-to-end framework that jointly identifies tool entities in both scientific text and Nextflow code and links them across modalities using established bioinformatics knowledge bases such as Bioconda and Bioweb. Trained on a manually annotated corpus, our approach achieves F1 scores of 84–89% for tool entity recognition in individual modules and an overall linking accuracy of 66%. By aligning descriptive content with computational implementation, CoPaLink significantly narrows the semantic gap between narrative methods descriptions and executable code, thereby enhancing support for workflow reproducibility.
This work addresses the lack of systematic evaluation benchmarks for AI agents in multi-step bioinformatics workflows, which hinders reliable assessment of their performance and robustness. We propose the first standardized evaluation framework encompassing end-to-end tasks such as RNA-seq analysis and variant calling. The framework integrates structured prompting, an automated LLM-based scoring mechanism, and perturbation tests—including corrupted inputs and decoy files—to holistically evaluate agents on workflow completeness, output correctness, and robustness, while also considering the applicability of open-source models in privacy-sensitive settings. Experimental results demonstrate that leading closed-source agents can reliably execute complex pipelines but exhibit reasoning vulnerabilities under perturbations, whereas open-source models, despite lower task completion rates, offer greater practical utility when data privacy constraints limit access to external systems.
Biomedical machine learning is often compromised by data leakage arising from repeated measurements, study heterogeneity, batch effects, or temporal dependencies, leading to biased model evaluation. This work proposes a leakage-aware resampling workflow that innovatively integrates leakage-safe data partitioning, training-set-only preprocessing, nested hyperparameter tuning, and post-hoc leakage auditing, culminating in an interactive HTML diagnostic report. Built upon R’s S4 class system, the framework supports classification, regression, and survival analysis tasks while ensuring reproducibility and task-specific evaluation rigor. Simulation studies and multi-study transcriptomic case analyses demonstrate that leakage-preventive pipelines substantially alter model performance and downstream conclusions, underscoring their necessity and practical utility in robust biomedical machine learning.
This work addresses the pervasive challenges in bioinformatics tooling—such as fragmentation, complex dependencies, inconsistent documentation, and irreproducible environments—that severely hinder method reuse and adaptation. To overcome these limitations, the authors propose PoSyMed, an open modular platform that integrates biomedical workflows through formalized tool descriptions, containerized execution, a persistent workflow engine, and a conversational interface. Innovatively, a large language model is incorporated as a semantic assistant within a typed, validated, and human-supervised framework to support tool discovery, pipeline construction, and parameter configuration. This design significantly enhances analytical transparency and reproducibility. The platform’s efficacy is demonstrated in representative biomedical use cases, and it has been released as open-source software.