Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

๐Ÿ“… 2026-06-10
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
AI-generated biomedical analyses often suffer from omissions of critical steps, methodological misapplications, or overinterpretation, undermining their reliability. This study presents the first systematic evaluation of skill-augmented large language models in a real-world non-small cell lung cancer transcriptomic biomarker discovery task. Using the OpenClaw framework, six large language models were equipped with autonomously invocable biomedical research skill modules, and their outputs were rigorously compared against native AI responses through a multi-reviewer double-blind assessment protocol. Results indicated a directional improvement in expert-rated quality for skill-augmented reports (mean score 5.50 vs. 5.11), though the difference did not reach statistical significance as assessed by bootstrap confidence intervals and Welchโ€™s t-test, suggesting the need for larger-scale validation. This work pioneers the investigation of domain-specific skill invocation as a means to enhance the scientific rigor of AI-assisted research.
๐Ÿ“ Abstract
Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical steps, misuse methods, or overstate conclusions. We evaluated whether autonomous access to a medical research skill package was associated with higher-quality AI-generated transcriptomic research-analysis outputs compared with native AI without skills. Methods. We conducted an exploratory multi-model human evaluation using a non-small cell lung cancer immunotherapy biomarker task. Six model backbones were tested. The evaluation included 21 anonymized outputs: 9 native-AI outputs and 12 skill-augmented outputs generated through an AI agent implementation represented by OpenClaw. Four non-expert biomedical reviewers and two blinded experts evaluated each output, with two ratings from each reviewer type. The primary outcome was expert-rated overall quality. Results. Skill-augmented outputs showed directionally higher expert overall quality than native-AI outputs (mean 5.50 vs 5.11; difference=0.39; bootstrap 95\% CI, -0.04 to 0.90; Welch p=0.156). Non-expert reviewer quality showed the same direction (mean 4.72 vs 4.47; difference=0.26; bootstrap 95\% CI, -0.25 to 0.80; Welch p=0.373). Expert agreement was limited (single-rating ICC=-0.15), and model-specific effects were descriptive and heterogeneous. Conclusions. Autonomous skill access showed a directional quality signal in this exploratory sample, but the signal was smaller than expert-rating noise and should not be interpreted as confirmatory evidence. The findings primarily motivate larger evaluations of skill-augmented AI agents with stronger reliability controls, platform replication, and biological-validity assessment.
Problem

Research questions and friction points this paper is trying to address.

AI agents
medical research analysis
transcriptomic biomarkers
large language models
output quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

skill-augmented AI agents
medical research analysis
transcriptomic biomarker
multi-model human evaluation
OpenClaw
Q
Qianyu Yao
AIPOCH PTE. LTD., Singapore
F
Fei Sun
AIPOCH PTE. LTD., Singapore
B
Bocheng Huang
AIPOCH PTE. LTD., Singapore
W
Wei Chen
AIPOCH PTE. LTD., Singapore
J
Jiarui Jiang
AIPOCH PTE. LTD., Singapore
S
Shu Quan
AIPOCH PTE. LTD., Singapore
Y
Yifei Chen
AIPOCH PTE. LTD., Singapore
Wenjie Xu
Wenjie Xu
Phd Student, Wuhan University
Knowledge GraphNLP
B
Bo Li
AIPOCH PTE. LTD., Singapore
L
Liping Su
AIPOCH PTE. LTD., Singapore
R
Ruoqiong Wu
AIPOCH PTE. LTD., Singapore
H
Huhai Hong
AIPOCH PTE. LTD., Singapore
H
Huimei Wang
AIPOCH PTE. LTD., Singapore