When Do Biological Reasoning Models Use Their Biological Inputs?

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether biological reasoning models genuinely leverage biological inputs such as DNA and proteins for inference or merely rely on textual information. To address this, we propose a systematic framework combining causal intervention and representation probing—incorporating input perturbation, evidence-conflict construction, linear probing, and reasoning trajectory analysis—to evaluate six models' actual dependence on biological data. Our findings reveal that current post-training strategies fail to ensure effective contributions from base model representations; notably, Evo2 and ESM3 exhibit minimal contribution within BioReason, with models predominantly following textual cues. This work demonstrates that accuracy improvements do not necessarily indicate enhanced utilization of biological inputs, thereby establishing a new paradigm for the mechanistic evaluation of multimodal biological models.
📝 Abstract
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
Problem

Research questions and friction points this paper is trying to address.

Biological reasoning models
Foundation model representations
Post-training
Large language models
Benchmark accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Biological Reasoning Models
Evidence Conflicts
Linear Probes
Post-training Strategies
Foundation Model Representations
🔎 Similar Papers
No similar papers found.