Does Learning Protein Folding Generalize to Broader Reasoning?

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of large language models (LLMs) in spatial-structural logical reasoning by proposing the Fold2Reason framework. This method constructs the FoldingCorpus dataset and, through multimodal post-training coupled with a shared representation decoding mechanism, deeply integrates the discrete topological structures and continuous geometric signals inherent in protein folding. It provides the first empirical evidence that non-linguistic scientific data can systematically enhance the general reasoning capabilities of LLMs. Experimental results demonstrate that the proposed framework improves structural prediction performance by 3.5-fold and yields an average accuracy gain of 3.23% across ten mainstream reasoning benchmarks. Ultimately, this work establishes a novel paradigm for cross-modal knowledge transfer, bridging structural biology and language model reasoning.
📝 Abstract
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.
Problem

Research questions and friction points this paper is trying to address.

protein folding
reasoning generalization
large language models
spatial reasoning
post-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Protein Folding
Post-training
Spatial Reasoning
3D Geometry Decoding
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yong Liu
Shanghai Jiao Tong University
Z
Zhanpeng Shi
Fudan University, Shanghai Innovation Institute
Yizhou Dang
Yizhou Dang
Software College, Northeastern University, China
Recommender Systems
Z
Zhongyue Zhang
Shanghai Jiao Tong University
X
Xiaoliang Shi
Shanghai Jiao Tong University
Z
Zhijian Wei
Shanghai Jiao Tong University
Shuangjia Zheng
Shuangjia Zheng
Shanghai Jiao Tong University
Generative AIDrug DiscoverySynthetic BiologyMulti-Agent System