🤖 AI Summary
This study addresses the limitations of large language models (LLMs) in spatial-structural logical reasoning by proposing the Fold2Reason framework. This method constructs the FoldingCorpus dataset and, through multimodal post-training coupled with a shared representation decoding mechanism, deeply integrates the discrete topological structures and continuous geometric signals inherent in protein folding. It provides the first empirical evidence that non-linguistic scientific data can systematically enhance the general reasoning capabilities of LLMs. Experimental results demonstrate that the proposed framework improves structural prediction performance by 3.5-fold and yields an average accuracy gain of 3.23% across ten mainstream reasoning benchmarks. Ultimately, this work establishes a novel paradigm for cross-modal knowledge transfer, bridging structural biology and language model reasoning.
📝 Abstract
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.