🤖 AI Summary
Current scientific AI models struggle to unify heterogeneous data, scientific laws, and expert knowledge within a single framework. This work proposes the first unified multimodal scientific reasoning model that maps natural language instructions and diverse scientific objects—including CIF files, SMILES strings, protein sequences, spectra, and scientific images—into a shared representation space. By integrating scientific priors during training and employing task-specific decoders, the model supports a wide range of scientific tasks. It achieves, for the first time, a unified approach to scientific understanding, prediction, and generation, outperforming GPT-5.5 and Gemini-3.1-Pro across more than 60 scientific benchmarks and matching or exceeding the performance of specialized models on multiple tasks.
📝 Abstract
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilities into a single, coherent scientific reasoning model. The architecture of S1-Omni is built upon three core components: unified representation of scientific data, natural-world knowledge alignment, and decoding for domain-specific tasks. First, S1-Omni maps natural-language instructions and scientific objects, including CIF, SMILES, protein sequences, spectra, and scientific images, into a shared representation space. Second, it incorporates scientific laws and expert knowledge into data construction and training, enabling the model to reason from scientific evidence. Third, it performs task-specific decoding to support a broad range of applications, including property prediction, spectrum-to-molecular generation, protein site and structure prediction, and scientific image generation and editing. S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks. It outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or surpasses domain-specific models on several benchmarks. Overall, S1-Omni provides a practical path toward unified scientific modeling.