🤖 AI Summary
This study addresses the limitations of existing scientific agents, which often suffer from excessive parameter counts yet insufficient capabilities, by proposing a 35B sparse model. The core innovation lies in introducing the first "science-aware refinement loop," wherein frontier agents diagnose failure cases and capability gaps to automatically generate targeted training data alongside executable environments. This framework jointly optimizes supervised fine-tuning, expert training, multi-teacher online distillation, and agent reinforcement learning. Consequently, the proposed model substantially improves efficiency across literature research, scientific coding, and multi-step workflows. By leveraging fewer parameters, it achieves competitive performance that surpasses mainstream open-source models, demonstrating that carefully designed sparse architectures combined with iterative, agent-driven refinement can effectively bridge the gap between model scale and scientific reasoning proficiency.
📝 Abstract
We introduce SAIL, an open model with 35B total and 3B active parameters for literature research, scientific coding, and multi-step research workflows. SAIL is developed through a science-aware improvement loop: agents built on frontier AI models analyze its task failures and construct training tasks that address the underlying capability gaps. The diagnosis examines search and evidence selection in literature tasks, scientific assumptions and reasoning in coding, and planning and revision in longer investigations. The agents draw on paper collections and scientific code repositories to build problems, interaction trajectories, and executable tasks with the required environments and tools. We repeat this loop over multiple development cycles and train SAIL through supervised fine-tuning, specialist training, multi-teacher on-policy distillation, and agentic reinforcement learning. SAIL achieves competitive performance across scientific research tasks with substantially fewer parameters than leading open-weight models.