π€ AI Summary
This study addresses the challenges of inefficient knowledge transfer and capability trade-offs during the post-training of multimodal large language models. Building upon Qwen, we construct dual-scale 27B and 9B models and propose an Interleaved Distillation Reinforcement Learning (IDRL) paradigm alongside an automated reasoning optimization framework, PAH. Through the dynamic alternation of multimodal supervised fine-tuning, online policy distillation, and agent trajectory reinforcement learning, our approach achieves synergistic optimization of reasoning and agentic capabilities. Experimental results demonstrate that both models significantly outperform their base counterparts, exhibiting superior performance on long-horizon planning and multimodal search tasks. Furthermore, generalization is enhanced without requiring parameter updates.
π Abstract
We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.