🤖 AI Summary
This work addresses key challenges in large language models—spanning multi-domain competence, long-context processing, reasoning stability, and expert specialization—by introducing a sparse Mixture-of-Experts decoder with 314 billion total parameters and 13.2 billion activated per token. The architecture features several core innovations: Grouped Differential Latent Attention, Expert-Specific PolyNorm, multi-token prediction, and an enhanced hyperconnected structure. Combined with MXFP8 quantization, window-aware context parallelism, and highly optimized fused kernels, the model enables efficient training and inference at a 256K-token context length. Further refined through multi-teacher distillation and reinforcement learning fine-tuning, the model achieves state-of-the-art performance among open-source models in mathematical reasoning, scientific knowledge, long-horizon agent tasks, and low-hallucination benchmarks.
📝 Abstract
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.