SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the memory, communication, and operator efficiency bottlenecks encountered when performing full-parameter post-training of trillion-parameter Mixture-of-Experts (MoE) models on non-GPU hardware. Focusing on the Ascend NPU SuperPOD platform, we propose an end-to-end hierarchical optimization framework encompassing model parallelism strategies, compute-communication co-scheduling, and customized low-level operators. By integrating solver-guided synthetic data to construct a domain-specific instruction tuning dataset, we achieve, for the first time, efficient full-parameter post-training of the DeepSeek-V4 model family on Ascend super-nodes. The system attains a Model FLOPs Utilization (MFU) of 34.22%, yielding a 2.93× speedup over open-source baselines. On operations research tasks, it achieves a zero-shot Pass@1 accuracy of 71.81%, significantly outperforming both GPT-5.4-Mini and the original DeepSeek-V4-Flash model.
📝 Abstract
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
Problem

Research questions and friction points this paper is trying to address.

post-training
trillion-parameter MoE models
distributed training
memory pressure
communication overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

full-parameter post-training
MoE models
Ascend SuperPOD
computation-communication orchestration
solver-grounded SFT
Dongfang Li
Dongfang Li
Harbin Institute of Technology, Shenzhen
Natural Language ProcessingLarge Language Models
Xiaodong Luo
Xiaodong Luo
sctu.edu.cn
Image generation Computer Vision
Ruoyu Sun
Ruoyu Sun
Chinese University of Hong Kong (Shenzhen), Shenzhen Institue of Big Data
Mathematical optimizationNeural networksMachine Learning
Xuhui Chen
Xuhui Chen
San Francisco State University
Computer Science
L
Linyuan Qiu
AI Training Platform Team, Shenzhen Loop Area Institute
Jian Meng
Jian Meng
Cornell Tech, Cornell University
Hardware-Efficient Machine LearningNeuromorphic ComputingDNN model compressionDNN hardware
Z
Zhengxuan Lu
AI Training Platform Team, Shenzhen Loop Area Institute
Yiting Wang
Yiting Wang
Graduate Student, University of Maryland
AI for EDAHardware Security
Y
Yucheng Xie
AI Training Platform Team, Shenzhen Loop Area Institute
T
Tao Guo
AI Training Platform Team, Shenzhen Loop Area Institute
T
Tianxiang Fang
AI Training Platform Team, Shenzhen Loop Area Institute
J
Jing Li
AI Training Platform Team, Shenzhen Loop Area Institute
S
Sihang Chen
AI Training Platform Team, Shenzhen Loop Area Institute
S
Shihao Hong
AI Training Platform Team, Shenzhen Loop Area Institute
C
Chang Liu
AI Training Platform Team, Shenzhen Loop Area Institute
W
Weihua Dai
AI Training Platform Team, Shenzhen Loop Area Institute
Z
Zirong Zeng
AI Training Platform Team, Shenzhen Loop Area Institute
Z
Ziwei Zhu
AI Training Platform Team, Shenzhen Loop Area Institute
Zhuohan Wang
Zhuohan Wang
Kings College London
Quantitative FinanceGenerative ModelGame Theory
Zhengjun Yue
Zhengjun Yue
Assistant Professor, TU Delft
Speech Technology for healthcarePathological speech recogntiion
Igor Vasilyev
Igor Vasilyev
Institute for System Dynamics and Control Theory of Siberian Branch of Russian Academy of Sciences
Operations ResearchCombinatorial OptimizationInteger Programming
M
Min Liu
AI Training Platform Team, Shenzhen Loop Area Institute
W
Weijian Sun
AI Training Platform Team, Shenzhen Loop Area Institute
Xin Chen
Xin Chen
深圳市腾讯计算机系统有限公司
machine learningdeep learninggraph neural networksrecommendation
Y
Yingmeng Gao
AI Training Platform Team, Shenzhen Loop Area Institute