IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low efficiency of real-time large model inference on resource-constrained edge devices by proposing a lightweight, edge-native language model. Methodologically, we design a hybrid attention architecture combined with an X-MTP multi-token prediction mechanism to eliminate KV cache replay. To streamline deployment, the model adopts an Instruct-Only paradigm, replaces RMSNorm with dynamic Tanh, and incorporates quantization-friendly components, while multi-domain online distillation preserves model capabilities. Experimental results demonstrate that the proposed approach achieves a 1.48× decoding speedup, delivers performance comparable to larger-scale models with more concise responses, and significantly enhances local inference efficiency on edge devices.
📝 Abstract
We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.
Problem

Research questions and friction points this paper is trying to address.

Edge AI
Compact Language Models
On-device Inference
Embodied Intelligence
Low Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-token prediction
shared-KV cache
on-policy distillation
edge inference
Dynamic Tanh
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Changdi Yang
Changdi Yang
PhD candidate, Northeastern University, Snap Inc.
Efficient Deep Learning
F
Fengquan Jiao
Robotics Foundation Model Team, Xpeng Inc.
H
Haochih Lin
Robotics Foundation Model Team, Xpeng Inc.
Haoran Yang
Haoran Yang
Central South University
Graph Neural NetworksData MiningRecommendation Systems
J
Jing Xiao
Robotics Foundation Model Team, Xpeng Inc.
L
Liangyu Huo
Robotics Foundation Model Team, Xpeng Inc.
S
Suxin Lu
Robotics Foundation Model Team, Xpeng Inc.
T
Tiance Chen
Robotics Foundation Model Team, Xpeng Inc.
W
Wei Liu
Robotics Foundation Model Team, Xpeng Inc.
Y
Yinggan Xu
Robotics Foundation Model Team, Xpeng Inc.
Y
Yunxiang Lu
Robotics Foundation Model Team, Xpeng Inc.
Z
Zai Zheng
Robotics Foundation Model Team, Xpeng Inc.
Z
Zhirui Xie
Robotics Foundation Model Team, Xpeng Inc.
Z
Zhongyang Che
Robotics Foundation Model Team, Xpeng Inc.
Z
Ziyan Tang
Robotics Foundation Model Team, Xpeng Inc.
Z
Zuoxiang Zhao
Robotics Foundation Model Team, Xpeng Inc.
Jian Yao
Jian Yao
Wuhan University
Computer VisionAI3DRoboticsSLAM