Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

πŸ“… 2026-07-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of enabling small-scale models to achieve strong agentic capabilities under strict parameter constraints by proposing a compact, general-purpose agent model with only 3 billion non-embedding parameters. The model is trained from scratch on 28 trillion tokens of diverse real and synthetic agent trajectories and employs an innovative Looped Transformer architecture that reuses layer stacks to increase effective capacity without expanding parameter count. Enhanced through hybrid-mode RLHF, length-controlled reinforcement learning, and a dual reward mechanism balancing process and outcome, the model demonstrates significantly improved performance in multi-task reasoning, code generation, and tool use. It outperforms larger counterparts such as Qwen3.5-9B and Gemma4-12B across multiple agent benchmarks, excelling in mathematical, programming, scientific reasoning, and alignment tasks, thereby serving as an efficient and lightweight local personal assistant.
πŸ“ Abstract
We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations show that Nanbeige4.2-3B outperforms larger models, including Qwen3.5-9B and Gemma4-12B, across diverse agentic benchmarks while remaining competitive on reasoning and alignment tasks. Performance with OpenClaw further supports its use as a compact local personal assistant.
Problem

Research questions and friction points this paper is trying to address.

agentic model
compact model
tool-use tasks
reasoning capabilities
personal assistant
Innovation

Methods, ideas, or system contributions that make the work stand out.

Looped Transformer
agentic scaffolds
mixed-mode RLHF
length-controlled reasoning RL
compact agentic model
C
Chen Yang
Nanbeige LLM Lab, Boss Zhipin
Chengrui Huang
Chengrui Huang
University of Electronic Science and Technology of China
Natural Language ProcessingTool Learning
F
Fufeng Lan
Nanbeige LLM Lab, Boss Zhipin
H
Hanhui Chen
Nanbeige LLM Lab, Boss Zhipin
H
Hao Zhou
Nanbeige LLM Lab, Boss Zhipin
Huatong Song
Huatong Song
GSAI, Renmin University of China
Large Language Models
Jiaqi Cao
Jiaqi Cao
Shanghai Jiao Tong University
Natural Language ProcessingLong-term Memory
J
Jiaying Zhu
Nanbeige LLM Lab, Boss Zhipin
J
Jinlin Niu
Nanbeige LLM Lab, Boss Zhipin
K
Kai Wang
Nanbeige LLM Lab, Boss Zhipin
Lisheng Huang
Lisheng Huang
δΈ­ε›½δΊΊζ°‘ε€§ε­¦ζœ¬η§‘η”Ÿ
θ‡ͺη„Άθ―­θ¨€ε€„η†γ€ε€§θ―­θ¨€ζ¨‘εž‹
Q
Qiliang Liang
Nanbeige LLM Lab, Boss Zhipin
R
Ran Le
Nanbeige LLM Lab, Boss Zhipin
R
Ruixiang Feng
Nanbeige LLM Lab, Boss Zhipin
S
Shuang Sun
Nanbeige LLM Lab, Boss Zhipin
T
Tao Gu
Nanbeige LLM Lab, Boss Zhipin
T
Tao Zhang
Nanbeige LLM Lab, Boss Zhipin
T
Tianyu Luo
Nanbeige LLM Lab, Boss Zhipin
Y
Yang Song
Nanbeige LLM Lab, Boss Zhipin
Yun Xing
Yun Xing
School of Computer Science and Engineering, Nanyang Technological University
Computer Vision
Y
Yuntao Wen
Nanbeige LLM Lab, Boss Zhipin
Ziyao Xu
Ziyao Xu
Peking University
Z
Zongchao Chen
Nanbeige LLM Lab, Boss Zhipin