StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

📅 2026-07-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of deploying accurate and efficient UI-centric vision-language models (UI-VLMs) on resource-constrained mobile devices, where high multi-task performance must coexist with stringent hardware limitations. The authors propose a holistic architecture-training-deployment co-design framework to build a 0.9B-parameter on-device UI-VLM, featuring a novel UI-aware layered visual encoder (ULVE) and a progressive dimensionality projection (PDP) connector. A five-stage curriculum training strategy enables effective multi-task optimization, while a module-differentiated quantization pipeline—combining PTQ-to-QAT with mixed precision (W4A16+KV8)—keeps accuracy degradation below 1%. The resulting model outperforms larger 2B–2.3B counterparts on ScreenQA (88.76 F1) and OCRBench v2 (57.25), and matches state-of-the-art general-purpose VLMs on RefCOCO (92.0%) and OCRBench v1 (831). When quantized, it runs stably on Snapdragon 8 Gen5 with a TTFT of ~0.84 seconds, decoding speed of 98 tokens/second, and peak memory usage of 1.4 GB.
📝 Abstract
Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen understanding, visual question answering, and element grounding; on the other is the strict compute, memory, and power budget of mobile chips. Existing work either trades one for the other, or stops at simulation without real-device validation. We present StepX-Edge, a 0.9B-parameter on-device UI vision-language model that resolves this tension through three-layer co-design of architecture, training, and deployment. Architecturally, UI-aware Layered Visual Encoding (ULVE) and a Progressive Dimensionality Projection (PDP) connector target the extreme aspect ratios and fine-grained perception of screens, while standard full attention throughout ensures native compatibility with mainstream mobile NPU operators. For training, the five-stage StepX-Curriculum framework is designed around our observation of mutual-promotion effects among UI subtasks, so that all four capabilities grow synergistically under a tight parameter budget rather than interfering. For deployment, a module-wise differentiated two-stage PTQ-to-QAT quantization scheme keeps the post-quantization accuracy loss within 1%. StepX-Edge achieves the strongest overall UI understanding among <=1B models, surpassing all 2B-2.3B baselines on ScreenQA (88.76 F1) and Chinese OCRBench v2 (57.25), and matching 1.3B-2.3B general VLMs on RefCOCO (92.0%) and OCRBench v1 (831) with far fewer parameters. After W4A16+KV8 quantization, the model runs stably on Snapdragon 8 Gen5 devices with ~0.84 s TTFT, 98 tok/s decode, and 1.4 GB peak memory. We will open-source the training data, the full training recipe, and the quantization deployment pipeline.
Problem

Research questions and friction points this paper is trying to address.

on-device deployment
vision-language model
UI understanding
accuracy-efficiency trade-off
mobile constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

on-device vision-language model
architecture-training-deployment co-design
UI-aware visual encoding
curriculum learning for UI tasks
differentiated quantization
💼 Related Jobs
No related jobs found.
Y
Yin Wang
Haotian Hu
Haotian Hu
浙江大学
自动驾驶
J
Jineng Han
W
Wentao Qiu
Z
Zhenhua Ge
L
Liujian Tang
F
Fanyi Wang