VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of on-device vision-language models (VLMs) in accuracy, efficiency, and behavioral reliability by proposing a diagnosis-driven post-training paradigm. The approach employs a teacher VLM to conduct stress testing, uncovering failure modes beyond human priors and converting them into supervisory signals for targeted optimization. Additionally, two visual token compression strategies are introduced to balance precision and latency. Experimental results demonstrate that NanoFull achieves an average score of 62.3 across 17 benchmarks—a 7.4-point improvement over the baseline—while maintaining a minimal infinite-loop rate. Furthermore, FlashFull preserves a competitive score of 61.4 while reducing the time-to-first-token on a Pixel 9 device from 138 seconds to 6.1 seconds, yielding a 23-fold speedup.
📝 Abstract
Sub-billion-parameter Vision-Language Models are increasingly viable for on-device deployment, yet compact model size alone does not guarantee usability. On a phone, such a model can still require more than two minutes to produce its first token. On-device usability depends on three axes: accuracy, efficiency, and behavioral reliability; standard benchmarks miss the third, with answers too short to expose doom loops and prompts too benign to probe adversarial safety. We introduce a diagnosis-driven post-training recipe in which a teacher VLM stress-tests the student, uncovers failure modes beyond human priors, and converts them into targeted supervision and preference alignment, supplementing generic data scaling with failure-driven optimization. Coupled with two visual-token policies, the recipe yields two accuracy-efficiency variants with improved behavioral reliability. \textbf{\NanoFull} attains a 62.3 normalized average over 17 benchmarks, the highest among openly released $\sim$0.5B models, +7.4 over its base at identical architecture and token budget, with doom-loop rates at or below the strongest baseline's. \textbf{\FlashFull} retains 61.4 while cutting warm time-to-first-token on a Pixel 9 from 138\,s to 6.1\,s (23$\times$). By jointly addressing all three axes, we move compact VLMs toward practical on-device usability.
Problem

Research questions and friction points this paper is trying to address.

On-device VLM
Behavioral reliability
Efficiency
Doom loops
Sub-billion-parameter models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
On-Device Deployment
Diagnosis-Driven Post-Training
Preference Alignment
Visual Token Policy