Small yet Assistive: Spatially-Aware Post-Training for Low Vision

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited spatial awareness of small vision-language models, which hinders their ability to assist visually impaired users in safe navigation. To this end, we propose Smol-VL-BLV, built upon a 500M-parameter decoder-only Transformer architecture. The method enhances spatial reasoning through teacher-student knowledge distillation and Group Relative Policy Optimization (GRPO) reinforcement learning. A composite reward mechanism tailored for visually impaired scenarios is designed, while mixed-precision quantization and lightweight fine-tuning are employed to mitigate catastrophic forgetting, enabling low-latency, offline on-device deployment. Experimental results demonstrate that the proposed model improves spatial understanding scores by 19.3% and OCR benchmark performance by 101.5%, offering an efficient and reliable lightweight visual assistance solution for visually impaired populations.
📝 Abstract
An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs offer competitive latency but lack spatial detail, directional cues, and hazard awareness for navigational assistance. We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric distances, and hazard detection. Because multi-stage post-training can induce catastrophic forgetting, we add a lightweight finetuning stage after the last stage GRPO finetuning to recover general descriptive quality while preserving BLV-specific spatial grounding. Our best model substantially outperforms the baseline across various benchmarks, including tasks: VQA, BLV captioning, OCR, and latency. Compared with the baseline for relative improvement, it improves the Spatial score gain of 19.3%, and the Social score gain of 14.8%. It also increases OCR-Bench by 101.5%, and raises TextVQA accuracy by 44.2%. These results show that BLV-focused post-training improves both accessibility-specific spatial grounding and general visual-text reasoning. Deployed on a mid-range Android smartphone via Mixed-Precision Quantization, the model remains approx. 450 MB and runs entirely on-device, offline and without network dependency, generating descriptions with latency dependent on host hardware capabilities. Our model, dataset, and code is publicly released at https://smol-vl-blv.github.io/Smol-VL-BLV-website/
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Blind and Low-Vision
Navigation Assistance
Spatial Awareness
On-device Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Model
Group Relative Policy Optimization
Knowledge Distillation
Catastrophic Forgetting
On-device Deployment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rishabh Choudhary
Indian Institute of Technology Mandi
S
Shreyansh Raj
Indian Institute of Technology Mandi
U
Umesh Goyal
Indian Institute of Technology Mandi
S
Shubh Kashyap
Indian Institute of Technology Mandi
S
Shrestha Kumar
Indian Institute of Technology Mandi
S
Sushovan Jena
Indian Institute of Technology Mandi
Komal Kumar
Komal Kumar
PhD MBZUAI
computer visionunified modelingmulti-agentReinforce LLMself supervised learning
Hisham Cholakkal
Hisham Cholakkal
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Computer VisionLarge Multimodal ModelsLLMHealthcare Foundation ModelConversational Assistant
Aditya Nigam
Aditya Nigam
Associate Professor in SCEE at Indian Institute of Technology Mandi
Deep LearningBiometricsPalmprint and Knuckleprint RecognitionIris RecognitionMulti Biometrics