OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance

πŸ“… 2026-07-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the privacy and latency challenges of existing cloud-based multimodal large language models (MLLMs) for visual assistance, which require uploading first-person images, as well as the computational limitations of wearable devices that hinder local deployment of large models. The authors propose a perception-computation decoupled architecture: an ESP32-based eyewear captures visual data and transmits it via Wi-Fi to a nearby consumer-grade device that runs the MLLM locally and provides spoken feedback, ensuring raw images never leave the user’s possession. Integrating lightweight compression, secure data redaction, and auditable logging, the system achieves median latencies of 993 ms (for thumbnails) and 1,625 ms (for 1280Γ—720 images), with over 93% of interactions completing under two seconds. It supports tasks such as obstacle detection and object querying, presenting the first open-source, local-first real-time visual assistant tailored for visually impaired users.
πŸ“ Abstract
We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary focus on blind and low-vision users. Cloud MLLM assistants offer strong visual understanding, but often require uploading first-person visual data and can suffer multi-second network delays; wearable glasses are ideal for sensing, but cannot host large models under tight compute and power budgets. OpenGlass addresses this gap with a sensing-computing split: an ESP32-based glasses-side unit captures visual context, while a nearby consumer-grade device performs local MLLM inference and local speech output, reducing cloud reliance and keeping raw egocentric visual data on user-controlled devices by default. We evaluate response quality, query-ready-to-audio latency, safety-aware abstention, and auditable logs. Under real ESP32 Wi-Fi capture, OpenGlass reaches 993 ms median user-to-audio latency with resized payloads and 1625 ms with raw 1280 x 720 payloads; 97.5% and 93.3% of trials fall below 2 s, respectively. OpenGlass is a user-initiated visual-assistance reference platform for obstacle/hazard awareness, sign/object queries, and image-quality self-checking, rather than a certified navigation aid. We release source code, hardware instructions, prompts, evaluation data, and logs.
Problem

Research questions and friction points this paper is trying to address.

visual assistance
privacy
latency
multimodal LLM
wearable computing
Innovation

Methods, ideas, or system contributions that make the work stand out.

sensing-computing split
local MLLM
privacy-preserving
low-latency visual assistance
wearable AI
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.