π€ AI Summary
This work addresses the privacy and latency challenges of existing cloud-based multimodal large language models (MLLMs) for visual assistance, which require uploading first-person images, as well as the computational limitations of wearable devices that hinder local deployment of large models. The authors propose a perception-computation decoupled architecture: an ESP32-based eyewear captures visual data and transmits it via Wi-Fi to a nearby consumer-grade device that runs the MLLM locally and provides spoken feedback, ensuring raw images never leave the userβs possession. Integrating lightweight compression, secure data redaction, and auditable logging, the system achieves median latencies of 993 ms (for thumbnails) and 1,625 ms (for 1280Γ720 images), with over 93% of interactions completing under two seconds. It supports tasks such as obstacle detection and object querying, presenting the first open-source, local-first real-time visual assistant tailored for visually impaired users.
π Abstract
We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary focus on blind and low-vision users. Cloud MLLM assistants offer strong visual understanding, but often require uploading first-person visual data and can suffer multi-second network delays; wearable glasses are ideal for sensing, but cannot host large models under tight compute and power budgets. OpenGlass addresses this gap with a sensing-computing split: an ESP32-based glasses-side unit captures visual context, while a nearby consumer-grade device performs local MLLM inference and local speech output, reducing cloud reliance and keeping raw egocentric visual data on user-controlled devices by default. We evaluate response quality, query-ready-to-audio latency, safety-aware abstention, and auditable logs. Under real ESP32 Wi-Fi capture, OpenGlass reaches 993 ms median user-to-audio latency with resized payloads and 1625 ms with raw 1280 x 720 payloads; 97.5% and 93.3% of trials fall below 2 s, respectively. OpenGlass is a user-initiated visual-assistance reference platform for obstacle/hazard awareness, sign/object queries, and image-quality self-checking, rather than a certified navigation aid. We release source code, hardware instructions, prompts, evaluation data, and logs.