On-Device Vision Training, Deployment, and Inference on a Thumb-Sized Microcontroller

📅 2026-04-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of performing end-to-end on-device training and inference for vision-based machine learning on extremely resource-constrained microcontrollers, where traditional approaches rely heavily on cloud infrastructure. The authors demonstrate a complete pipeline—including data acquisition, training of a two-layer convolutional neural network (CNN), and real-time inference—on an ESP32-S3 XIAO ML Kit featuring only 8 MB of PSRAM and no external dependencies. By leveraging batch-level gradient accumulation, precomputed scaling lookup tables, a three-priority weight loading scheme, and PSRAM-aware memory management, the system achieves a full training cycle in just 9 minutes and inference at 6.3 frames per second on a 64×64 three-class classification task. Implemented in only 1,750 lines of C++ code, the system includes a custom Adam optimizer and CNN compatible with the Arduino IDE, and is released under the MIT license.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionComputer Vision: Learning & Optimization for CVCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
This paper presents a complete, end-to-end on-device vision machine learning pipeline, comprising data acquisition, two-layer CNN training with Adam optimization, and real-time inference, executing entirely on a microcontroller-class device costing $15-40 USD. Unlike cloud-based workflows that require external infrastructure and conceal the computational pipeline from the practitioner, this system implements every step of the core ML lifecycle in approximately 1,750 lines of readable C++ that compiles in under one minute using the Arduino IDE, with no external ML dependencies. Running on the Seeed Studio ESP32-S3 XIAO ML Kit (8 MB PSRAM), the firmware achieves three-class 64x64 image classification in approximately 9 minutes per training run, with real-time inference at 6.3 FPS. Key contributions include: correct batch-level gradient accumulation; pre-computed resize lookup tables for inference; dual-format weight export for SD-free baked-in deployment; a three-tier weight priority system (SD binary > baked-in header > He-initialization) resolved automatically at boot; a single-constant network reconfiguration interface; and PSRAM-aware memory management suited to microcontroller constraints. All source code and reference datasets are released under the MIT License at https://github.com/webmcu-ai/on-device-vision-ai
Problem

Research questions and friction points this paper is trying to address.

On-Device Learning
Microcontroller
Vision AI
Edge Computing
TinyML
Innovation

Methods, ideas, or system contributions that make the work stand out.

on-device training
microcontroller vision
gradient accumulation
PSRAM-aware memory management
embedded ML deployment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jeremy Ellis
High School Robotics Educator, British Columbia, Canada