Ternarization of Vision Language Models for use on edge devices

📅 2025-04-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To enable efficient deployment of vision-language models (VLMs) on resource-constrained edge devices, this work proposes a direct ternarization compression paradigm for pre-trained VLMs—bypassing costly from-scratch training. Methodologically: (1) we introduce a k-means–driven weight clustering initialization strategy to accelerate ternarization convergence and improve accuracy; (2) we implement high-precision ternary matrix multiplication and develop custom TensorFlow Lite–native ternary operators. Our key contributions include the first end-to-end ternarization framework for pre-trained VLMs; achieving ~13× memory reduction over full-precision counterparts while maintaining near-original perplexity (increase < 0.5%) and outperforming binary models in token generation latency; and establishing the state-of-the-art trade-off between accuracy and efficiency among ternary VLM compression methods.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionComputer Vision: Large Vision ModelsNatural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
We propose a process to compress a pre-trained Vision Language Model into a ternary version of itself instead of training a ternary model from scratch. A new initialization scheme from pre-trained weights based on the k-means algorithm is proposed to reduce the ternarization time. We implement different custom operators for executing the ternary model on the TensorFlow Lite Engine. We compare the original model with its ternary and binary versions in terms of memory consumption, inference speed and perplexity. We find that the ternary model using our custom ternary matrix multiplication operator provides a good compromise in term of memory usage and perplexity, while having the fastest token generation speed.
Problem

Research questions and friction points this paper is trying to address.

Compress Vision Language Models for edge devices
Reduce ternarization time using k-means initialization
Optimize memory, speed, perplexity in ternary models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compress Vision Language Model into ternary version
K-means based initialization for faster ternarization
Custom ternary operators for TensorFlow Lite Engine
💼 Related Jobs
No related jobs found.
Ben Crulis
Ben Crulis
Unknown affiliation
AIdeep learningcomputer vision
C
Cyril de Runz
LIFAT, University of Tours
B
Barthélémy Serres
LIFAT, University of Tours
G
Gilles Venturini
LIFAT, University of Tours