Samsone: A Family of Open Small Audio Language Models for On-Device Inference

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决隐私保护和低延迟处理需求,通过开发Samsone系列小型音频语言模型并优化其在设备端执行性能,实现了与大模型相当的效果。
📝 Abstract
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone.
Problem

Research questions and friction points this paper is trying to address.

Small Audio Language Models
on-device inference
edge computing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Small Audio Language Models
on-device inference
edge computing
open-source
🔎 Similar Papers
No similar papers found.
P
Piotr Masztalski
Samsung R&D Institute Poland; AGH University of Kraków, Poland
M
Michał K. Grzeszczyk
Samsung R&D Institute Poland
O
Olaf Sikorski
Samsung R&D Institute Poland