arXiv · 2609.21666
Samsone: A Family of Open Small Audio Language Models for On-Device Inference
Abstract
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski. 2026-09-18. Samsone: A Family of Open Small Audio Language Models for On-Device Inference. https://arxiv.org/abs/2609.21666
Cite the original work for its findings. Save a collection to share your selection of sources.