Audio--Image Alignment as a Continued-Pretraining Stage in Automatic Speech Recognition
Thousands of languages are spoken worldwide, yet many remain under-resourced for Automatic Speech Recognition (ASR) due to the limited availability of high-quality transcribed speech data. Collecting accurate transcriptions is often costly and labor-intensive, particularly for low-resource languages. In this work, we introduce a representation-alignment stage between large-scale pretraining and supervised ASR fine-tuning, in which image representations extracted from pretrained vision encoders are aligned with audio representations to adapt a pretrained audio encoder using only paired audio--image data, with no transcriptions. On FLEURS, alignment reduces overall WER across 11 Indian languages from 66.42% to 49.77%, a 25.07% relative reduction that is statistically significant in every language. Two controls establish that the gain comes from the image signal itself: a compute-matched continued-pretraining baseline yields no improvement (66.42% vs. 65.99%), and shuffling the audio--image pairs removes the benefit entirely (72.60%). The aligned encoder also improves over the baseline at every fine-tuning budget from 10 to 100 hours on Vaani and LibriSpeech, with the largest gains in the lowest-resource settings. These findings highlight audio--image representation alignment as an effective transcription-free adaptation strategy for ASR.