arXiv · 2506.11089
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
Abstract
Automatic speech recognition (ASR) models rely on high-quality transcribed data for effective training. Generating pseudo-labels for large unlabeled audio datasets often relies on complex pipelines that combine multiple ASR outputs through multi-stage processing, leading to error propagation, information loss and disjoint optimization. We propose a unified multi-ASR prompt-driven framework using postprocessing by either textual or speech-based large language models (LLMs), replacing voting or other arbitration logic for reconciling the ensemble outputs. We perform a comparative study of multiple architectures with and without LLMs, showing significant improvements in transcription accuracy compared to traditional methods. Furthermore, we use the pseudo-labels generated by the various approaches to train semi-supervised ASR models for different datasets, again showing improved performance with textual and speechLLM transcriptions compared to baselines.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jeena Prakash, Blessingh Kumar, Kadri Hacioglu, Bidisha Sharma, Sindhuja Gopalan, Malolan Chetlur, Shankar Venkatesan, Andreas Stolcke. 2025-06-05. Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM. https://doi.org/10.21437/interspeech.2025-1707
Cite the original work for its findings. Save a collection to share your selection of sources.