arXiv Science⌕ Search

arXiv · 2610.08994

Phoneme-Guided Initialization for LLM-based Speech Recognition

Abstract

Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ryo Magoshi, Shinsuke Sakai, Tatsuya Kawahara. 2026-10-06. Phoneme-Guided Initialization for LLM-based Speech Recognition. https://arxiv.org/abs/2610.08994

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR

Modern automatic speech recognition (ASR) systems excel at transcribing lexical content but often omit nonverbal vocalizations (NVs), such as laughter, breaths, coughs, and cries, that carry conversational and affective information. Modeling NVs in ASR is challenging because NV annotations are sparse and highly long-tailed, with frequent categories such as breaths and laughter dominating rarer events such as cries and coughs. We study three data-centric strategies for improving low-resource NV recognition: (1) a two-stage curriculum that first maps all NV events to a generic token and then fine-tunes on target categories; (2) inter-token transfer from high-resource events, such as laughter and breath, to rare events, such as crying; and (3) voice-conversion augmentation with class balancing. Experiments show that shared acoustic structure across vocal events can be exploited to improve rare-category detection while preserving lexical ASR quality.

eess.AS↗

CASM: Context-Aware Semi-Markov Post-Processor for Beat Tracking

In music beat tracking, model-predicted activations must be decoded into a discrete, musically coherent event sequence. Direct peak picking closely follows local evidence but can retain spurious peaks or miss weak beats. Dynamic Bayesian networks (DBNs), a widely used structured post-processor, improve sequence consistency under predefined global tempo, meter, and transition constraints, but their behavior can depend strongly on these settings. We introduce CASM, a context-aware semi-Markov decoder that instead conditions its temporal constraint on local activation evidence. CASM also accounts for ambiguity among competing periodic interpretations, including half- and double-tempo alternatives. Deterministic safeguards prevent implausible outputs and preserve beat-downbeat consistency. Applied to fixed activations from three neural beat trackers (BeatThis, MSCNN, and TCN), CASM improves temporal continuity while preserving event-level F1 across the GTZAN and SMC datasets, without backbone retraining or dataset-specific retuning. Further analysis shows that CASM is less sensitive than the DBN baseline to the composition of the calibration data.

eess.AS↗

VM-ARRAYDPS: Virtual Microphone Augmented Diffusion Posterior Sampling for Unsupervised Blind Speech Separation

Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recently, diffusion-based approaches have emerged as a promising alternative by leveraging powerful generative priors. Among them, ArrayDPS formulates BSS problem as a posterior sampling problem, and utilizes a pretrained speech diffusion model to guide the recovery of clean source signals. A key factor behind its separation capability is the multi-channel consistency (MC) objective, which enforces the estimated source signals to reconstruct the observed microphone mixtures through the estimated acoustic transfer functions. However, the number of microphones in the array is often limited, which constrains the performance of ArrayDPS. To address this issue, we propose VM-ArrayDPS, a novel method that augments the microphone array with virtual microphones with higher-SNR, these microphones can offer extra MC constraints to enhance the separation performance. Experimental results demonstrate that VM-ArrayDPS significantly outperforms ArrayDPS on both 2-speaker and 3-speaker datasets, showcasing the effectiveness of virtual microphone augmentation in improving BSS performance. We also did ablation studies to show the influence of the number of virtual microphones and weight of the MC objective brought by virtual microphones.

eess.AS↗