arXiv Science⌕ Search

arXiv · 2610.09211

BanglaBox: A Phonetically-Balanced Corpus and Data-Efficient Foundation-Model Adaptation for Bangla Text-to-Speech with Zero-Shot Voice Cloning

Abstract

We present a recipe for adapting English-pretrained autoregressive TTS foundation models to underrepresented languages, demonstrated on Bangladeshi Bangla. Existing Bangla TTS corpora are small and single-speaker, and to our knowledge no open zero-shot voice-cloning system is available for the Bangladeshi register. We contribute a phonetically- and gender-balanced two-tier Bangladeshi Bangla corpus balanced via a tiered Jensen-Shannon divergence objective over conjunct clusters (juktakkhor), together with three fine-tuning changes: a merge-consistent tokenizer extension, Bangla text normalization, and a prompt-masked dual-loss objective. These changes preserve zero-shot cloning across the language switch. Our BanglaEval protocol applies Wilcoxon signed-rank tests with Bonferroni correction over native-speaker ratings. BanglaBox attains near-natural Naturalness, outperforms commercial and open-source baselines on Naturalness, Speaker Similarity, and Clarity, and reaches speaker similarity comparable to prior few-shot results while using approximately 7x less Bangla fine-tuning audio. We further validate the recipe beyond our own test split using a seven-category stress set of difficult real-world text, the public BnTTS evaluation benchmarks, and naturally occurring Bangla that no language model wrote. All artifacts, including the corpus, weights, code, and complete evaluation materials, are released publicly and unconditionally.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Emtiaz Uddin Ahmed, Araf Mahmud, Sajib Hossain, Tarikul Islam Tamiti, Sajid Fardin Dipto, Anomadarshi Barua. 2026-10-06. BanglaBox: A Phonetically-Balanced Corpus and Data-Efficient Foundation-Model Adaptation for Bangla Text-to-Speech with Zero-Shot Voice Cloning. https://arxiv.org/abs/2610.09211

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency representations. Existing models commonly use harmonic constant-Q transform (HCQT)-based representations to provide frequency-adaptive resolution, at the cost of expensive feature extraction when training mixtures are generated on the fly. We revisit this design and compare HCQT with a linear short-time Fourier transform (STFT), whose frequency bins are directly provided as model inputs. Despite its fixed frequency resolution and the absence of a pitch-aligned input grid, the linear STFT outperforms HCQT while substantially reducing feature-extraction cost. Further analysis shows that a longer analysis window or broader spectral coverage provides no additional improvement, while restricting the input to the predicted pitch range reduces the advantage of the linear STFT. These results suggest that finer frequency resolution does not necessarily improve vocal-ensemble MPE, and that shorter analysis windows can be more effective for time-varying vocal pitches.

cs.SD↗

SEA-LM: Egocentric Spatial Audio Understanding for Wearable Microphone Arrays

Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentanglement of overlapping sound sources. To address this, we present SEA-LM, a Spatial Audio Understanding model. First, we introduce FOACODER, a layout-flexible spatial audio encoder trained on source localization and ego-centric voice activity detection objectives to encode First Order Ambisonics derived from variable-count, variable-position smart-glasses arrays via beamforming. We train a Multimodal Large Language Model (MLLM) to understand these spatial audio embeddings through a two-stage curriculum spanning six tasks, including sound localization and spatially selective transcription in settings with multiple speakers and overlapping sounds. To prevent the transcription outputs from dominating the next token prediction loss and overwhelming the direction predictions, we introduce a Spatio-temporal Weighted Cross-Entropy Loss. On our evaluation set, SEA-LM achieves lower azimuth and elevation MAE, higher temporal IoU, lower external-source hallucination and missing-source rates, and lower WER on most transcription tasks than compared baselines, while remaining robust across 1,211 smart-glasses array configurations with 4 to 9 microphones.

cs.SD↗

ImpactMat: Continuous Material Estimation for Inverse Impact Sound Rendering

Impact sound rendering synthesizes the sound produced when a 3D object is struck, but practical renderers often rely on fixed material presets such as wood, plastic, or steel. These presets limit the range of impact sounds a renderer can express, while manually adjusting the underlying material parameters remains difficult without expertise in material acoustics. We therefore study inverse impact sound rendering: predicting material parameters from a reference impact sound so that a simulator can recreate a similar material response. To support this task, we introduce ImpactMat, a dataset and benchmark of single and blended material impact sounds paired with ground-truth material parameters. We further propose a feed-forward model that predicts these parameters from one or more recordings, using blended materials to learn smooth transitions between material types. Experiments show that our method outperforms competitive baselines and enables re-rendering from real recordings without manual parameter tuning. The project page is available https://material-from-impact.github.io/material-from-impact/.

cs.SD↗