arXiv Science⌕ Search

arXiv · 2610.09225

Lyapunov-Inspired LyRIC Activation and GLARE Attention in Chaos-Guided State Space Modeling for EMG-To-Speech (ETS) Synthesis

Abstract

Electromyography-to-Speech (ETS) synthesis is typically a non-linear, chaotic dynamical system. However, no prior work has studied the chaotic behavior of ETS synthesis to date. Yet, prior works strictly rely on standard reconstruction metrics with parameter-heavy transformers that systematically over-smooth natural acoustic dynamics. To close this gap, for the first time, we propose a chaos-inspired Lyapunov-derived activation function (LyRIC) with two novel chaotic loss functions, Lyapunov Exponent Regularization and Multi-Scale Detrended Fluctuation Analysis, to explicitly capture the deterministic chaos of human phonation. In addition, we introduce a compressed novel encoder, GLAME, which synergizes global Mamba state-space modeling with localized GLARE attention. We comprehensively perform frame-level acoustic evaluation in a multilingual and multi-speaker setup using English and Mandarin datasets. The proposed system outperforms the established baseline with a 4.69x increase in objective intelligibility (STOI: 0.61 vs. 0.13) and a 2.08x improvement in spectral reconstruction (LSD: 1.08 vs. 2.25). Importantly, this improvement is achieved with 73.49% fewer parameters (14.34M vs. 54.10M), establishing a new baseline for ETS synthesis. To the best of our knowledge, this is the first work demonstrating that integrating non-linear chaotic physics into neural networks yields superior yet compact inductive biases for real-time ETS synthesis.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sajid Fardin Dipto, Tarikul Islam Tamiti, Luke Baja-Ricketts, David Vergano, Anomadarshi Barua. 2026-10-06. Lyapunov-Inspired LyRIC Activation and GLARE Attention in Chaos-Guided State Space Modeling for EMG-To-Speech (ETS) Synthesis. https://arxiv.org/abs/2610.09225

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Auditing generative audio calls for known-task audio-llm evaluation

Speech and audio LLMs are evaluated by comparing waveform predictions with predictions from an automatic speech recognition (ASR) transcript. For fixed closed-set tasks, this conflates acoustic evidence with the need to invoke a generative audio model. We estimate incremental call value with matched selectors sharing pre-call evidence. Each policy may retain the transcript label, use a local encoder, or invoke a generative model; matched control removes generative actions but preserves pre-call evidence and development selection. On VocalSound, transcript-only accuracy is 0.296, while supervised CLAP and WavLM controls reach 0.850 and 0.854 without calls. Full selector reaches 0.925 at 12.5% calls versus 0.921 for matched No-call selector (difference 0.004; 95% CI [-0.025, 0.033]). Thus, results do not show a call gain after transcript and encoder evidence are available. Relevant quantity is incremental accuracy from allowing calls, not the waveform-transcript gap.

cs.SD↗

Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

Audiobook narration, conversational agents, and audiovisual dubbing require speech that conveys changing emotions and adapts its pacing within a single utterance. But most existing TTS systems typically rely on utterance-level style conditioning, making such fine-grained control difficult to achieve. In light of this, and inspired by the success of post-training in large language models, we propose a unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration. Supervised fine-tuning establishes instruction-conditioned speech generation, while reinforcement learning with group relative policy optimization refines control accuracy using emotion and duration rewards alongside content and speaker preservation objectives. By reusing the pretrained architecture, our approach avoids additional inference-time control modules. Experiments demonstrate significantly improved fine-grained controllability while maintaining speech intelligibility and speaker identity, highlighting post-training as a practical approach to extending existing speech synthesis models.

cs.SD↗

Do Language Models Need Music Supervision? Verifiable Rewards for Multi-Constraint Symbolic Music Generation

Language models now generate symbolic music from text, and research has focused on musicality. However, many applications require a score that meets explicit constraints, which models struggle to satisfy jointly: on MusicConstraintBench, our benchmark of 2,180 items over eight families of programmatically verifiable constraints, Llama-3.1-70B satisfies 0.630 of single-constraint items but only 0.044 of four-constraint ones. As a remedy, we introduce MusicRLVR, which trains a language model with group relative policy optimisation (GRPO) on verifier rewards alone, needing no human annotation, reward model or music-domain supervised fine-tuning. MusicRLVR incorporates (1) a hard validation gate that rejects malformed scores, (2) graded per-family credit that, unlike a binary reward, separates partially correct outputs, and (3) an all-satisfied bonus for meeting every constraint at once. Extensive experiments show that, in under four hours of training, MusicRLVR raises Qwen3-4B-Instruct-2507 from 0.160 to 0.797 on mixed constraints, outperforming Llama-3.1-70B, and generalises to unseen property combinations, out-of-range parameters and more constraints than any training prompt. The recipe transfers to Qwen3-8B, and neither trained model loses significant accuracy on general benchmarks.

cs.SD↗