arXiv · 2610.03058
Unsupervised Instantaneous Phase and Frequency Tracking by Inverse Voice Synthesis
Abstract
Knowledge-driven neural vocoders struggle to learn reliable fundamental frequency end-to-end, because spectral objectives provide weak supervision of periodic structure and lack phase information. We address this with a source-filter model whose alias-free additive source makes the instantaneous phase of the glottal cycle explicit; differentiating it yields the instantaneous frequency, and thus $F_0$, without an external tracker. Waveform error supervises only the deterministic harmonic path, while a spectral loss covers the full signal. On M4Singer and LM-SSD, the reconstruction is phase-aligned, reaching a signal-to-reconstruction-error ratio of 8.1 dB, while neural baselines remain negative. However, GOLF, given an external $F_0$, still reaches lower spectral distortion. On LM-SSD, the recovered $F_0$ attains the highest overall accuracy of any method tested, including supervised neural pitch trackers applied off the shelf, and the glottal closure instants come within 0.53 points of REAPER's identification rate, without any $F_0$ label.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chin-Yun Yu, György Fazekas. 2026-10-02. Unsupervised Instantaneous Phase and Frequency Tracking by Inverse Voice Synthesis. https://arxiv.org/abs/2610.03058
Cite the original work for its findings. Save a collection to share your selection of sources.