arXiv ScienceSearch

arXiv subjects

Yunda Chen

Publications and source records attributed to Yunda Chen.

2 recordsLinked to original sources

Local chord corruption is not recognizer replay: structure-matched calibration

Synthetic chord substitutions offer controlled tests of music generation, but their effects can differ from those of a complete recognized chord sequence. We propose structure-matched calibration, which constructs synthetic chord sequences that match the changed positions and harmonic-relation composition of recognizer replay. Paired generation compares both target-response magnitude and output-chord agreement with replay. On 29 of 30 MUSDB18-HQ songs, central four-second tritone corruption produces a larger target response than complete replay in MIDI-SAG. On 24 held-out MoisesDB songs, structure matching reduces target-response distance to replay by 81% for MIDI-SAG and 77% for MusicGen-Chord. Distance decreases on every song in both models with CNN--CRF. Joint matching also reduces output-chord mismatch with replay by 8--17 percentage points relative to temporal or relational matching alone.

cs.SD

Deep Learning for Personalized Binaural Audio Reproduction

Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep learning for this task and organizes them by generation mechanism into two paradigms: explicit personalized filtering and end-to-end rendering. Explicit methods predict personalized head-related transfer functions (HRTFs) from sparse measurements, morphological features, or environmental cues, and then use them in the conventional rendering pipeline. End-to-end methods map source signals directly to binaural signals, aided by other inputs such as visual, textual, or parametric guidance, and they learn personalization within the model. We also summarize the field's main datasets and evaluation metrics to support fair and repeatable comparison. Finally, we conclude with a discussion of key applications enabled by these technologies, current technical limitations, and potential research directions for deep learning-based spatial audio systems.

eess.AS