KeyBound: Keyed and Host-Bound Learned Audio Watermarking for Speech Provenance
Audio watermarking is a proactive route to attributing synthetic speech to its source. Learned audio watermarks are typically judged by payload recovery after a fixed catalog of signal distortions such as noise, compression, filtering, and resampling. That test is necessary but not sufficient for provenance. A mark offered as evidence of origin should not be readable by an unauthorized party, should not be transferable to unrelated audio, and should not vanish when the recording is re-synthesized by a modern generative model. We present KeyBound, a learned audio watermark that restores the two ingredients classical watermarking supplied and learned schemes set aside, a secret key and a host-aware carrier. KeyBound masks the payload with a secret key and embeds the masked bits through a carrier modulated by a frozen spectral representation of the host, so the key governs payload access while the host-conditioned carrier resists direct transplantation. A key-independent presence head lets any party detect a mark, whereas only a key holder reads its attribution, and under the single-sample uniformity assumption a wrong-key decode clears our verification rule with probability at most $2.1\times10^{-3}$. On LibriSpeech against WavMark, AudioSeal, and Timbre, KeyBound holds 1.00 detection accuracy and 0.98 bit accuracy under a spectral denoiser that costs every baseline its detection, decodes at chance without the key, and rejects transplanted carriers. Detection further transfers to held-out DAC and BigVGAN re-synthesis, though exact payload recovery degrades. Speech provenance is thus better posed as a keyed, host-bound attribution problem than as the recovery of a payload under a catalog of signal distortions fixed in advance.