arXiv Science⌕ Search

arXiv · 2610.08236

Geometric Representations for Transformed Pattern Matching in Music

Abstract

We review the notion of representing music using point sets and argue that such representations are better adapted than sequential representations for matching patterns in unvoiced, polyphonic music, such as keyboard music. Musical transformations such as transposition, inversion, diminution, augmentation and retrograde can be modelled by geometric transformations in pitch-time representations that combine translation with scaling parallel to and reflection in the time axis. We identify eight types of geometric pitch-time representation that use chromatic pitch, morphetic pitch, morph or chroma to represent pitch and either onset time or midtime to represent time. We illustrate how these types of representation allow us to characterise different types of musical transformation. For example, by using midtime instead of onset time, we can precisely characterise certain retrograde relationships; and by using morph and chroma pitch representations we can characterise transformations involving octave displacements and duplications. We present the concept of a transformation class and consider the three specific classes, $F_{\mathrm{2STR}}$, $F_{\mathrm{2STRMod7}}$ and $F_{\mathrm{2STRMod12}}$. We introduce the notion of an inter-pattern transformation graph for a set of patterns, $S$, and a transformation class, $F$. Each vertex in such a graph represents a pattern in $S$ and there is an edge in the graph from pattern $P_1$ to pattern $P_2$ if and only if $P_1$ can be mapped onto $P_2$ by a transformation in $F$. We show, with the aid of such graphs, that the musical relationships between the occurrences of the HAYDN theme in Ravel's Menuet sur le nom d'Haydn can be precisely described in terms of transformations in $F_{\mathrm{2STR}}$, $F_{\mathrm{2STRMod7}}$ and $F_{\mathrm{2STRMod12}}$ within the pitch-time representations considered.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David Meredith. 2026-10-06. Geometric Representations for Transformed Pattern Matching in Music. https://arxiv.org/abs/2610.08236

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard-like speech in a zero-shot manner without requiring Lombard-specific training data. Our approach extends F5-TTS with a learned style embedding representation and analyzes the resulting latent space using principal component analysis (PCA) to identify directions associated with Lombard-related attributes. By manipulating these directions, we obtain interpretable control over vocal effort and articulation and generate speech at different Lombard levels. Experimental results show that the proposed method preserves speaker identity and naturalness, improves intelligibility under noisy conditions, and generalizes to previously unseen speakers. These findings demonstrate that style-embedding manipulation provides an effective and scalable framework for controllable zero-shot Lombard speech synthesis.

cs.SD↗

Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement

Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Model weights are available for download at: https://huggingface.co/nvidia/RE-USE.

cs.SD↗

Voice "Cloning" is Style Transfer

Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the voices of individuals with speech loss. However, in our work, we find that despite the term, voice cloning does not faithfully ''clone'' an individual's voice. Instead, we find that widely-used voice cloning models systematically apply style transfer to source voices. As rated by human annotators, cloned voices are perceived as more authoritative, warm, customer-service-like, and human-like compared to their sources. Human annotators also report greater trust in cloned voices than source voices, and a greater willingness to disclose sensitive personal information to them. Our work furthermore shows that voice cloning leads to homogenization of speaker characteristics, as measured by reduced variance in accent, speaking rate, and the audio embedding space. Together, our results highlight a new set of limitations and risks of voice cloning technology and their potential impact on human behavior.

cs.SD↗