arXiv Science⌕ Search

arXiv · 2610.03398

Harmonic Eigenspace: A Web-based Application for Navigating and Composing Microtonal Harmony

Abstract

This paper presents a web-based application for navigating and composing microtonal harmony, built on the Harmonic Eigenspace, a four-dimensional psychoacoustically grounded space in which tetrad chord types are located by their spectral dissonance profiles, computed with Sethares's roughness/dissonance model. The coordinate system is transposition-invariant: a coordinate triple (α, \b{eta}, γ) locates the three upper notes in relation to the root, so a chord quality corresponds to a direction in the space, the invariant ray along which transposition acts, while the root frequency sets the scale. The dissonance field over these coordinates can be computed at any register; the locations of its local minima are register-invariant, as they arise from partial-coincidence ratio conditions. The dissonance volume contains 100 local minima that align with just-intonation intervals and act as landmarks, organising the space into basins around the most consonant tetrads. We embed tetrads from three tonal equal temperaments as discrete lattices within this continuous volume. The application presents this space through two components: the Harmonic Eigenspace as a navigable 3D visualisation of the dissonance volume in which all nodes are playable, and a Modal Studio that extends modal interchange logic to the ten-gradation interval vocabulary of 53-TET. The application also functions as a MIDI controller with MIDI Polyphonic Expression support, usable in any digital audio workstation that supports this format. A listening study with 31 participants used both scenes of the application: listeners first rated isolated 53-TET chords alongside chords familiar from Western practice, such as the maj7 and the m7; they then rated chord progressions composed in the Modal Studio, measuring their acceptance or rejection of microtonal progressions heard for the first time.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David Dalmazzo, Ken Déguernel. 2026-10-02. Harmonic Eigenspace: A Web-based Application for Navigating and Composing Microtonal Harmony. https://arxiv.org/abs/2610.03398

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A mathematical model of the vowel space

The articulatory-acoustic relationship is many-to-one and non linear and this is a great limitation for studying speech production. A simplification is proposed to set a bijection between the vowel space (f1, f2) and the parametric space of different vocal tract models. The generic area function model is based on mixtures of cosines allowing the generation of main vowels with two formulas. Then the mixture function is transformed into a coordination function able to deal with articulatory parameters. This is shown that the coordination function acts similarly with the Fant's model and with the 4-Tube DRM derived from the generic model.

cs.SD↗

Re-purposing Multimodal Large Language Models for Audio-Text Retrieval

Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentally limited by the capacity of small-scale encoders. Specifically, the text encoders struggle to understand complex queries that require reasoning or world knowledge. In this paper, we propose AuroLA, a novel contrastive language-audio pre-trained model that re-purposes Multimodal Large Language Models (MLLMs) as a unified backbone for audio-text retrieval. Specifically, we make the following contributions: (i) we construct a scalable data pipeline that curates diverse audio from multiple sources and generates multi-granular captions, ranging from long descriptions to structured tags, via automated annotation; (ii) we adapt an MLLM for retrieval by prompting it to summarise the audio/text input and using the hidden state of a special token as audio/text embeddings. (iii) extensive experiments demonstrate that AuroLA consistently outperforms state-of-the-art dual-encoder models, including the recent PE-AV. This validates the effectiveness of MLLM as a unified backbone for audio-text retrieval.

cs.SD↗

Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement

Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Model weights are available for download at: https://huggingface.co/nvidia/RE-USE.

cs.SD↗