arXiv Science⌕ Search

arXiv · 2609.37540

Rate-Agnostic Bioacoustics: Heterogeneous Multi-Taxa Classification with Continuous Filterbanks and Fourier Neural Operators

Abstract

Conventional bioacoustic classification models rely on fixed-rate spectral representations, requiring recordings acquired at heterogeneous sampling rates to be resampled before analysis. We propose a Sampling-Frequency-Independent (SFI) frontend that processes each recording directly at its native sampling rate, coupled with a Fourier Neural Operator (FNO) backbone featuring progressive temporal-scale fusion. This framework avoids fixed-rate resampling and high-frequency information loss while producing fixed-size representations across sampling rates. Mild training-time sampling-rate (\textit{sr}) augmentation further improves robustness to unseen rate variations. Evaluated on a multi-taxa corpus comprising 84 classes and 60 sampling rates, the proposed SFI-FNO configuration outperforms fixed-rate and corpus-maximum-rate baselines, achieving .906 accuracy, .921 balanced accuracy, and a Macro-F1 score of .899.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Stefano Ciapponi, Francesco Ardan Dal Rı, Nicola Conci, Elisabetta Farella. 2026-09-29. Rate-Agnostic Bioacoustics: Heterogeneous Multi-Taxa Classification with Continuous Filterbanks and Fourier Neural Operators. https://arxiv.org/abs/2609.37540

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Trigger Sound Suppression for Misophonia

Misophonia, a disorder of decreased tolerance to specific sounds, affects 5-20% of the population, yet sufferers have no good options: therapy helps a minority, and earplugs or noise cancellation silence everything. We present a study for neural trigger sound suppression for misophonia, selectively removing trigger sounds. We curate a dataset covering the 10 most common trigger classes. Using streaming dual-path networks operating on 6 ms audio chunks, we explore both one-hot and multi-hot-conditioned models that suppress 1-3 triggers from the acoustic scene. We validate our model outputs in a listening study with 30 adults with clinically elevated misophonia impairment. Participants reported significantly lower distress and arousal, and improved valence, for suppressed audio.

cs.SD↗

RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition

Audio is increasingly used for human activity recognition (HAR) because it captures object interactions, environmental events, and contextual cues in everyday environments. High-resolution (HR) audio provides rich acoustic information for model development but incurs substantial energy and storage costs and may expose sensitive speech content. Low-resolution (LR) audio offers a more privacy-preserving and resource-efficient alternative for deployment, but reduced sampling rates can remove acoustic cues essential for activity recognition, leading to significant performance degradation. We formulate this training-deployment mismatch as sensor-resolution privileged learning, in which HR audio is available during training, while inference relies exclusively on LR audio. We propose RAST, a resolution-aware transfer framework that compresses HR teacher representations by preserving token-level information and neighborhood structure before performing localized HR-LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR-only training and direct teacher-transfer baselines, improving LR-only recognition by up to approximately 7.8% while requiring only LR audio at inference.

cs.SD↗

Audio Token Attention Is Predictable Before the Language Model Runs

A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map, fitted in closed form without labels, predicts this all-layer attention ranking at $ρ\geq .69$ on eleven of thirteen LALMs. Our method, Triage, cuts audio tokens by this prediction and, on multiple choice, cuts again at layer 2, correcting the prediction with the attention observed there. Triage sets its compression without labels, under two budgets that limit how far its output may differ from the model's own full-audio output. At the conservative budget, its word error rate and accuracy stay within .04 of full audio. At the aggressive budget, Triage beats every baseline in all twelve transcription cases. On multiple choice, at 2.2-5x compression, it outperforms DART, the strongest baseline on average, by .043 in mean accuracy. Because it cuts before the language model, it raises the audio that fits in Qwen2.5-Omni-3B's context window from 21.8 to about 62 minutes. At its most compressive point, Triage lets one GPU serve 4x as many concurrent 5-minute streams of that model. Project page: https://audio-triage.github.io

cs.SD↗