arXiv ScienceSearch

arXiv subjects

Li Su

Publications and source records attributed to Li Su.

At least 19 recordsLinked to original sources

Separate-and-Detect: Unified Drum Transcription and Stem Generation via Latent Diffusion

Automatic Drum Transcription (ADT) is commonly formulated as a direct mapping from a music mixture to symbolic drum events. While effective for transcription, this formulation discards the acoustic stems that are useful for editing, remixing, and production. We revisit an alternative separate-and-detect formulation, where a drum source separation front end first produces five editable drum stems, and a fixed onset detector then converts each stem into symbolic events. The separator is built on a five-stem latent diffusion model that jointly generates kick, snare, toms, hi-hats, and cymbals in a compact VAE latent space. We further study two training-only auxiliary branches--an onset branch (OB) and a timbre branch (TB)--which shape the separator during learning but are discarded at inference. Trained on synthetic drum multitracks and evaluated on MDB Drums and ENST-Drums, the proposed pipeline consistently improves over a strong U-Net-based drum separation baseline in overall transcription F1. It also outperforms a representative end-to-end ADT system on kick and snare F1 under our evaluation protocol, while additionally providing separated audio stems. The ablation results show that OB gives the most stable transcription gains, whereas TB changes the trade-off between reconstruction, acoustic stem quality, and onset detection. These results suggest that generative drum demixing can serve not only as a source separation model, but also as a practical front end for interpretable drum transcription.

cs.SD

Improving Longitudinal Targeted Maximum Likelihood Estimation in Target Trial Emulation using Joint Calibrated Weights

In target trial emulation (TTE), marginal structural models (MSMs) can be used to characterise per-protocol treatment effects over time. The MSM parameters are often estimated by inverse probability weighting (IPW), with weights estimated by maximum likelihood. However, IPW-based estimators can be unstable in small samples and are sensitive to misspecification of the weight models. An alternative method for estimating the MSM parameters is longitudinal targeted maximum likelihood estimation (LTMLE). LTMLE is double robust and potentially more efficient than IPW. Nevertheless, LTMLE also relies on inverse probability weights and may therefore share the instability of IPW-based estimators. We propose joint calibrated LTMLE, which integrates LTMLE with joint calibrated weights tailored for per-protocol effect estimation in TTE. This calibration of weights improves finite-sample performance by enforcing covariate balance in both the treatment and censoring processes simultaneously. Simulations show that the proposed method has improved efficiency and robustness to weight model misspecification, compared to standard LTMLE. We illustrate the method using a case study to evaluate the effect of highly active antiretroviral therapy on CD4 cell count among HIV-positive women.

stat.ME

Probability-Invariant Random Walk Learning on Gyral Folding-Based Cortical Similarity Networks for Alzheimer's and Lewy Body Dementia Diagnosis

Alzheimer's disease (AD) and Lewy body dementia (LBD) present overlapping clinical features yet require distinct diagnostic strategies. While neuroimaging-based brain network analysis is promising, atlas-based representations may obscure individualized anatomy. Gyral folding-based networks using three-hinge gyri provide a biologically grounded alternative, but inter-individual variability in cortical folding results in inconsistent landmark correspondence and highly irregular network sizes, violating the fixed-topology and node-alignment assumptions of most existing graph learning methods, particularly in clinical datasets where pathological changes further amplify anatomical heterogeneity. We therefore propose a probability-invariant random-walk-based framework that classifies individualized gyral folding networks without explicit node alignment. Cortical similarity networks are built from local morphometric features and represented by distributions of anonymized random walks, with an anatomy-aware encoding that preserves permutation invariance. Experiments on a large clinical cohort of AD and LBD subjects show consistent improvements over existing gyral folding and atlas-based models, demonstrating robustness and potential for dementia diagnosis.

q-bio.NC

STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning Models

Large Reasoning Models (LRMs) have advanced automated multi-step reasoning, but their ability to generate complex Chain-of-Thought (CoT) trajectories introduces severe privacy risks, as sensitive information may be deeply embedded throughout the reasoning process. Existing Large Language Models (LLMs) unlearning approaches that typically focus on modifying only final answers are insufficient for LRMs, as they fail to remove sensitive content from intermediate steps, leading to persistent privacy leakage and degraded security. To address these challenges, we propose Sensitive Trajectory Regulation (STaR), a parameter-free, inference-time unlearning framework that achieves robust privacy protection throughout the reasoning process. Specifically, we first identify sensitive content via semantic-aware detection. Then, we inject global safety constraints through secure prompt prefix. Next, we perform trajectory-aware suppression to dynamically block sensitive content across the entire reasoning chain. Finally, we apply token-level adaptive filtering to prevent both exact and paraphrased sensitive tokens during generation. Furthermore, to overcome the inadequacies of existing evaluation protocols, we introduce two metrics: Multi-Decoding Consistency Assessment (MCS), which measures the consistency of unlearning across diverse decoding strategies, and Multi-Granularity Membership Inference Attack (MIA) Evaluation, which quantifies privacy protection at both answer and reasoning-chain levels. Experiments on the R-TOFU benchmark demonstrate that STaR achieves comprehensive and stable unlearning with minimal utility loss, setting a new standard for privacy-preserving reasoning in LRMs.

cs.AI

Timed text extraction from Taiwanese Kua-\'a-h\`i TV series

Taiwanese opera (Kua-\'a-h\`i), a major form of local theatrical tradition, underwent extensive television adaptation notably by pioneers like I\^unn L\=e-hua. These videos, while potentially valuable for in-depth studies of Taiwanese opera, often have low quality and require substantial manual effort during data preparation. To streamline this process, we developed an interactive system for real-time OCR correction and a two-step approach integrating OCR-driven segmentation with Speech and Music Activity Detection (SMAD) to efficiently identify vocal segments from archival episodes with high precision. The resulting dataset, consisting of vocal segments and corresponding lyrics, can potentially supports various MIR tasks such as lyrics identification and tune retrieval. Code is available at https://github.com/z-huang/ocr-subtitle-editor .

cs.SD

Erkang-Diagnosis-1.1 Technical Report

This report provides a detailed introduction to Erkang-Diagnosis-1.1 model, our AI healthcare consulting assistant developed using Alibaba Qwen-3 model. The Erkang model integrates approximately 500GB of high-quality structured medical knowledge, employing a hybrid approach combining enhanced pre-training and retrieval-enhanced generation to create a secure, reliable, and professional AI health advisor. Through 3-5 efficient interaction rounds, Erkang Diagnosis can accurately understand user symptoms, conduct preliminary analysis, and provide valuable diagnostic suggestions and health guidance. Designed to become users intelligent health companions, it empowers primary healthcare and health management. To validate, Erkang-Diagnosis-1.1 leads GPT-4 in terms of comprehensive medical exams.

cs.AI

LargeSHS: A large-scale dataset of music adaptation

Recent advances in AI-based music generation have focused heavily on text-conditioned models, with less attention given to reference-based generation such as song adaptation. To support this line of research, we introduce LargeSHS, a large-scale dataset derived from SecondHandSongs, containing over 1.7 million metadata entries and approximately 900k publicly accessible audio links. Unlike existing datasets, LargeSHS includes structured adaptation relationships between musical works, enabling the construction of adaptation trees and performance clusters that represent cover song families. We provide comprehensive statistics and comparisons with existing datasets, highlighting the unique scale and richness of LargeSHS. This dataset paves the way for new research in cover song generation, reference-based music generation, and adaptation-aware MIR tasks.

cs.SD

Spectral extremal problems for the $(p,Q)$-spectral radius of hypergraphs

Let $Q$ be an $s$-vertex $r$-uniform hypergraph, and let $H$ be an $n$-vertex $r$-uniform hypergraph. Denote by $\mathcal{N}(Q,H)$ the number of isomorphic copies of $Q$ in $H$. For a hereditary family $\mathcal{P}$ of $r$-uniform hypergraphs, define $$\pi(Q,\mathcal{P}):=\lim\limits_{n\to \infty}\binom{n}{s}^{-1}\max\{\mathcal{N}(Q,H): H\in \mathcal{P}~~\mbox{and}~~|V(H)|=n\}.$$ For $p\geq1$, the $(p,Q)$-spectral radius of $H$ is defined as $$\lambda^{(p)}(Q,H):=\max_{\|\mathbf{x}\|_{p}=1}s!\sum_{\{i_{1},\ldots,i_{s}\}\in \binom{[n]}{s}}\mathcal{N}(Q,H[\{i_{1},\ldots,i_{s}\}])x_{i_{1}}\cdots x_{i_{s}}.$$ In this paper, we present a systematically investigation of the parameter $\lambda^{(p)}(Q,H)$. First, we prove that the limit $$\lambda^{(p)}(Q,\mathcal{P}):=\lim\limits_{n\to \infty}n^{s/p-s}\max\{\lambda^{(p)}(Q,H): H\in \mathcal{P}~~\mbox{and}~~|V(H)|=n\}$$ exists, and for $p>1$, it satisfies $$\pi(Q,\mathcal{P})=\lambda^{(p)}(Q,\mathcal{P}).$$ Second, we study spectral generalized Tur\'an problems. Specifically, we establish a spectral stability result and apply it to derive a spectral version of the Erd\H{o}s Pentagon Problem: for $p\geq1$ and sufficiently large $n$, the balanced blow-up of $C_{5}$ maximizes $\lambda^{(p)}(C_{5},H)$ among all $n$-vertex triangle-free graphs $H$, thereby improving a result of Liu \cite{Liu2025}. Furthermore, we show that for $p\geq1$ and sufficiently large $n$, the $l$-partite Tur\'an graph $T_{l}(n)$ attains the maximum $\lambda^{(p)}(K_{s},H)$ among all $n$-vertex F-free graphs $H$, where $F$ is an edge-critical graph with $\chi(F)=l+1$. This provides a spectral analogue of a theorem due to Ma and Qiu \cite{MQ2020}.

math.CO

SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control

Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particularly challenging. Recent approaches to timbre transfer often rely on spectral objectives or implicit style matching, offering limited control over envelope shaping. Moreover, public synthesizer datasets rarely provide diverse coverage of timbres and ADSR envelopes. To address these gaps, we present SynthCloner, a factorized codec model that disentangles audio into three attributes: ADSR envelope, timbre, and content. This separation enables expressive audio transfer with independent control over these attributes. Additionally, we introduce SynthCAT, a new synthesizer dataset with a task-specific rendering pipeline covering 250 timbres, 120 ADSR envelopes, and 100 MIDI sequences. Experiments show that SynthCloner outperforms baselines on both objective and subjective metrics, while enabling independent attribute control. The code, model checkpoint, and audio examples are available at https://buffett0323.github.io/synthcloner/.

eess.AS

Spectral Tur\'an-type problems for the $\alpha$-spectral radius of hypergraphs with degree stability

An $r$-pattern $P$ is an ordered pair $P=([l],E)$, where $l$ is a positive integer and $E$ is a set of $r$-multisets with elements from $[l]$. An $r$-graph $H$ is said to be $P$-colorable if there is a homomorphism $\phi$: $V(H)\rightarrow [l]$ such that $\{\phi(v_{1}),\ldots,\phi(v_{r})\}\in E$ for every edge $\{v_{1},\ldots,v_{r}\}\in E(H)$. Let $\mathrm{Col}(P)$ denote the family of all $P$-colorable $r$-graphs. This paper studies spectral extremal problems for $\alpha$-spectral radius of hypergraphs via analytic techniques. We first prove that for any $r$-pattern $P$, the hypergraph attaining the maximum $\alpha$-spectral radius in $\mathrm{Col}(P)$ is asymptotically regular. Specifically, we establish asymptotically tight lower bounds for the minimum component of the principal eigenvector and the minimum degree of the spectral extremal hypergraphs in $\mathrm{Col}(P)$. Building on this regularity, we further show that for any family $\mathcal{F}$ of $r$-graphs that is degree-stable with respect to $\mathrm{Col}(P)$, spectral Tur\'an-type problems can be completely reduced to spectral extremal problems within $\mathrm{Col}(P)$. As an application, we determine the maximum $\alpha$-spectral radius ($\alpha\geq1$) among all $n$-vertex $F^{(r)}$-free $r$-graphs, where $F^{(r)}$ is the $r$-expansion of the color-critical graph $F$. This provides a powerful reduction tool for handling spectral Tur\'{a}n-type problems in hypergraphs. Finally, leveraging the spectral method, we derive a corresponding edge Tur\'an extremal result. More precisely, we show that if $\mathcal{F}$ is degree-stable with respect to $\mathrm{Col}(P)$, then every $\mathcal{F}$-free edge extremal hypergraph must be a $P$-colorable hypergraph.

math.CO

VioPTT: Violin Technique-Aware Transcription from Synthetic Data Augmentation

While automatic music transcription is well-established in music information retrieval, most models are limited to transcribing pitch and timing information from audio, and thus omit crucial expressive and instrument-specific nuances. One example is playing technique on the violin, which affords its distinct palette of timbres for maximal emotional impact. Here, we propose VioPTT (Violin Playing Technique-aware Transcription), a lightweight cascade model that directly transcribes violin playing technique in addition to pitch onset and offset. Furthermore, we release MOSA-VPT, a novel, high-quality synthetic violin playing technique dataset to circumvent the need for manually labeled annotations. Leveraging this dataset, our model demonstrated strong generalization to real-world note-level violin technique recordings in addition to achieving state-of-the-art transcription performance. To our knowledge, VioPTT is the first to jointly combine violin transcription and playing technique prediction within a unified framework.

cs.SD

Enhancing Automatic Chord Recognition through LLM Chain-of-Thought Reasoning

Music Information Retrieval (MIR) encompasses a broad range of computational techniques for analyzing and understanding musical content, with recent deep learning advances driving substantial improvements. Building upon these advances, this paper explores how large language models (LLMs) can serve as an integrative bridge to connect and integrate information from multiple MIR tools, with a focus on enhancing automatic chord recognition performance. We present a novel approach that positions text-based LLMs as intelligent coordinators that process and integrate outputs from diverse state-of-the-art MIR tools-including music source separation, key detection, chord recognition, and beat tracking. Our method converts audio-derived musical information into textual representations, enabling LLMs to perform reasoning and correction specifically for chord recognition tasks. We design a 5-stage chain-of-thought framework that allows GPT-4o to systematically analyze, compare, and refine chord recognition results by leveraging music-theoretical knowledge to integrate information across different MIR components. Experimental evaluation on three datasets demonstrates consistent improvements across multiple evaluation metrics, with overall accuracy gains of 1-2.77% on the MIREX metric. Our findings demonstrate that LLMs can effectively function as integrative bridges in MIR pipelines, opening new directions for multi-tool coordination in music information retrieval tasks.

cs.SD

Is Transfer Learning Necessary for Violin Transcription?

Automatic music transcription (AMT) has achieved remarkable progress for instruments such as the piano, largely due to the availability of large-scale, high-quality datasets. In contrast, violin AMT remains underexplored due to limited annotated data. A common approach is to fine-tune pretrained models for other downstream tasks, but the effectiveness of such transfer remains unclear in the presence of timbral and articulatory differences. In this work, we investigate whether training from scratch on a medium-scale violin dataset can match the performance of fine-tuned piano-pretrained models. We adopt a piano transcription architecture without modification and train it on the MOSA dataset, which contains about 30 hours of aligned violin recordings. Our experiments on URMP and Bach10 show that models trained from scratch achieved competitive or even superior performance compared to fine-tuned counterparts. These findings suggest that strong violin AMT is possible without relying on pretrained piano representations, highlighting the importance of instrument-specific data collection and augmentation strategies.

cs.SD

Bridging Brain Connectomes and Clinical Reports for Early Alzheimer's Disease Diagnosis

Integrating brain imaging data with clinical reports offers a valuable opportunity to leverage complementary multimodal information for more effective and timely diagnosis in practical clinical settings. This approach has gained significant attention in brain disorder research, yet a key challenge remains: how to effectively link objective imaging data with subjective text-based reports, such as doctors' notes. In this work, we propose a novel framework that aligns brain connectomes with clinical reports in a shared cross-modal latent space at both the subject and connectome levels, thereby enhancing representation learning. The key innovation of our approach is that we treat brain subnetworks as tokens of imaging data, rather than raw image patches, to align with word tokens in clinical reports. This enables a more efficient identification of system-level associations between neuroimaging findings and clinical observations, which is critical since brain disorders often manifest as network-level abnormalities rather than isolated regional alterations. We applied our method to mild cognitive impairment (MCI) using the ADNI dataset. Our approach not only achieves state-of-the-art predictive performance but also identifies clinically meaningful connectome-text pairs, offering new insights into the early mechanisms of Alzheimer's disease and supporting the development of clinically useful multimodal biomarkers.

cs.CV

Whole-brain Transferable Representations from Large-Scale fMRI Data Improve Task-Evoked Brain Activity Decoding

A fundamental challenge in neuroscience is to decode mental states from brain activity. While functional magnetic resonance imaging (fMRI) offers a non-invasive approach to capture brain-wide neural dynamics with high spatial precision, decoding from fMRI data -- particularly from task-evoked activity -- remains challenging due to its high dimensionality, low signal-to-noise ratio, and limited within-subject data. Here, we leverage recent advances in computer vision and propose STDA-SwiFT, a transformer-based model that learns transferable representations from large-scale fMRI datasets via spatial-temporal divided attention and self-supervised contrastive learning. Using pretrained voxel-wise representations from 995 subjects in the Human Connectome Project (HCP), we show that our model substantially improves downstream decoding performance of task-evoked activity across multiple sensory and cognitive domains, even with minimal data preprocessing. We demonstrate performance gains from larger receptor fields afforded by our memory-efficient attention mechanism, as well as the impact of functional relevance in pretraining data when fine-tuning on small samples. Our work showcases transfer learning as a viable approach to harness large-scale datasets to overcome challenges in decoding brain activity from fMRI data.

eess.IV

Improving BERT for Symbolic Music Understanding Using Token Denoising and Pianoroll Prediction

We propose a pre-trained BERT-like model for symbolic music understanding that achieves competitive performance across a wide range of downstream tasks. To achieve this target, we design two novel pre-training objectives, namely token correction and pianoroll prediction. First, we sample a portion of note tokens and corrupt them with a limited amount of noise, and then train the model to denoise the corrupted tokens; second, we also train the model to predict bar-level and local pianoroll-derived representations from the corrupted note tokens. We argue that these objectives guide the model to better learn specific musical knowledge such as pitch intervals. For evaluation, we propose a benchmark that incorporates 12 downstream tasks ranging from chord estimation to symbolic genre classification. Results confirm the effectiveness of the proposed pre-training objectives on downstream tasks.

cs.SD

Domain-Adaptive Diagnosis of Lewy Body Disease with Transferability Aware Transformer

Lewy Body Disease (LBD) is a common yet understudied form of dementia that imposes a significant burden on public health. It shares clinical similarities with Alzheimer's disease (AD), as both progress through stages of normal cognition, mild cognitive impairment, and dementia. A major obstacle in LBD diagnosis is data scarcity, which limits the effectiveness of deep learning. In contrast, AD datasets are more abundant, offering potential for knowledge transfer. However, LBD and AD data are typically collected from different sites using different machines and protocols, resulting in a distinct domain shift. To effectively leverage AD data while mitigating domain shift, we propose a Transferability Aware Transformer (TAT) that adapts knowledge from AD to enhance LBD diagnosis. Our method utilizes structural connectivity (SC) derived from structural MRI as training data. Built on the attention mechanism, TAT adaptively assigns greater weights to disease-transferable features while suppressing domain-specific ones, thereby reducing domain shift and improving diagnostic accuracy with limited LBD data. The experimental results demonstrate the effectiveness of TAT. To the best of our knowledge, this is the first study to explore domain adaptation from AD to LBD under conditions of data scarcity and domain shift, providing a promising framework for domain-adaptive diagnosis of rare diseases.

cs.LG

Robust estimation of optimal dynamic treatment regimes with nonignorable missing covariates

Estimating optimal dynamic treatment regimes (DTRs) using observational data is often challenged by nonignorable missing covariates arsing from informative monitoring of patients in clinical practice. To address nonignorable missingness of pseudo-outcomes induced by nonignorable missing covariates, a weighted Q-learning approach using parametric Q-function models and a semiparametric missingness propensity model has recently been proposed. However, misspecification of parametric Q-functions at later stages of a DTR can propagate estimation errors to earlier stages via the pseudo-outcomes themselves and indirectly through biased estimation of the missingness propensity of the pseudo-outcomes. This robustness concern motivates us to develop a direct-search-based optimal DTR estimator built on a robust and efficient value estimator, where nonparametric methods are employed for treatment propensity and Q-function estimation, and inverse probability weighting is applied using missingness propensity estimated with the aid of nonresponse instrumental variables. Specifically, in our value estimator, we replace weights estimated by prediction models of treatment propensity with stable weights estimated by balancing covariate functions in a reproducing-kernel Hilbert space (RKHS). Augmented by Q-functions estimated by RKHS-based smoothing splines, our value estimator mitigates the misspecification risk of the weighted Q-learning approach while maintaining the efficiency gain from employing pseudo-outcomes in missing data scenarios. The asymptotic properties of the proposed estimator are derived, and simulations demonstrate its superior performance over weighted Q-learning under model misspecification. We apply the proposed methods to investigate the optimal fluid strategy for sepsis patients using data from the MIMIC database.

stat.ME