arXiv ScienceSearch

arXiv subjects

Ruimin Wang

Publications and source records attributed to Ruimin Wang.

9 recordsLinked to original sources

Stabilizing Instruction Supervision for Instruct-TTS via Controllable Diversification and Drift Filtering

Instruct-TTS systems expand structured style labels into natural-language training instructions through LLM rewriting, yet we find that over 40% of unconstrained rewrites contain semantic drift that corrupts supervision and weakens generalization. We formalize this problem as instruction supervision instability and propose a data-centric stabilization recipe that jointly improves coverage and fidelity through three mechanisms: controllable instruction diversification for systematic expansion, LLM-based drift filtering for quality control, and attribute-aligned supervision that grounds prosody control in acoustic perturbations. On the Chinese split of InstructTTSEval, our recipe raises instruction-following from 34.5% without fine-tuning and 51.0% with naive fine-tuning to 56.4%, while constrained rewriting reduces drift from 40.4% to 15.4%. Ablations confirm the three mechanisms are complementary, and the drift taxonomy may generalize to instruction-driven generation beyond TTS.

cs.SD

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains. To address this gap, we introduce MeloCodec, a novel framework designed to effectively incorporate melodic priors, a critical form of acoustic information for singing. To address the optimization instability typically caused by the direct fusion of such explicit priors, we propose a Tokenize-then-Fuse paradigm that pre-trains a discrete melodic branch to lock in structures before feature fusion. To robustly realize this paradigm, we further propose a two-stage training strategy that prevents codebook collapse and ensures stable convergence. Experiments show that MeloCodec outperforms baselines in singing voice representation, improving pitch consistency and enabling controllable pitch manipulation with minimal timbre degradation.

cs.SD

CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis

Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict absent from ordinary reconstruction: training pairs reference cues with original lyrics, whereas inference asks revised lyrics to override source-lyric-correlated cues; one source-following patch can propagate through AR history. We introduce CLASVS. Its State-Control-Transition (SCT) routing keeps target-lyric and reference-melody controls persistent, returns semantic feedback on phonetic progress to the causal planner, and confines the previous latent patch to the local Transition. Progressive State-Control Grounding (PSCG) learns this contract through paired-edit-free, content-consistent Mandarin reconstruction. On two Mandarin benchmarks, CLASVS improves all four operations over discrete-AR Vevo2 and reduces macro-PER by 46.2%, while maintaining melody, singer similarity, and perceptual quality. Together, these results establish a strong continuous-AR operating point for score-annotation-free lyric edits and a basis for broader stepwise control. Audio demonstrations are available on our project page: https://piedpiperg.github.io/clasvs-demo/.

cs.SD

UniGeo: A Unified 3D Indoor Object Detection Framework Integrating Geometry-Aware Learning and Dynamic Channel Gating

The growing adoption of robotics and augmented reality in real-world applications has driven considerable research interest in 3D object detection based on point clouds. While previous methods address unified training across multiple datasets, they fail to model geometric relationships in sparse point cloud scenes and ignore the feature distribution in significant areas, which ultimately restricts their performance. To deal with this issue, a unified 3D indoor detection framework, called UniGeo, is proposed. To model geometric relations in scenes, we first propose a geometry-aware learning module that establishes a learnable mapping from spatial relationships to feature weights, which enabes explicit geometric feature enhancement. Then, to further enhance point cloud feature representation, we propose a dynamic channel gating mechanism that leverages learnable channel-wise weighting. This mechanism adaptively optimizes features generated by the sparse 3D U-Net network, significantly enhancing key geometric information. Extensive experiments on six different indoor scene datasets clearly validate the superior performance of our method.

cs.CV

Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention

Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce annotated data. To tackle these challenges, we propose a multi-loss learning (MLL) framework integrating an energy-adaptive mixup (EAM) method and a frame-level attention module (FLAM). The EAM method leverages SNR-based augmentation to generate diverse speech samples capturing subtle emotional variations. FLAM enhances frame-level feature extraction for multi-frame emotional cues. Our MLL strategy combines Kullback-Leibler divergence, focal, center, and supervised contrastive loss to optimize learning, address class imbalance, and improve feature separability. We evaluate our method on four widely used SER datasets: IEMOCAP, MSP-IMPROV, RAVDESS, and SAVEE. The results demonstrate our method achieves state-of-the-art performance, suggesting its effectiveness and robustness.

cs.SD

Data-Driven Prognosis of Failure Detection and Prediction of Lithium-ion Batteries

Battery prognostics and health management predictive models are essential components of safety and reliability protocols in battery management system frameworks. Overall, developing a robust and efficient battery model that aligns with the current literature is a useful step in ensuring the safety of battery function. For this purpose, a multi-physics, multi-scale deterministic data-driven prognosis (DDP) is proposed that only relies on in situ measurements of data and estimates the failure based on the curvature information extracted from the system. Unlike traditional applications that require explicit expression of conservation principle, the proposed method devices a local conservation functional in the neighborhood of each data point which is represented as the minimization of curvature in the system. By eliminating the need for offline training, the method can predict the onset of instability for a variety of systems over a prediction horizon. The prediction horizon to prognosticate the instability, alternatively, is considered as the remaining useful life (RUL) metric. The framework is then employed to analyze the health status of Li-ion batteries. Based on the results, it has demonstrated that the DDP technique can accurately predict the onset of failure of Li-ion batteries.

physics.data-an

Modelling Protagonist Goals and Desires in First-Person Narrative

Many genres of natural language text are narratively structured, a testament to our predilection for organizing our experiences as narratives. There is broad consensus that understanding a narrative requires identifying and tracking the goals and desires of the characters and their narrative outcomes. However, to date, there has been limited work on computational models for this problem. We introduce a new dataset, DesireDB, which includes gold-standard labels for identifying statements of desire, textual evidence for desire fulfillment, and annotations for whether the stated desire is fulfilled given the evidence in the narrative context. We report experiments on tracking desire fulfillment using different methods, and show that LSTM Skip-Thought model achieves F-measure of 0.7 on our corpus.

cs.AI

Lattice thermal expansion and anisotropic displacements in urea, bromomalonic aldehyde, pentachloropyridine and naphthalene

Anisotropic displacement parameters (ADPs) are commonly used in crystallography, chemistry and related fields to describe and quantify thermal motion of atoms. Within the very recent years, these ADPs have become predictable by lattice dynamics in combination with first-principles theory. Here, we study four very different molecular crystals, namely urea, bromomalonic aldehyde, pentachloropyridine, and naphthalene, by first-principles theory to assess the quality of ADPs calculated in the quasi-harmonic approximation. In addition, we predict both thermal expansion and thermal motion within the quasi-harmonic approximation and compare the predictions with experimental data. Very reliable ADPs are calculated within the quasi-harmonic approximation for all four cases up to at least 200 K, and they turn out to be in better agreement with experiment than the harmonic ones. In one particular case, ADPs can even reliably be predicted up to room temperature. Our results also hint at the importance of normal-mode anharmonicity in the calculation of ADPs.

cond-mat.mtrl-sci

Multicharged optical vortices induced in a dissipative atomic vapor system

We investigate numerically the dynamics of optical vortex beams carrying different topological charges, launched in a dissipative three level ladder type nonlinear atomic vapor. We impose the electromagnetically induced transparency (EIT) condition on the medium. Linear, cubic, and quintic susceptibilities, considered simultaneously with the dressing effect, are included in the analysis. Generally, the beams slowly expand during propagation and new vortices are induced, commonly appearing in oppositely charged pairs. We demonstrate that not only the form and the topological charge of the incident beam, but also its growing size in the medium greatly affect the formation and evolution of vortices. We formulate common rules for finding the number of induced vortices and the corresponding rotation directions, stemming from the initial conditions of various incident beams, as well as from the dynamical aspects of their propagation. The net topological charge of the vortex is conserved during propagation, as it should be, but the total number of charges is not necessarily same as the initial number, because of the complex nature of the system. When the EIT condition is lifted, an enhancement region of beam dynamics if reached, in which the dynamics and the expansion of the beam greatly accelerate. In the end, we discuss the liquid like behavior of light evolution in this dissipative system and propose a potential experimental scheme for observing such a behavior.

physics.optics