arXiv ScienceSearch

arXiv subjects

Weijie Wu

Publications and source records attributed to Weijie Wu.

At least 19 recordsLinked to original sources

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

In current zero-shot text-to-speech systems, conventional semantic tokenizers are typically optimized using supervised automatic speech recognition or self-supervised learning objectives. However, due to the inherent nature of speech, semantic and acoustic information cannot be completely decoupled, and ASR-based tokenizers discard acoustic details to focus on linguistic content; models relying on them usually struggle to achieve optimal speaker similarity. Furthermore, these tokenizers are optimized independently and lack direct supervision from downstream acoustic generation tasks. This isolated training creates a feature gap between the extracted discrete tokens and the continuous space required by acoustic models, fundamentally bottlenecking the upper bound of synthesis quality. To bridge this gap, we propose Phoenix TTS, a unified framework that tightly couples representation learning with generative acoustic modeling. Specifically, our speech tokenizer is optimized to reconstruct self-supervised features to maintain semantic richness, while simultaneously receiving direct supervision from a Flow Matching training loss. Through this joint training paradigm, the extracted discrete tokens successfully preserve essential semantic information and natively align with the feature space of the downstream Flow Matching model. Comprehensive evaluations highlight the efficiency and effectiveness of Phoenix TTS. Trained on 110K hours of data, the system achieves excellent speech intelligibility, yielding WER that consistently falls below that of ground-truth recordings. Simultaneously, it maintains robust zero-shot speaker similarity, rivaling or outperforming several prominent large-scale baselines. Furthermore, as an advantageous byproduct of this unified training, the learned tokenizer can be seamlessly adapted to zero-shot voice conversion tasks without requiring task-specific fine-tuning.

cs.SD

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.

cs.SD

Precision quantum simulation of magnon spectra and interactions

Quantum simulation promises to advance materials discovery by accurately simulating complex states of matter, their microscopic excitations, and macroscopic response functions. The central challenge in resolving the underlying interacting dynamics is to combine high-fidelity evolution with the sophisticated control necessary to manipulate individual quasi-particles in quantum many-body states. Here, we report on high-precision simulation of both linear and non-linear response functions in a 2D XY spin-1/2 magnet using an analog-digital superconducting processor of up to 97 qubits. By interleaving digital gates with analog evolution precisely characterized via Hamiltonian learning, we selectively excite magnons at tunable energy densities. Measuring first the linear magnon response -- a central probe in neutron-scattering experiments -- we extract temperature-dependent spectra and lifetimes. Our results reveal stark variations in magnon decay rates across the Brillouin zone, with enhancement near van Hove singularities and suppression for edge-localized modes. Next, we perform a suite of nonlinear measurements, including the study of self-scattering mechanisms, as well as pump-probe spectroscopy to directly characterize the magnon interactions. While matrix-product state simulations capture the dynamics well in either small systems or at low temperatures, their predictions become inaccurate away from these limits. This work demonstrates precise simulation of the interacting dynamics in quantum magnets, and provides key insights into quasi-particles and their microscopic scattering mechanisms.

quant-ph

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audio but lack linguistic constraints, causing content errors during generation, whereas semantic tokens from self-supervised learning (SSL) models ensure precise text alignment but discard some acoustic information. To bridge this gap, we propose SARA, a dual-stream VAE that directly fuses a frozen SSL semantic anchor with a dedicated residual acoustic encoder. This effectively mitigates the dilemma, creating an efficient and compact latent space without relying on complex regularizers. SARA achieves superior reconstruction quality over strong baselines. Furthermore, in downstream zero-shot TTS tasks, it yields highly natural and expressive synthesis quality, and maintains robust generation performance even under accelerated inference, offering a favorable trade-off between synthesis speed and computational cost.

cs.SD

EarthEmbeddingExplorer: A Web Application for Cross-Modal Retrieval of Global Satellite Images

While the Earth observation community has witnessed a surge in high-impact foundation models and global Earth embedding datasets, a significant barrier remains in translating these academic assets into freely accessible tools. This tutorial introduces EarthEmbeddingExplorer, an interactive web application designed to bridge this gap, transforming static research artifacts into dynamic, practical workflows for discovery. We will provide a comprehensive hands-on guide to the system, detailing its cloud-native software architecture, demonstrating cross-modal queries (natural language, visual, and geolocation), and showcasing how to derive scientific insights from retrieval results. By democratizing access to precomputed Earth embeddings, this tutorial empowers researchers to seamlessly transition from state-of-the-art models and data archives to real-world application and analysis. The web application is available at https://modelscope.ai/studios/Major-TOM/EarthEmbeddingExplorer.

cs.CV

Theory of Scalable Spin Squeezing with Disordered Quantum Dipoles

Spin squeezed entanglement enables metrological precision beyond the classical limit. Understood through the lens of continuous symmetry breaking, dipolar spin systems exhibit the remarkable ability to generate spin squeezing via their intrinsic quench dynamics. To date, this understanding has primarily focused on lattice spin systems; in practice however, dipolar spin systems$\unicode{x2014}$ranging from ultracold molecules to nuclear spin ensembles and solid-state color centers$\unicode{x2014}$often exhibit significant amounts of positional disorder. Here, we develop a theory for scalable spin squeezing in a two-dimensional randomly diluted lattice of quantum dipoles, which naturally realize a dipolar XXZ model. Via extensive quantum Monte Carlo simulations, we map out the phase diagram for finite-temperature XY order, and by extension scalable spin squeezing, as a function of both disorder and Ising anisotropy. As the disorder increases, we find that scalable spin squeezing survives only near the Heisenberg point. We show that this behavior is due to the presence of rare tightly-coupled dimers, which effectively heat the system post-quench. In the case of strongly-interacting nitrogen-vacancy centers in diamond, we demonstrate that an experimentally feasible strategy to decouple the problematic dimers from the dynamics is sufficient to enable scalable spin squeezing.

quant-ph

Elucidating the Inter-system Crossing of the Nitrogen-Vacancy Center up to Megabar Pressures

The integration of Nitrogen-Vacancy color centers into diamond anvil cells has opened the door to quantum sensing at megabar pressures. Despite a multitude of experimental demonstrations and applications ranging from quantum materials to geophysics, a detailed microscopic understanding of how stress affects the NV center remains lacking. In this work, using a combination of first principles calculations as well as high-pressure NV experiments, we develop a complete description of the NV's optical properties under general stress conditions. In particular, our ab initio calculations reveal the complex behavior of the NV's inter-system crossing rates under stresses that both preserve and break the defect's symmetry. Crucially, our proposed framework immediately resolves a number of open questions in the field, including: (i) the microscopic origin of the observed contrast-enhancement in (111)-oriented anvils, and (ii) the surprising observation of NV contrast-inversion in certain high-pressure regimes. Our work lays the foundation for optimizing the performance of NV high-pressure sensors by controlling the local stress environment, and more generally, suggests that symmetry-breaking stresses can be utilized as a novel tuning knob for generic solid-state spin defects.

quant-ph

SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks.

eess.AS

Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction

Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we introduce Phoenix-VAD, an LLM-based model that enables streaming semantic endpoint detection. Specifically, Phoenix-VAD leverages the semantic comprehension capability of the LLM and a sliding window training strategy to achieve reliable semantic endpoint detection while supporting streaming inference. Experiments on both semantically complete and incomplete speech scenarios indicate that Phoenix-VAD achieves excellent and competitive performance. Furthermore, this design enables the full-duplex prediction module to be optimized independently of the dialogue model, providing more reliable and flexible support for next-generation human-computer interaction.

eess.AS

Patterning programmable spin arrays on DNA origami for quantum technologies

The controlled assembly of solid-state spins with nanoscale spatial precision is an outstanding challenge for quantum technology. Here, we combine DNA-based patterning with nitrogen-vacancy (NV) ensemble quantum sensors in diamond to form and sense programmable 2D arrays of spins. We use DNA origami to control the spacing of chelated Gd$^{3+}$ spins, as verified by the observed linear relationship between proximal NVs' relaxation rate, $1/T_1$, and the engineered number of Gd$^{3+}$ spins per origami unit. We further show that DNA origami provides a robust way of functionalizing the diamond surface with spins as it preserves the charge state and spin coherence of proximal, shallow NV centers. Our work enables the formation and interrogation of ordered, strongly interacting spin networks with applications in quantum sensing and quantum simulation. We quantitatively discuss the prospects of entanglement-enhanced metrology and high-throughput proteomics.

quant-ph

Multiscale Coupled Polarization and BKT Transitions in Tow-Dimensional Hybrid Organic-Inorganic Perovskites

We present an extended two-dimensional XY rotor model specifically designed to capture the polarization dynamics of hybrid organic-inorganic perovskite monolayers. This framework integrates nearest and next-nearest neighbor couplings, crystalline anisotropy inherent to perovskite lattice symmetries, external bias fields, and long-range dipolar interactions that are prominent in layered perovskite architectures. Through a combination of analytical coarse-graining and large-scale Monte Carlo simulations on 64*64 lattices, we identify two distinct thermodynamic regimes: a low-temperature quasi-ferroelectric state characterized by finite polarization and domain wall formation, and a higher-temperature Berezinskii-Kosterlitz-Thouless (BKT) crossover associated with vortex-antivortex unbinding and the suppression of long-range order. Our results reproduce key experimental signatures observed in quasi-two-dimensional perovskites, including dual peaks in dielectric susceptibility, enhanced vortex density near the transition, multistable polarization hysteresis under applied fields, and the scaling behavior of domain wall widths. This minimal yet realistic model provides a unifying perspective on how topological transitions and ferroelectric ordering coexist in layered perovskite systems, offering quantitative guidance for interpreting the emergent polar vortex lattices and complex phase behavior recently reported in hybrid perovskite thin films.

cond-mat.mtrl-sci

Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion

Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive speaking style of the target speaker, thereby limiting the controllability of voice conversion. In this work, we propose Discl-VC, a novel voice conversion framework that disentangles content and prosody information from self-supervised speech representations and synthesizes the target speaker's voice through in-context learning with a flow matching transformer. To enable precise control over the prosody of generated speech, we introduce a mask generative transformer that predicts discrete prosody tokens in a non-autoregressive manner based on prompts. Experimental results demonstrate the superior performance of Discl-VC in zero-shot voice conversion and its remarkable accuracy in prosody control for synthesized speech.

cs.SD

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech tokenizers has become increasingly important. This paper introduces DS-Codec, a novel neural speech codec featuring a dual-stage training framework with mirror and non-mirror architectures switching, designed to achieve superior speech reconstruction. We conduct extensive experiments and ablation studies to evaluate the effectiveness of our training strategy and compare the performance of the two architectures. Our results show that the mirrored structure significantly enhances the robustness of the learned codebooks, and the training strategy balances the advantages between mirrored and non-mirrored structures, leading to improved high-fidelity speech reconstruction.

cs.SD

InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition

Language-Guided object recognition in remote sensing imagery is crucial for large-scale mapping and automated data annotation. However, existing open-vocabulary and visual grounding methods rely on explicit category cues, limiting their ability to handle complex or implicit queries that require advanced reasoning. To address this issue, we introduce a new suite of tasks, including Instruction-Oriented Object Counting, Detection, and Segmentation (InstructCDS), covering open-vocabulary, open-ended, and open-subclass scenarios. We further present EarthInstruct, the first InstructCDS benchmark for earth observation. It is constructed from two diverse remote sensing datasets with varying spatial resolutions and annotation rules across 20 categories, necessitating models to interpret dataset-specific instructions. Given the scarcity of semantically rich labeled data in remote sensing, we propose InstructSAM, a training-free framework for instruction-driven object recognition. InstructSAM leverages large vision-language models to interpret user instructions and estimate object counts, employs SAM2 for mask proposal, and formulates mask-label assignment as a binary integer programming problem. By integrating semantic similarity with counting constraints, InstructSAM efficiently assigns categories to predicted masks without relying on confidence thresholds. Experiments demonstrate that InstructSAM matches or surpasses specialized baselines across multiple tasks while maintaining near-constant inference time regardless of object count, reducing output tokens by 89% and overall runtime by over 32% compared to direct generation approaches. We believe the contributions of the proposed tasks, benchmark, and effective approach will advance future research in developing versatile object recognition systems.

cs.CV

Spin squeezing in an ensemble of nitrogen-vacancy centers in diamond

Spin squeezed states provide a seminal example of how the structure of quantum mechanical correlations can be controlled to produce metrologically useful entanglement. Such squeezed states have been demonstrated in a wide variety of artificial quantum systems ranging from atoms in optical cavities to trapped ion crystals. By contrast, despite their numerous advantages as practical sensors, spin ensembles in solid-state materials have yet to be controlled with sufficient precision to generate targeted entanglement such as spin squeezing. In this work, we present the first experimental demonstration of spin squeezing in a solid-state spin system. Our experiments are performed on a strongly-interacting ensemble of nitrogen-vacancy (NV) color centers in diamond at room temperature, and squeezing (-0.5 $\pm$ 0.1 dB) is generated by the native magnetic dipole-dipole interaction between NVs. In order to generate and detect squeezing in a solid-state spin system, we overcome a number of key challenges of broad experimental and theoretical interest. First, we develop a novel approach, using interaction-enabled noise spectroscopy, to characterize the quantum projection noise in our system without directly resolving the spin probability distribution. Second, noting that the random positioning of spin defects severely limits the generation of spin squeezing, we implement a pair of strategies aimed at isolating the dynamics of a relatively ordered sub-ensemble of NV centers. Our results open the door to entanglement-enhanced metrology using macroscopic ensembles of optically active spins in solids.

quant-ph

A Universal Protocol for Quantum-Enhanced Sensing via Information Scrambling

We introduce a novel protocol, which enables Heisenberg-limited quantum-enhanced sensing using the dynamics of any interacting many-body Hamiltonian. Our approach - dubbed butterfly metrology - utilizes a single application of forward and reverse time evolution to produce a coherent superposition of a "scrambled" and "unscrambled" quantum state. In this way, we create metrologically-useful long-range entanglement from generic local quantum interactions. The sensitivity of butterfly metrology is given by a sum of local out-of-time-order correlators (OTOCs) - the prototypical diagnostic of quantum information scrambling. Our approach broadens the landscape of platforms capable of performing quantum-enhanced metrology; as an example, we provide detailed blueprints and numerical studies demonstrating a route to scalable quantum-enhanced sensing in ensembles of solid-state spin defects.

quant-ph

Impurity-level induced broadband photoelectric response in wide-band semiconductor SrSnO3

Broadband spectrum detectors exhibit great promise in fields such as multispectral imaging and optical communications. Despite significant progress, challenges like materials instability, complex manufacturing process and high costs still hinder further application. Here we present a method that achieves broadband spectral detect by impurity-level in SrSnO3. We report over 200 mA/W photo-responsivity at 275 nm (ultraviolet C solar-bind) and 367 nm (ultraviolet A) and ~ 1 mA/W photo-responsivity at 532 nm and 700 nm (visible) with a voltage bias of 5V. Further transport and photoluminescence results indicate that the broadband response comes from the impurity levels and mutual interactions. Additionally, the photodetector demonstrates excellent robustness and stability under repeated tests and prolonged exposure in air. These findings show the potential of SSO photodetectors and propose a method to achieve broadband spectrum detection, creating new possibility for the development of single-phase, low-cost, simple structure and high-efficiency photodetectors.

physics.app-ph

A strongly interacting, two-dimensional, dipolar spin ensemble in (111)-oriented diamond

Systems of spins with strong dipolar interactions and controlled dimensionality enable new explorations in quantum sensing and simulation. In this work, we investigate the creation of strong dipolar interactions in a two-dimensional ensemble of nitrogen-vacancy (NV) centers generated via plasma-enhanced chemical vapor deposition (PECVD) on (111)-oriented diamond substrates. We find that diamond growth on the (111) plane yields high incorporation of spins, both nitrogen and NV centers, where the density of the latter is tunable via the miscut of the diamond substrate. Our process allows us to form dense, preferentially aligned, 2D NV ensembles with volume-normalized AC sensitivity down to $\eta_{AC}$ = 810 pT um$^{3/2}$ Hz$^{-1/2}$. Furthermore, we show that (111) affords maximally positive dipolar interactions amongst a 2D NV ensemble, which is crucial for leveraging dipolar-driven entanglement schemes and exploring new interacting spin physics.

quant-ph