arXiv ScienceSearch

arXiv subjects

Xiang Hao

Publications and source records attributed to Xiang Hao.

At least 19 recordsLinked to original sources

Decoherence of ergotropy of a relativistic battery as a probe of motion-selected Unruh thermality in Minkowski spacetime

We put forward a physical model of a relativistic Unruh-DeWitt battery moving along a general accelerated trajectory, including a linear motion and a circular motion with a reflecting boundary. The maximal amount of quantum work extraction, defined as the ergotropy, serves as a witness to Unruh effect modified by motion trajectories. It is found out that for a very low Unruh temperature, linear motion yields a high amount of ergotropy, while for a high temperature, circular motion becomes optimal for estimating the Unruh effect. For a specific acceleration, the ergotropy is the same for two different trajectories. The behavior demonstrates that, in a certain condition, one can simulate the work extraction of the accelerated battery in a linear motion by exploring the battery in circular motion. The observed ergotropy for the thermality is closely related to quantum coherence affected by spacetime vacuum fluctuations. In the presence of a reflecting boundary plane, we study different ways for the dynamics of the ergotropy. When the battery moves near the boundary, the prominent oscillation of the ergotropy will happen and exhibit some large instantaneous peaks. The interesting phenomenon results from the protection of quantum coherence for the accelerated battery in the vicinity of the boundary. Far away from the boundary, the oscillation behavior can be suppressed and the ergotropy rapidly arrives at a steady value. By contrast, the circular motion contributes to prolonging the oscillation evolution. The asymptotic amount of the ergotropy can rise progressively up to a saturation value with increasing the distance from the circular trajectory to the boundary plane. From the perspective of energy transfer, optimized quantum work extraction for the accelerated battery moving along a selected motion with a boundary plane is advantageous for probing Unruh thermality.

gr-qc

Quantum thermodynamics of ergotopy for a relativistic battery as a witness to Unruh-Hawking thermality in curved (A)dS spacetimes

We propose a relativistic quantum battery model consisting of an accelerated Unruh-DeWitt detector coupled to a massless scalar field in a de Sitter and anti-de Sitter spacetimes. The maximal amount of quantum extractable work, defined as the ergotropy, is used to probe Unruh-Hawking thermality induced by acceleration and spacetime curvature. Using the open quantum system approach, we study the dynamics of the ergotropy with respect to the Kubo-Martin-Schwinger temperature, spacetime boundary conditions, and dimensionality. It has been found that the asymptotic value of quantum work extraction is determined solely by the acceleration and curvature. The steady behavior in the two spacetimes can be unified to witness the global thermality which is independent of boundary conditions and dimensionality. From a local perspective, we investigate how the ergotropy evolves through different pathways as the battery gradually reaches the same thermal equilibrium state characterized by a certain KMS temperature. In dS spacetime, the evolution at large acceleration exhibits pronounced oscillations and differs from the fast thermal relaxation observed at low acceleration. Varying the boundary condition in AdS spacetime can improve the energy storage of the moving battery. When the dimension of AdS spacetime is increased, vacuum fluctuations can modestly amplify the ergotropy in the initial stage and facilitate the rapid thermalization. From the perspective of energy transfer, the relativistic quantum battery helps explore the thermal vacuum in curved spacetimes.

hep-th

AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation

We propose AugCodec, a low-bitrate disentangled neural speech codec that leverages data augmentation to decompose speech into three distinct components: semantic, speaker, and prosody tokens. Specifically, we employ tailored augmenta tion strategies to transform speech into distinct variants, each serving as input for extracting tokens that preserve the target attribute while suppressing others. This disentanglement strategy enables substantial reduction in token rate. Further more, we introduce an augmentation loss that aligns semantic encoder outputs between source and voice-converted speech, encouraging speaker-agnostic embeddings while mitigating the acoustic mismatch induced by voice conversion. Experiments on LibriSpeech test-clean demonstrate that AugCodec significantly outperforms state-of-the-art methods in both reconstruction quality and disentanglement, while operating at only 12.5Hz with three token streams.

cs.SD

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding

Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens or hundreds of tokens, the limited context length of large language models (LLMs) forces the VLMs to perceive the frames sparsely and lose temporal information. To address this, we explore extreme video token compression towards one token per frame at the final LLM layer. Our key insight is that heuristic-based compression, widely adopted by previous methods, is prone to information loss, and this necessitates supervising LLM layers into learnable and progressive modules for token-level compression (LP-Comp). Such compression enables our VLM to digest 2x-4x more frames with improved performance. To further increase the token efficiency, we investigate frame-level compression, which selects the frames most relevant to the queries via the internal attention scores of the LLM layers, named question-conditioned compression (QC-Comp). As a notable distinction from previous studies, we mitigate the position bias of LLM attention in long contexts, i.e., the over-concentration on the beginning and end of a sequence, by splitting long videos into short segments and employing local attention. Collectively, our combined token-level and frame-level leads to an extreme compression model for long video understanding, named XComp, achieving a significantly larger compression ratio and enabling denser frame sampling. Our XComp is finetuned from VideoChat-Flash with a data-efficient supervised compression tuning stage that only requires 2.5% of the supervised fine-tuning data, yet boosts the accuracy from 42.9% to 46.2% on LVBench and enhances multiple other long video benchmarks.

cs.CV

Labeled TrustSet Guided: Batch Active Learning with Reinforcement Learning

Batch active learning (BAL) is a crucial technique for reducing labeling costs and improving data efficiency in training large-scale deep learning models. Traditional BAL methods often rely on metrics like Mahalanobis Distance to balance uncertainty and diversity when selecting data for annotation. However, these methods predominantly focus on the distribution of unlabeled data and fail to leverage feedback from labeled data or the model's performance. To address these limitations, we introduce TrustSet, a novel approach that selects the most informative data from the labeled dataset, ensuring a balanced class distribution to mitigate the long-tail problem. Unlike CoreSet, which focuses on maintaining the overall data distribution, TrustSet optimizes the model's performance by pruning redundant data and using label information to refine the selection process. To extend the benefits of TrustSet to the unlabeled pool, we propose a reinforcement learning (RL)-based sampling policy that approximates the selection of high-quality TrustSet candidates from the unlabeled data. Combining TrustSet and RL, we introduce the Batch Reinforcement Active Learning with TrustSet (BRAL-T) framework. BRAL-T achieves state-of-the-art results across 10 image classification benchmarks and 2 active fine-tuning tasks, demonstrating its effectiveness and efficiency in various domains.

cs.LG

Full-Field Metasurface Characterization with Polarization Sensitive Coherent Modulation Imaging

Characterizing the intensity, phase, and polarization of engineered light is fundamental to understanding and applying metasurfaces. However, existing characterization frameworks are hindered by several limitations, most notably their inability to account for the polarization of the field. Here, we report polarization sensitive coherent modulation imaging (PS-CMI), a light-weight but robust, high-resolution platform for the full-field characterization of metasurface-modulated light. By supplementing the orthogonal x- and y- complex amplitude components with an additional 45°-component, this approach calculates the retardance between two orthogonal polarization components while eliminating phase offsets, thereby enabling the subsequent recovery of the complete polarization state. We demonstrate the versatility of our method by characterizing light fields produced by a United States Air Force (USAF) target, two kinds of complex polarization field, and a metalens. This compact solution addresses a critical gap in metasurface metrology and is broadly applicable to other fields requiring the mapping of complex, polarized light distributions.

physics.optics

CSPR-Net: Self-supervised Curved Surface Projection Rectification Network for Geometric Distortion Correction in Non-planar Projections

Projecting images onto non-planar surfaces inevitably introduces geometric distortions that degrade visual quality. Traditional correction methods often require tedious manual calibration or structured light sequences to establish pixel-wise correspondences. In this paper, we develop the Curved Surface Projection Rectification Network (CSPR-Net), a self-supervised deep learning framework for automated distortion correction. Our approach employs dual coordinate-based neural networks to learn the bi-directional mapping between the projector and camera spaces. By enforcing a robust cycle-consistency constraint, CSPR-Net autonomously resolves complex geometric transformations without requiring ground-truth deformation fields. Furthermore, a gradient-based loss function is introduced to mitigate the impact of complex ambient light interference and accurately capture high-frequency geometric variations. Quantitative evaluations in physical experimental scenarios demonstrate that CSPR-Net achieves a 20.7% improvement in end-to-end fidelity (SSIM) and outperforms the polynomial baseline by 3.8% and 5.4% in forward and inverse mapping in terms of SSIM respectively, effectively generating high-precision pre-warped images for seamless projection.

physics.optics

Coherent quantum work extraction of a relativistic battery as a probe for acceleration-induced Unruh thermality

We propose a physical scheme of a uniformly accelerated Unruh-DeWitt battery and utilize quantum work extraction as a probe to witness the thermal nature of the Unruh effect induced by the accelerated motion. By employing the open quantum system approach, we analyze the coherent and incoherent components of the ergotropy which is the maximum amount of quantum work extraction of the relativistic battery driven by a coherent field. It has been proved that the steady values of coherent quantum work extraction in the asymptotic condition is only determined by the acceleration-dependent Unruh temperature. The asymptotic behavior of coherent ergotropy can demonstrate the thermal nature of the Unruh effect with respect to the Kubo-Martin-Schwinger condition. Under the circumstance of the Unruh effect, we explore the effect of the phase of the coherent charging field on the dynamics of coherent ergotropy when the battery approaches to the same thermal equilibrium state. The variation in the phase of the coherent driving field can improve the energy storage capacity of a relativistic battery. From viewpoint of energy transfer, the relativistic battery is helpful to examine the Unruh thermality.

gr-qc

RefTok: Reference-Based Tokenization for Video Generation

Effectively handling temporal redundancy remains a key challenge in learning video models. Prevailing approaches often treat each set of frames independently, failing to effectively capture the temporal dependencies and redundancies inherent in videos. To address this limitation, we introduce RefTok, a novel reference-based tokenization method capable of capturing complex temporal dynamics and contextual information. Our method encodes and decodes sets of frames conditioned on an unquantized reference frame. When decoded, RefTok preserves the continuity of motion and the appearance of objects across frames. For example, RefTok retains facial details despite head motion, reconstructs text correctly, preserves small patterns, and maintains the legibility of handwriting from the context. Across 4 video datasets (K600, UCF-101, BAIR Robot Pushing, and DAVIS), RefTok significantly outperforms current state-of-the-art tokenizers (Cosmos and MAGVIT) and improves all evaluated metrics (PSNR, SSIM, LPIPS) by an average of 36.7% at the same or higher compression ratios. When a video generation model is trained using RefTok's latents on the BAIR Robot Pushing task, the generations not only outperform MAGVIT-B but the larger MAGVIT-L, which has 4x more parameters, across all generation metrics by an average of 27.9%.

cs.CV

Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate

Most neural speech codecs achieve bitrate adjustment through intra-frame mechanisms, such as codebook dropout, at a Constant Frame Rate (CFR). However, speech segments inherently have time-varying information density (e.g., silent intervals versus voiced regions). This property makes CFR not optimal in terms of bitrate and token sequence length, hindering efficiency in real-time applications. In this work, we propose a Temporally Flexible Coding (TFC) technique, introducing variable frame rate (VFR) into neural speech codecs for the first time. TFC enables seamlessly tunable average frame rates and dynamically allocates frame rates based on temporal entropy. Experimental results show that a codec with TFC achieves optimal reconstruction quality with high flexibility, and maintains competitive performance even at lower frame rates. Our approach is promising for the integration with other efforts to develop low-frame-rate neural speech codecs for more efficient downstream tasks.

eess.AS

Quantum work extraction of a moving battery as a witness to Unruh thermality in high-dimensional spacetimes

We put forward a physical model of a uniformly accelerated Unruh-DeWitt battery and use quantum work extraction as a probe to witness the thermal nature of the Unruh effect in a high dimensional Minkowski spacetime. By means of the open quantum system approach, we investigate the maximal amount of quantum work extraction with respect to the acceleration-induced Unruh temperature, spacetime dimensionality and field mass. It has been found that the steady amount of quantum work extraction in the asymptotic condition is just determined by the Unruh temperature in arbitrary dimensional spacetimes. The asymptotic behavior can demonstrate the global feature of Unruh thermality dependent on the Kubo-Martin-Schwinger condition. From a local viewpoint of Unruh effect, we study the different ways for the dynamics of quantum work extraction when the battery gradually arrives at the same steady state. In the massless scalar field, the evolution with a small acceleration takes on a unique monotonicity in $D=3$ dimensional spacetime and changes to a decaying oscillation for other higher dimensions. The increase in spacetime dimensionality can increase the energy storage capacity of the moving battery. If the mass of the scalar field is considered, the related quantum work extraction is so robust against the Unruh decoherence that the high values can keep for a very long time. The persistence of quantum work extraction is strengthened in higher dimensional spacetime.

hep-th

NowYouSee Me: Context-Aware Automatic Audio Description

Audio Description (AD) plays a pivotal role as an application system aimed at guaranteeing accessibility in multimedia content, which provides additional narrations at suitable intervals to describe visual elements, catering specifically to the needs of visually impaired audiences. In this paper, we introduce $\mathrm{CA^3D}$, the pioneering unified Context-Aware Automatic Audio Description system that provides AD event scripts with precise locations in the long cinematic content. Specifically, $\mathrm{CA^3D}$ system consists of: 1) a Temporal Feature Enhancement Module to efficiently capture longer term dependencies, 2) an anchor-based AD event detector with feature suppression module that localizes the AD events and extracts discriminative feature for AD generation, and 3) a self-refinement module that leverages the generated output to tweak AD event boundaries from coarse to fine. Unlike conventional methods which rely on metadata and ground truth AD timestamp for AD detection and generation tasks, the proposed $\mathrm{CA^3D}$ is the first end-to-end trainable system that only uses visual cue. Extensive experiments demonstrate that the proposed $\mathrm{CA^3D}$ improves existing architectures for both AD event detection and script generation metrics, establishing the new state-of-the-art performances in the AD automation.

cs.CV

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning

In various video-language learning tasks, the challenge of achieving cross-modality alignment with multi-grained data persists. We propose a method to tackle this challenge from two crucial perspectives: data and modeling. Given the absence of a multi-grained video-text pretraining dataset, we introduce a Granularity EXpansion (GEX) method with Integration and Compression operations to expand the granularity of a single-grained dataset. To better model multi-grained data, we introduce an Iterative Approximation Module (IAM), which embeds multi-grained videos and texts into a unified, low-dimensional semantic space while preserving essential information for cross-modal alignment. Furthermore, GEXIA is highly scalable with no restrictions on the number of video-text granularities for alignment. We evaluate our work on three categories of video tasks across seven benchmark datasets, showcasing state-of-the-art or comparable performance. Remarkably, our model excels in tasks involving long-form video understanding, even though the pretraining dataset only contains short video clips.

cs.CV

Fast generation of arbitrary optical focus array

We report a novel method to generate arbitrary optical focus arrays (OFAs). Our approach rapidly produces computer-generated holograms (CGHs) to precisely control the positions and the intensities of the foci. This is achieved by replacing the fast Fourier transform (FFT) operation in the conventional iterative Fourier-transform algorithm (IFTA) with a linear algebra one, identifying/removing zero elements from the matrices, and employing a generalized weighting strategy. On the premise of accelerating the calculation speed by >70 times, we demonstrate OFA with 99% intensity precision in the experiment. Our method proves effective and is applicable for the systems in which real-time OFA generation is essential.

physics.optics

Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Humans can easily isolate a single speaker from a complex acoustic environment, a capability referred to as the "Cocktail Party Effect." However, replicating this ability has been a significant challenge in the field of target speaker extraction (TSE). Traditional TSE approaches predominantly rely on voiceprints, which raise privacy concerns and face issues related to the quality and availability of enrollment samples, as well as intra-speaker variability. To address these issues, this work introduces a novel text-guided TSE paradigm named LLM-TSE. In this paradigm, a state-of-the-art large language model, LLaMA 2, processes typed text input from users to extract semantic cues. We demonstrate that textual descriptions alone can effectively serve as cues for extraction, thus addressing privacy concerns and reducing dependency on voiceprints. Furthermore, our approach offers flexibility by allowing the user to specify the extraction or suppression of a speaker and enhances robustness against intra-speaker variability by incorporating context-dependent textual information. Experimental results show competitive performance with text-based cues alone and demonstrate the effectiveness of using text as a task selector. Additionally, they achieve a new state-of-the-art when combining text-based cues with pre-registered cues. This work represents the first integration of LLMs with TSE, potentially establishing a new benchmark in solving the cocktail party problem and expanding the scope of TSE applications by providing a versatile, privacy-conscious solution.

eess.AS

Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNet

Speech enhancement is critical for improving speech intelligibility and quality in various audio devices. In recent years, deep learning-based methods have significantly improved speech enhancement performance, but they often come with a high computational cost, which is prohibitive for a large number of edge devices, such as headsets and hearing aids. This work proposes an ultra-low-power speech enhancement system based on the brain-inspired spiking neural network (SNN) called Spiking-FullSubNet. Spiking-FullSubNet follows a full-band and sub-band fusioned approach to effectively capture both global and local spectral information. To enhance the efficiency of computationally expensive sub-band modeling, we introduce a frequency partitioning method inspired by the sensitivity profile of the human peripheral auditory system. Furthermore, we introduce a novel spiking neuron model that can dynamically control the input information integration and forgetting, enhancing the multi-scale temporal processing capability of SNN, which is critical for speech denoising. Experiments conducted on the recent Intel Neuromorphic Deep Noise Suppression (N-DNS) Challenge dataset show that the Spiking-FullSubNet surpasses state-of-the-art methods by large margins in terms of both speech quality and energy efficiency metrics. Notably, our system won the championship of the Intel N-DNS Challenge (Algorithmic Track), opening up a myriad of opportunities for ultra-low-power speech enhancement at the edge. Our source code and model checkpoints are publicly available at https://github.com/haoxiangsnr/spiking-fullsubnet.

eess.AS

In situ fully vectorial tomography and pupil function retrieval of tightly focused fields

Tightly focused optical fields are essential in nano-optics, but their applications have been limited by the challenges of accurate yet efficient characterization. In this article, we develop an in situ method for reconstructing the fully vectorial information of tightly focused fields in three-dimensional (3D) space, while simultaneously retrieving the pupil functions. Our approach encodes these fields using phase-modulated focusing and polarization-split detection, followed by decoding through an algorithm based on least-sampling matrix-based Fourier transform and analytically derived gradient. We further employ a focus scanning strategy. When combined with our decoding algorithm, this strategy mitigates the imperfections in the detection path. This approach requires only 10 frames of 2D measurements to realize approximate 90% accuracy in tomography and pupil function retrieval within 10s. Thus, it serves as a robust and convenient tool for the precise characterization and optimization of light at the nanoscale. We apply this technique to fully vectorial field manipulation, adaptive-optics-assisted nanoscopy, and addressing mixed-state problems.

physics.optics

Quantum coherence effects on inelastic thermoelectric devices: From diodes to transistors

We present a study on inelastic thermoelectric devices, wherein charge currents and electronic and phononic heat currents are intricately interconnected. The employment of double quantum dots in conjunction with a phonon bath positions them as promising candidates for quantum thermoelectric diodes and transistors. Within this study, we illustrate that quantum coherence effects yield significant charge and Seebeck rectification effects. It's worth noting that, while the thermal transistor effect is observable in the linear response regime, especially when phonon-assisted inelastic processes dominate the transport, quantum coherence does not enhance thermal amplification. Our work provides valuable insights for the optimization of general thermoelectric devices.

cond-mat.mes-hall