arXiv ScienceSearch

arXiv subjects

Min Hu

Publications and source records attributed to Min Hu.

At least 19 recordsLinked to original sources

HUSH-Bench: Measuring Memory-Use Boundaries for Sensitive History in Conversational Agents

Long-term memory helps conversational agents maintain continuity across sessions, while relevance and current-turn warrant remain distinct decisions. We study this boundary under a stated conservative policy in which sensitive history shapes a response when the current turn supplies a reason to use it. We introduce HUSH-Bench, a controlled benchmark of 2,400 benign prompts paired with histories containing one marked sensitive disclosure and matched no-memory references. HUSH-Bench measures unsolicited history integration with the Unsolicited History Integration Score (UIS; 0--100, higher is worse), records whether the marked disclosure reaches the generator, and includes paired prompts that differ only in whether the user asks the assistant to use earlier context. We evaluate four models under no-memory, full-context, and three retrieval-based memory settings. Memory access raises UIS from near zero to 8.9--26.6 for one model and 51.3--83.0 for the other three. Retrieval systems expose the marked disclosure in 23.0\%--30.3\% of cases, while related sensitive entries or summaries remain available and three models continue to show high UIS. Across four generators, an explicit invitation increases target-memory uptake scores by 27.0--41.3; measured helpfulness remains stable while mean over-scope rises. These results motivate treating memory storage, retrieval, warrant, and per-turn scope as separate design decisions.

cs.AI

Energy threshold in Smith-Purcell radiation

Smith Purcell radiation has emerged as a crucial platform for investigating light-matter interactions and developing compact, tunable light sources that span from microwaves to X-rays. In classical theory, it is believed that Cherenkov radiation exhibits an energy threshold for electrons, while Smith Purcell radiation is considered free of such a threshold. Although quantum theory suggests there is an emission cutoff in Smith-Purcell radiation, the behavior of this radiation near the threshold remains understudied. In this article, we address this gap by examining the behavior of Smith-Purcell radiation near the threshold from quantum perspectives. Specifically, we derive a quantum energy threshold based on energy-momentum conservation, providing a rigorous limit for the onset of Smith Purcell radiation. Furthermore, we find that around the threshold the incident electron emits a photon and subsequently reverses its direction of motion. Additionally, we establish a classical energy threshold below which the classical theory breakdown by applying the Duane Hunt limit to Smith Purcell radiation. Accordingly, quantum theory is required when the electron energy falls between the classical and quantum thresholds. Our findings enrich the understanding of Smith Purcell radiation and provide valuable insights for developing low energy driven and heralded quantum light sources.

physics.optics

Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model

We present daVinci-MagiHuman, an open-source audio-video generative foundation model for human-centric generation. daVinci-MagiHuman jointly generates synchronized video and audio using a single-stream Transformer that processes text, video, and audio within a unified token sequence via self-attention only. This single-stream design avoids the complexity of multi-stream or cross-attention architectures while remaining easy to optimize with standard training and inference infrastructure. The model is particularly strong in human-centric scenarios, producing expressive facial performance, natural speech-expression coordination, realistic body motion, and precise audio-video synchronization. It supports multilingual spoken generation across Chinese (Mandarin and Cantonese), English, Japanese, Korean, German, and French. For efficient inference, we combine the single-stream backbone with model distillation, latent-space super-resolution, and a Turbo VAE decoder, enabling generation of a 5-second 256p video in 2 seconds on a single H100 GPU. In automatic evaluation, daVinci-MagiHuman achieves the highest visual quality and text alignment among leading open models, along with the lowest word error rate (14.60%) for speech intelligibility. In pairwise human evaluation, it achieves win rates of 80.0% against Ovi 1.1 and 60.9% against LTX 2.3 over 2000 comparisons. We open-source the complete model stack, including the base model, the distilled model, the super-resolution model, and the inference codebase.

cs.CV

Enhancing Zero-Shot Time Series Forecasting in Off-the-Shelf LLMs via Noise Injection

Large Language Models (LLMs) have demonstrated effectiveness as zero-shot time series (TS) forecasters. The key challenge lies in tokenizing TS data into textual representations that align with LLMs' pre-trained knowledge. While existing work often relies on fine-tuning specialized modules to bridge this gap, a distinct, yet challenging, paradigm aims to leverage truly off-the-shelf LLMs without any fine-tuning whatsoever, relying solely on strategic tokenization of numerical sequences. The performance of these fully frozen models is acutely sensitive to the textual representation of the input data, as their parameters cannot adapt to distribution shifts. In this paper, we introduce a simple yet highly effective strategy to overcome this brittleness: injecting noise into the raw time series before tokenization. This non-invasive intervention acts as a form of inference-time augmentation, compelling the frozen LLM to extrapolate based on robust underlying temporal patterns rather than superficial numerical artifacts. We theoretically analyze this phenomenon and empirically validate its effectiveness across diverse benchmarks. Notably, to fully eliminate potential biases from data contamination during LLM pre-training, we introduce two novel TS datasets that fall outside all utilized LLMs' pre-training scopes, and consistently observe improved performance. This study provides a further step in directly leveraging off-the-shelf LLMs for time series forecasting.

cs.AI

MAGI-1: Autoregressive Video Generation at Scale

We present MAGI-1, a world model that generates videos by autoregressively predicting a sequence of video chunks, defined as fixed-length segments of consecutive frames. Trained to denoise per-chunk noise that increases monotonically over time, MAGI-1 enables causal temporal modeling and naturally supports streaming generation. It achieves strong performance on image-to-video (I2V) tasks conditioned on text instructions, providing high temporal consistency and scalability, which are made possible by several algorithmic innovations and a dedicated infrastructure stack. MAGI-1 facilitates controllable generation via chunk-wise prompting and supports real-time, memory-efficient deployment by maintaining constant peak inference cost, regardless of video length. The largest variant of MAGI-1 comprises 24 billion parameters and supports context lengths of up to 4 million tokens, demonstrating the scalability and robustness of our approach. The code and models are available at https://github.com/SandAI-org/MAGI-1 and https://github.com/SandAI-org/MagiAttention. The product can be accessed at https://sand.ai.

cs.CV

Dual-Splitting Conformal Prediction for Multi-Step Time Series Forecasting

Time series forecasting is crucial for applications like resource scheduling and risk management, where multi-step predictions provide a comprehensive view of future trends. Uncertainty Quantification (UQ) is a mainstream approach for addressing forecasting uncertainties, with Conformal Prediction (CP) gaining attention due to its model-agnostic nature and statistical guarantees. However, most variants of CP are designed for single-step predictions and face challenges in multi-step scenarios, such as reliance on real-time data and limited scalability. This highlights the need for CP methods specifically tailored to multi-step forecasting. We propose the Dual-Splitting Conformal Prediction (DSCP) method, a novel CP approach designed to capture inherent dependencies within time-series data for multi-step forecasting. Experimental results on real-world datasets from four different domains demonstrate that the proposed DSCP significantly outperforms existing CP variants in terms of the Winkler Score, achieving a performance improvement of up to 23.59% compared to state-of-the-art methods. Furthermore, we deployed the DSCP approach for renewable energy generation and IT load forecasting in power management of a real-world trajectory-based application, achieving an 11.25% reduction in carbon emissions through predictive optimization of data center operations and controls.

cs.LG

A compact unshielded optically-pumped magnetic gradiometer

Optically-pumped magnetic gradiometers (OPGs) play a crucial role in applications such as magnetic anomaly detection and bio-magnetic measurements. This study classifies current OPGs into four types based on their differential modes: voltage, frequency, optical rotation, and magnetic field differential modes. We introduce the concept of inherent Common-Mode Rejection Ratio (CMRR) and analyze the differences between the inherent CMRR and the measured CMRR, as well as the upper limit of inherent CMRR. We point out that although magnetic field differential method has the potential to increase inherent CMRR by a factor of 1+AF, the difference between the feedback gains is often neglected, which may set the limit of inherent CMRR. We designed and fabricated a compact, unshielded OPG with a specially designed scheme to minimize the distance between the sensing heads and the magnetic source. Measurement results demonstrate a measured CMRR of 1200@1Hz and a sensitivity of approximately 5 pT/cm/\sqrt{Hz} from 1 Hz to 100 Hz.

physics.optics

ReverseNER: A Self-Generated Example-Driven Framework for Zero-Shot Named Entity Recognition with Large Language Models

This paper presents ReverseNER, a method aimed at overcoming the limitation of large language models (LLMs) in zero-shot named entity recognition (NER) tasks, arising from their reliance on pre-provided demonstrations. ReverseNER tackles this challenge by constructing a reliable example library composed of dozens of entity-labeled sentences, generated through the reverse process of NER. Specifically, while conventional NER methods label entities in a sentence, ReverseNER features reversing the process by using an LLM to generate entities from their definitions and subsequently expand them into full sentences. During the entity expansion process, the LLM is guided to generate sentences by replicating the structures of a set of specific \textsl{feature sentences}, extracted from the task sentences by clustering. This expansion process produces dozens of entity-labeled task-relevant sentences. After constructing the example library, the method selects several semantically similar entity-labeled examples for each task sentence as references to facilitate the LLM's entity recognition. We also propose an entity-level self-consistency scoring mechanism to improve NER performance with LLMs. Experiments show that ReverseNER significantly outperforms other zero-shot NER methods with LLMs, marking a notable improvement in NER for domains without labeled data, while declining computational resource consumption.

cs.CL

Enhancing heat transfer in X-ray tube by van der heterostructures-based thermionic emission

Van der Waals (vdW) heterostructures have attracted much attention due to their distinctive optical, electrical, and thermal properties, demonstrating promising potential in areas such as photocatalysis, ultrafast photonics, and free electron radiation devices. Particularly, they are promising platforms for studying thermionic emission. Here, we illustrate that using vdW heterostructure-based thermionic emission can enhance heat transfer in vacuum devices. As a proof of concept, we demonstrate that this approach offers a promising solution to the long-standing overheating issue in X-ray tubes. Specifically, we show that the saturated target temperature of a 2000 W X-ray tube can be reduced from around 1200 celsius to 490 celsius. Additionally, our study demonstrates that by reducing the height of the Schottky barrier formed in the vdW heterostructures, the thermionic cooling performance can be enhanced. Our findings pave the way for the development of high-power X-ray tubes.

physics.app-ph

Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering

Audio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are prone to overlearning dataset biases, resulting in poor robustness. Furthermore, current datasets may not provide a precise diagnostic for these methods. To tackle these challenges, firstly, we propose a novel dataset, MUSIC-AVQA-R, crafted in two steps: rephrasing questions within the test split of a public dataset (MUSIC-AVQA) and subsequently introducing distribution shifts to split questions. The former leads to a large, diverse test space, while the latter results in a comprehensive robustness evaluation on rare, frequent, and overall questions. Secondly, we propose a robust architecture that utilizes a multifaceted cycle collaborative debiasing strategy to overcome bias learning. Experimental results show that this architecture achieves state-of-the-art performance on MUSIC-AVQA-R, notably obtaining a significant improvement of 9.32%. Extensive ablation experiments are conducted on the two datasets mentioned to analyze the component effectiveness within the debiasing strategy. Additionally, we highlight the limited robustness of existing multi-modal QA methods through the evaluation on our dataset. We also conduct experiments combining various baselines with our proposed strategy on two datasets to verify its plug-and-play capability. Our dataset and code are available at https://github.com/reml-group/MUSIC-AVQA-R.

cs.CV

AR Visualization System for Ship Detection and Recognition Based on AI

Augmented reality technology has been widely used in industrial design interaction, exhibition guide, information retrieval and other fields. The combination of artificial intelligence and augmented reality technology has also become a future development trend. This project is an AR visualization system for ship detection and recognition based on AI, which mainly includes three parts: artificial intelligence module, Unity development module and Hololens2AR module. This project is based on R3Det algorithm to complete the detection and recognition of ships in remote sensing images. The recognition rate of model detection trained on RTX 2080Ti can reach 96%. Then, the 3D model of the ship is obtained by ship categories and information and generated in the virtual scene. At the same time, voice module and UI interaction module are added. Finally, we completed the deployment of the project on Hololens2 through MRTK. The system realizes the fusion of computer vision and augmented reality technology, which maps the results of object detection to the AR field, and makes a brave step toward the future technological trend and intelligent application.

cs.CV

A general approach to improve the bias stability of NMR gyroscope

In recent years, progress in improving the bias stability of NMR gyroscopes has been hindered. Taking inspiration from the core idea of rotation modulation in the strapdown inertial navigation system, we propose a general approach to enhancing the bias stability of NMR gyroscopes that does not require consideration of the actual physical sources. The method operates on the fact that the sign of the bias does not follow that of the sensing direction of the NMR gyroscope, which is much easier to modulate than with other types of gyroscopes. We conducted simulations to validate the method's feasibility.

physics.ins-det

Adaptive loose optimization for robust question answering

Question answering methods are well-known for leveraging data bias, such as the language prior in visual question answering and the position bias in machine reading comprehension (extractive question answering). Current debiasing methods often come at the cost of significant in-distribution performance to achieve favorable out-of-distribution generalizability, while non-debiasing methods sacrifice a considerable amount of out-of-distribution performance in order to obtain high in-distribution performance. Therefore, it is challenging for them to deal with the complicated changing real-world situations. In this paper, we propose a simple yet effective novel loss function with adaptive loose optimization, which seeks to make the best of both worlds for question answering. Our main technical contribution is to reduce the loss adaptively according to the ratio between the previous and current optimization state on mini-batch training data. This loose optimization can be used to prevent non-debiasing methods from overlearning data bias while enabling debiasing methods to maintain slight bias learning. Experiments on the visual question answering datasets, including VQA v2, VQA-CP v1, VQA-CP v2, GQA-OOD, and the extractive question answering dataset SQuAD demonstrate that our approach enables QA methods to obtain state-of-the-art in- and out-of-distribution performance in most cases. The source code has been released publicly in \url{https://github.com/reml-group/ALO}.

cs.CL

Futures Quantitative Investment with Heterogeneous Continual Graph Neural Network

This study aims to address the challenges of futures price prediction in high-frequency trading (HFT) by proposing a continuous learning factor predictor based on graph neural networks. The model integrates multi-factor pricing theories with real-time market dynamics, effectively bypassing the limitations of existing methods that lack financial theory guidance and ignore various trend signals and their interactions. We propose three heterogeneous tasks, including price moving average regression, price gap regression and change-point detection to trace the short-, intermediate-, and long-term trend factors present in the data. In addition, this study also considers the cross-sectional correlation characteristics of future contracts, where prices of different futures often show strong dynamic correlations. Each variable (future contract) depends not only on its historical values (temporal) but also on the observation of other variables (cross-sectional). To capture these dynamic relationships more accurately, we resort to the spatio-temporal graph neural network (STGNN) to enhance the predictive power of the model. The model employs a continuous learning strategy to simultaneously consider these tasks (factors). Additionally, due to the heterogeneity of the tasks, we propose to calculate parameter importance with mutual information between original observations and the extracted features to mitigate the catastrophic forgetting (CF) problem. Empirical tests on 49 commodity futures in China's futures market demonstrate that the proposed model outperforms other state-of-the-art models in terms of prediction accuracy. Not only does this research promote the integration of financial theory and deep learning, but it also provides a scientific basis for actual trading decisions.

cs.LG

Speaker Change Detection for Transformer Transducer ASR

Speaker change detection (SCD) is an important feature that improves the readability of the recognized words from an automatic speech recognition (ASR) system by breaking the word sequence into paragraphs at speaker change points. Existing SCD solutions either require additional ensemble for the time based decisions and recognized word sequences, or implement a tight integration between ASR and SCD, limiting the potential optimum performance for both tasks. To address these issues, we propose a novel framework for the SCD task, where an additional SCD module is built on top of an existing Transformer Transducer ASR (TT-ASR) network. Two variants of the SCD network are explored in this framework that naturally estimate speaker change probability for each word, while allowing the ASR and SCD to have independent optimization scheme for the best performance. Experiments show that our methods can significantly improve the F1 score on LibriCSS and Microsoft call center data sets without ASR degradation, compared with a joint SCD and ASR baseline.

eess.AS

Ultrahigh ion diffusion in oxide crystal by engineering the interfacial transporter channels

The mass storage and removal in solid conductors always played vital role on the technological applications such as modern batteries, permeation membranes and neuronal computations, which were seriously lying on the ion diffusion and kinetics in bulk lattice. However, the ions transport was kinetically limited by the low diffusional process, which made it a challenge to fabricate applicable conductors with high electronic and ionic conductivities at room temperature. It was known that at essentially all interfaces, the existed space charge layers could modify the charge transport, storage and transfer properties. Thus, in the current study, we proposed an acid solution/WO3/ITO structure and achieved an ultrafast hydrogen transport in WO3 layer by interfacial job-sharing diffusion. In this sandwich structure, the transport pathways of the protons and electrons were spatially separated in acid solution and ITO layer respectively, resulting the pronounced increasing of effective hydrogen diffusion coefficient (Deff) up to 106 times. The experiment and theory simulations also revealed that this accelerated hydrogen transport based on the interfacial job-sharing diffusion was universal and could be extended to other ions and oxide materials as well, which would potentially stimulate systematic studies on ultrafast mixed conductors or faster solid-state electrochemical switching devices in the future.

cond-mat.mtrl-sci

Robust Anomaly Detection for Time-series Data

Time-series anomaly detection plays a vital role in monitoring complex operation conditions. However, the detection accuracy of existing approaches is heavily influenced by pattern distribution, existence of multiple normal patterns, dynamical features representation, and parameter settings. For the purpose of improving the robustness and guaranteeing the accuracy, this research combined the strengths of negative selection, unthresholded recurrence plots, and an extreme learning machine autoencoder and then proposed robust anomaly detection for time-series data (RADTD), which can automatically learn dynamical features in time series and recognize anomalies with low label dependency and high robustness. Yahoo benchmark datasets and three tunneling engineering simulation experiments were used to evaluate the performance of RADTD. The experiments showed that in benchmark datasets RADTD possessed higher accuracy and robustness than recurrence qualification analysis and extreme learning machine autoencoder, respectively, and that RADTD accurately detected the occurrence of tunneling settlement accidents, indicating its remarkable performance in accuracy and robustness.

cs.LG

Atomic origin for hydrogenation promoted bulk oxygen vacancies removal in vanadium dioxide

Oxygen vacancies (VO), a common type of point defects in metal oxides materials, play important roles on the physical and chemical properties. To obtain stoichiometric oxide crystal, the pre-existing VO is always removed via careful post-annealing treatment at high temperature in air or oxygen atmosphere. However, the annealing conditions is difficult to control and the removal of VO in bulk phase is restrained due to high energy barrier of VO migration. Here, we selected VO2 crystal film as the model system and developed an alternative annealing treatment aided by controllable hydrogen doping, which can realizes effective removal of VO defects in VO2-{\delta} crystal at lower temperature. This finding is attributed to the hydrogenation accelerated oxygen vacancies recovery in VO2-{\delta} crystal. Theoretical calculations revealed that the H-doping induced electrons are prone to accumulate around the oxygen defects in VO2-{\delta} film, which facilitates the diffusion of VO and thus makes it easier to be removed. The methodology is expected to be applied to other metal oxides for oxygen-related point defects control.

cond-mat.mtrl-sci