arXiv ScienceSearch

arXiv subjects

Zheng Qin

Publications and source records attributed to Zheng Qin.

At least 19 recordsLinked to original sources

Full Inseparability and Genuine Multipartite Entanglement Coincide for Finite-Mode Gaussian States

For general mixed states, entanglement across every bipartition need not imply genuine multipartite entanglement (GME), because a biseparable decomposition may switch the separable cut from term to term. We prove that this convex ambiguity disappears for Gaussian states of finitely many bosonic modes. More generally, for any finite family of partitions, a Gaussian density operator in the trace-norm-closed convex class generated by states separable across those partitions is already separable across one fixed partition in the family. Only the target is Gaussian; a valid decomposition may be continuous and may contain arbitrary non-Gaussian states. Thus full inseparability and GME coincide, Gaussian k-separability and k-producibility reduce to fixed-partition tests, and party-wise tensor powers cannot activate GME from a biseparable Gaussian state. The proof combines a spectral selector with a holomorphic rigidity argument that converts one product vector in the square-root range of a Gaussian state into a block-local covariance certificate. The result shows that partition mixing, a generic mixed-state mechanism, adds no new exact finite-mode Gaussian states.

quant-ph

Large Scale Entanglement Structure Detection in 100-Qubit Systems via Local Joint Measurements

Identifying the entanglement structure of a many-body quantum state, namely how its constituents partition into unentangled blocks, is a central task in quantum information science, yet conventional tomography scales exponentially with system size. Here we introduce a scalable framework that recognizes large-scale entanglement structures directly from local correlation fingerprints. By choosing a representative local Pauli basis that satisfies a boundary-matching condition p_1 = p_R, the entire chain is read out in a single measurement configuration, keeping the measurement effort independent of system size. In noisy simulations, this single-basis protocol classifies GHZ-, W-, and cluster-type structures among 30 candidate partitions with a mean accuracy exceeding 95% for systems of up to 100 qubits. We further validate the protocol on a superconducting quantum processor, where it reliably classifies block structures for systems of up to 13 qubits before noise- and depth-induced degradation sets in at larger sizes. By mapping these failure modes explicitly, our results delineate the boundary of hardware-level scalability and point to a concrete strategy for characterizing entanglement structure on near-term quantum devices.

quant-ph

DynaMOMA: Instantaneous Prediction of Grasp Poses for Mobile Manipulation of Dynamic Objects

Mobile manipulation is a fundamental robotics task and has advanced rapidly in recent years, enabling robots to navigate, reach, and interact with objects in complex environments. However, mobile manipulation of dynamic objects remains highly challenging, as robots must coordinate the mobile base and arm while adapting to continuously evolving target poses. A key challenge lies in predicting temporally consistent short-horizon grasp trajectories from dynamic observations. In this work, we propose \ours{}, a dynamic mobile manipulation framework that couples instantaneous grasp trajectory prediction with whole-body control policy. Our predictor uses an anchor-based diffusion model to generate temporally consistent short-horizon grasp trajectories conditioned on historical observations. The predicted trajectories are then encoded as compact features and fed to a whole-body reinforcement learning policy, which controls the mobile manipulator for dynamic grasping. We further introduce a anticipation-guided reward that equips the policy with an anticipatory grasping horizon by adaptively shifting the target from the current grasp observation to the instantaneously predicted grasp trajectory. Through extensive experiments in Isaac Gym simulation, we show that our method achieves strong performance in mobile manipulation of dynamic objects across diverse settings and grasping metrics. Furthermore, our predictor and policy demonstrate strong generalizability in real-world experiments.

cs.RO

SAM-Flow: Source-Anchored Masked Flow for Training-Free Image Editing

Training-free image editing has recently attracted increasing attention due to its ability to modify real images using powerful pre-trained diffusion and flow-matching models without additional training. However, existing inversion-based and differential-flow-based methods usually perform global latent transport, which inevitably propagates editing effects to non-target regions and leads to background leakage. To address this problem, we propose SAM-Flow, a source-anchored masked flow framework for localized training-free image editing. Instead of updating the whole latent representation, SAM-Flow first uses a scout image and token-grounded attention maps to localize the editable semantic regions. It then applies differential velocity updates only within these regions, while anchoring the remaining areas to the source-image latent trajectory. To further improve spatial stability and boundary naturalness, we introduce a time-varying source-anchored projection mechanism with dynamic soft masks, transition regions, and temporal mask accumulation. The proposed method is plug-and-play and can be integrated with mainstream flow-matching backbones such as Stable Diffusion 3 and FLUX without any fine-tuning. Extensive qualitative and quantitative experiments demonstrate that SAM-Flow achieves accurate semantic editing while significantly improving background preservation, providing a simple and general localized editing paradigm for training-free image editing. Code is available at: https://github.com/chwbob/Sam-Flow.

cs.CV

The odd-parity altermagnetism induced reconstruction of the Chern-insulating phase in Haldane-Hubbard model

Odd-parity altermagnetism(ALM) extends compensated collinear magnetism beyond the even-parity spin splitting of conventional altermagnets, but its role in correlated topological phases remains largely unexplored. Using the cluster slave-spin method, we show that the odd-parity ALM appearing in the ALM Chern-insulating phase of Haldane-Hubbard model significantly reconstructs the local topology in the conventional Chern-insulating phase, while the total Chern number remains unchanged compared to the Chern-insulating phase. The Berry curvature becomes spin and valley selective; zigzag ribbons develop chiral-symmetry-breaking edge states; while armchair ribbons remain inversion symmetric. The optical response mirrors this separation between the local reconstruction and the global topology: low-energy spectra are governed by quasiparticles near the gap, whereas the low-frequency Hall conductivity stays quantized, $\sigma_{\rm T\uparrow}(\Omega\to 0)=\sigma_{\rm T\downarrow}(\Omega\to 0)=e^2/h$. These results establish the Haldane-Hubbard model as a minimal correlated platform for odd-parity altermagnetic topology.

cond-mat.str-el

Calibrating the Role of Entanglement in Variational Quantum Algorithms from a Geometric Perspective

Calibrating the role of entanglement in quantum algorithms is a crucial task in the development of quantum computing. Most existing studies have primarily focused on how the static properties of entanglement-such as its magnitude and phase-affect key performance metrics. In this work, we instead explore the relationship between the dynamical behaviors of entanglement and the execution of variational quantum algorithms from a geometric perspective. We find that, in contrast to conventional Hamiltonian dynamics where the evolution process is dominated by the dynamical phase, quantum state evolution in quantum algorithms is primarily governed by the geometric phase with the trajectory determined by the parameter-dependent Hilbert space geometry. In the problem-agnostic Hardware-Efficient Ansatz (HEA), entanglement dynamics and state evolution are decoupled. Conversely, in the problem-inspired Hamiltonian Variational Ansatz (HVA), the dynamical phase contribution is enhanced, allowing entanglement to function as a dynamical resource: more entanglement consumption correlates directly with faster quantum state evolution.

quant-ph

Chiral Superconductivity in Periodically Driven Altermagnet/Superconductor Heterostructures

The interplay between magnetism and superconductivity provides a fertile ground for engineering exotic topological phases, while dynamical control via periodic driving offers a unique avenue to access quantum states that are inaccessible in static equilibrium. Here, we propose a strategy to achieve the Floquet chiral topological superconductivity in an altermagnet-superconductor heterostructure driven by elliptically polarized light. We show that for $s$-wave pairing, the system undergoes a transition from a trivial to a chiral topological superconducting phase. More strikingly, with the introduction of mixed $s+d$-wave pairing, we find that the system can access Floquet chiral topological superconducting phases with highly tunable Chern numbers up to N=4. These exotic phases are attributed to the intertwining of altermagnetism, superconducting pairing, and the periodic driving field. Our work establishes the light-driven altermagnetic heterostructure as a versatile platform for exploring and manipulating high-Chern-number chiral topological superconductivity.

cond-mat.supr-con

A Modular, Data-Free Pipeline for Multi-Label Intention Recognition in Transportation Agentic AI Applications

In this study, a modular, data-free pipeline for multi-label intention recognition is proposed for agentic AI applications in transportation. Unlike traditional intent recognition systems that depend on large, annotated corpora and often struggle with fine-grained, multi-label discrimination, our approach eliminates the need for costly data collection while enhancing the accuracy of multi-label intention understanding. Specifically, the overall pipeline, named DMTC, consists of three steps: 1) using prompt engineering to guide large language models (LLMs) to generate diverse synthetic queries in different transport scenarios; 2) encoding each textual query with a Sentence-T5 model to obtain compact semantic embeddings; 3) training a lightweight classifier using a novel online focal-contrastive (OFC) loss that emphasizes hard samples and maximizes inter-class separability. The applicability of the proposed pipeline is demonstrated in an agentic AI application in the maritime transportation context. Extensive experiments show that DMTC achieves a Hamming loss of 5.35% and an AUC of 95.92%, outperforming state-of-the-art multi-label classifiers and recent end-to-end SOTA LLM-based baselines. Further analysis reveals that Sentence-T5 embeddings improve subset accuracy by at least 3.29% over alternative encoders, and integrating the OFC loss yields an additional 0.98% gain compared to standard contrastive objectives. In conclusion, our system seamlessly routes user queries to task-specific modules (e.g., ETA information, traffic risk evaluation, and other typical scenarios in the transportation domain), laying the groundwork for fully autonomous, intention-aware agents without costly manual labelling.

cs.LG

UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features

The task of synthesizing novel views from a single image is highly ill-posed due to multiple explanations for unobserved areas. Most current methods tend to generate unseen regions from ambiguity priors and interpolation near input views, which often lead to severe distortions. To address this limitation, we propose a novel model dubbed as UniView, which can leverage reference images from a similar object to provide strong prior information during view synthesis. More specifically, we construct a retrieval and augmentation system and employ a multimodal large language model (MLLM) to assist in selecting reference images that meet our requirements. Additionally, a plug-and-play adapter module with multi-level isolation layers is introduced to dynamically generate reference features for the target views. Moreover, in order to preserve the details of an original input image, we design a decoupled triple attention mechanism, which can effectively align and integrate multi-branch features into the synthesis process. Extensive experiments have demonstrated that our UniView significantly improves novel view synthesis performance and outperforms state-of-the-art methods on the challenging datasets.

cs.CV

Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion

Generating 3D human motions from text is a challenging yet valuable task. The key aspects of this task are ensuring text-motion consistency and achieving generation diversity. Although recent advancements have enabled the generation of precise and high-quality human motions from text, achieving diversity in the generated motions remains a significant challenge. In this paper, we aim to overcome the above challenge by designing a simple yet effective text-to-motion generation method, \textit{i.e.}, Diverse-T2M. Our method introduces uncertainty into the generation process, enabling the generation of highly diverse motions while preserving the semantic consistency of the text. Specifically, we propose a novel perspective that utilizes noise signals as carriers of diversity information in transformer-based methods, facilitating a explicit modeling of uncertainty. Moreover, we construct a latent space where text is projected into a continuous representation, instead of a rigid one-to-one mapping, and integrate a latent space sampler to introduce stochastic sampling into the generation process, thereby enhancing the diversity and uncertainty of the outputs. Our results on text-to-motion generation benchmark datasets~(HumanML3D and KIT-ML) demonstrate that our method significantly enhances diversity while maintaining state-of-the-art performance in text consistency.

cs.CV

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of empathetic, context-aware responses. Here we introduce HumanSense, a comprehensive benchmark designed to evaluate the human-centered perception and interaction capabilities of MLLMs, with a particular focus on deep understanding of extended multimodal contexts and the formulation of rational feedback. Our evaluation reveals that leading MLLMs still have considerable room for improvement, particularly for advanced interaction-oriented tasks. Supplementing visual input with audio and text information yields substantial improvements, and Omni-modal models show advantages on these tasks.Furthermore, grounded in the observation that appropriate feedback stems from a contextual analysis of the interlocutor's needs and emotions, we posit that reasoning ability serves as the key to unlocking it. We devise a multi-stage, modality-progressive reinforcement learning approach, resulting in HumanSense-Omni-Reasoning, which substantially enhances performance on higher-level understanding and interactive tasks. Additionally, we observe that successful reasoning processes appear to exhibit consistent thought patterns. By designing corresponding prompts, we also enhance the performance of non-reasoning models in a training-free manner.Project page: \textcolor{brightpink}{https://digital-avatar.github.io/ai/HumanSense/}

cs.CV

Quantum Spin Hall Effect with Extended Topologically Protected Features in Altermangetic Multilayers

Conventional topological classification theory dictates that time-reversal symmetry confines the quantum spin Hall (QSH) effect to a $\mathbb{Z}_2$ classification, permitting only a single pair of gapless helical edge states. Here, we utilize the recently discovered altermagnetism to circumvent this fundamental constraint. We demonstrate the realization of a unique QSH phase possessing multiple pairs of gapless helical edge states in altermagnetic multilayers. This exotic QSH phase, characterized by a mirror-spin Chern number, emerges from the interplay of spin-orbit coupling and $d$-wave altermagnetic ordering. Moreover, using first-principles calculations, we identify altermagnetic Fe$_2$Se$_2$O multilayers as promising material candidates, in which the number of gapless helical edge states scales linearly with the number of layers, leading to a correspondingly large, exactly quantized, and experimentally accessible spin-Hall conductance. Our findings unveil a new mechanism for stabilizing multiple pairs of gapless helical edge states, significantly expanding the scope of QSH effects, and provide a blueprint for utilizing altermagnetism to engineer desired topological phases.

cond-mat.mes-hall

Light-induced Odd-parity Magnetism in Conventional Collinear Antiferromagnets

Recent studies have drawn growing attention on non-relativistic odd-parity magnetism in the wake of altermagnets. Nevertheless, odd-parity spin splitting is often believed to appear in non-collinear magnetic configurations. Here, using symmetry arguments and effective model analysis, we show that Floquet engineering offers a universal strategy for achieving odd-parity magnetism in two-dimensional (2D) collinear antiferromagnets under irradiation of periodic driving light fields such as circularly polarized light, elliptically polarized light, and bicircular light. A comprehensive classification of potential candidates for collinear monolayer or bilayer antiferromagnets is established. Strikingly, the light-induced odd-parity spin splitting can be flexibly controlled by adjusting the crystalline symmetry or the polarization state of incident light, enabling the reversal or conversion of spin-splitting. By combining first-principles calculations and Floquet theorem, we present illustrative examples of 2D collinear antiferromagnetic (AFM) materials to verify the light-induced odd-parity magnetism. Our work not only offers a powerful approach for uniquely achieving odd-parity spin-splitting with high tunability, but also expands the potential of Floquet engineering in designing unconventional compensated magnetism.

cond-mat.mtrl-sci

The odd-parity altermagnetism: A spin group study

Following recent intensive studies on altermagnetism(ALM) characterized by non-relativistic even-parity spin splitting, realizing unconventional odd-parity magnetism has also attracted increasing interest. Here, using symmetry arguments based on spin-group analyses, we elucidate sufficient conditions for the emergence of odd-parity spin splitting in collinear antiferromagnetic systems, which is further established as the standard odd-parity ALM. It is derived that the odd-parity ALM arises from the following criteria: (i)the breaking nonmagnetic time reversal symmetry(TRS), i.e., the breaking real-space TRS; (ii)the long-range collinear compensated magnetism; (iii)the symmetry $[C_{2}||\bar{E}]$ or $[C_{2}||M]$ connecting opposite-spin sublattices, where $C_{2}$, $\bar{E}$, and $M$ respectively represent a $180^{\circ}$ rotation around the axis perpendicular to spins, the inversion, and the mirror reflection separating opposite-spin sublattices, directly reflecting the high-order harmonic($l\ge3$) and the $p$-wave($l=1$) odd-parity ALM, respectively. Moreover, we utilize the well-known Haldane-Hubbard model to identify odd-parity spin splitting in the collinear ALM ground state, where (i)the nonmagnetic TRS is broken by opposite sublattice currents coming from the Haldane hopping; (ii)the symmetry $[C_{2}||\bar{E}]$ is ensured because the currents flowing on opposite-spin sublattices are reversed.

cond-mat.str-el

Towards Imperceptible JPEG Image Hiding: Multi-range Representations-driven Adversarial Stego Generation

Image hiding fully explores the hidden potential of deep learning-based models, aiming to conceal image-level messages within cover images and reveal them from stego images to achieve covert communication. Existing hiding schemes are easily detected by the naked eyes or steganalyzers due to the cover type confined to the spatial domain, single-range feature extraction and attacks, and insufficient loss constraints. To address these issues, we propose a multi-range representations-driven adversarial stego generation framework called MRAG for JPEG image hiding. This design stems from the fact that steganalyzers typically combine local-range and global-range information to better capture hidden traces. Specifically, MRAG integrates the local-range characteristic of the convolution and the global-range modeling of the transformer. Meanwhile, a features angle-norm disentanglement loss is designed to launch multi-range representations-driven feature-level adversarial attacks. It computes the adversarial loss between covers and stegos based on the surrogate steganalyzer's classified features, i.e., the features before the last fully connected layer. Under the dual constraints of features angle and norm, MRAG can delicately encode the concatenation of cover and secret into subtle adversarial perturbations from local and global ranges relevant to steganalysis. Therefore, the resulting stego can achieve visual and steganalysis imperceptibility. Moreover, coarse-grained and fine-grained frequency decomposition operations are devised to transform the input, introducing multi-grained information. Extensive experiments demonstrate that MRAG can achieve state-of-the-art performance.

cs.CV

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning

Dense video captioning is a challenging task that aims to localize and caption multiple events in an untrimmed video. Recent studies mainly follow the transformer-based architecture to jointly perform the two sub-tasks, i.e., event localization and caption generation, in an end-to-end manner. Based on the general philosophy of detection transformer, these methods implicitly learn the event locations and event semantics, which requires a large amount of training data and limits the model's performance in practice. In this paper, we propose a novel dense video captioning framework, named PR-DETR, which injects the explicit position and relation prior into the detection transformer to improve the localization accuracy and caption quality, simultaneously. On the one hand, we first generate a set of position-anchored queries to provide the scene-specific position and semantic information about potential events as position prior, which serves as the initial event search regions to eliminate the implausible event proposals. On the other hand, we further design an event relation encoder to explicitly calculate the relationship between event boundaries as relation prior to guide the event interaction to improve the semantic coherence of the captions. Extensive ablation studies are conducted to verify the effectiveness of the position and relation prior. Experimental results also show the competitive performance of our method on ActivityNet Captions and YouCook2 datasets.

cs.CV

Parameter inference of millilensed gravitational waves using neural spline flows

When gravitational waves (GWs) propagate near massive objects, they undergo gravitational lensing that imprints lens model dependent modulations on the waveform. This effect provides a powerful tool for cosmological and astrophysical studies. Due to the added parameters of lenses and the uncertainty of lens models, parameter inference for lensed GW events using traditional methods is extremely time-consuming, thus requiring more efficient parameter inference methods. In this work, we explore the use of neural spline flows (NSFs) for posterior inference of millilensed GWs, and successfully apply NSFs to the inference of 11-dimensional lens parameters. Our results demonstrate that compared with traditional methods like Bilby dynesty that rely on Bayesian inference, the NSF network we built not only achieves inference accuracy comparable to traditional methods for most parameters, but also can reduce the inference time from approximately 3 days to 0.8 s on average. Additionally, the network exhibits strong generalization for the spin parameters of GW sources. It is anticipated to become a powerful tool for future low-latency searches for lensed GW signals.

gr-qc

From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval

Composed Image Retrieval (CIR) is a challenging multimodal task that retrieves a target image based on a reference image and accompanying modification text. Due to the high cost of annotating CIR triplet datasets, zero-shot (ZS) CIR has gained traction as a promising alternative. Existing studies mainly focus on projection-based methods, which map an image to a single pseudo-word token. However, these methods face three critical challenges: (1) insufficient pseudo-word token representation capacity, (2) discrepancies between training and inference phases, and (3) reliance on large-scale synthetic data. To address these issues, we propose a two-stage framework where the training is accomplished from mapping to composing. In the first stage, we enhance image-to-pseudo-word token learning by introducing a visual semantic injection module and a soft text alignment objective, enabling the token to capture richer and fine-grained image information. In the second stage, we optimize the text encoder using a small amount of synthetic triplet data, enabling it to effectively extract compositional semantics by combining pseudo-word tokens with modification text for accurate target image retrieval. The strong visual-to-pseudo mapping established in the first stage provides a solid foundation for the second stage, making our approach compatible with both high- and low-quality synthetic data, and capable of achieving significant performance gains with only a small amount of synthetic data. Extensive experiments were conducted on three public datasets, achieving superior performance compared to existing approaches.

cs.CV