arXiv ScienceSearch

arXiv subjects

Xuyang Wang

Publications and source records attributed to Xuyang Wang.

At least 19 recordsLinked to original sources

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typically separate geometry reconstruction, part reasoning, and articulation estimation into different stages. This separation can weaken consistency between shape, active parts, and motion, while also incurring substantial inference cost. We introduce Artic-O, an end-to-end, feed-forward framework for articulated object reconstruction via latent geometry learning. Instead of fitting geometry in image or view space, Artic-O maps sparse multi-state observations into a pretrained latent geometry space, where a frozen flow-matching decoder provides a complete-shape prior for recovering visible and occluded structures. To connect geometry with articulation, Artic-O fuses visual tokens, geometry latents, and point-wise decoder features in an image-grounded part-reasoning module for active-part segmentation and articulation prediction. We further train the model with a geometry-to-articulation curriculum and a decoupled two-pass strategy to balance reconstruction and part-level supervision. On PartNet-Mobility, Artic-O achieves strong reconstruction quality while being substantially more efficient than LARM, a strong prior method. It reduces Chamfer Distance, improves F-score, and achieves comparable or better articulation accuracy across most joint metrics, while reducing inference time from 9 minutes to about 0.3 seconds per object.

cs.CV

Design and commissioning of a windowless gas-target system for high-current beams at JUNA

Windowless gas targets avoid the beam-energy loss and straggling introduced by entrance foils and are therefore well suited for direct measurements of low-energy nuclear reactions. A windowless gas-target system designed for operation with milliampere beams has been developed for the Jinping Underground Nuclear Astrophysics facility (JUNA). The system combines three-stage differential pumping, closed-loop gas recovery and purification, a constant-temperature power-compensation calorimeter, and a position-resolved target-thickness monitor based on secondary elastic scattering. Stable operation was achieved over a target-pressure range of 1-3 mbar, with pressure fluctuations below 1% during 8 h of continuous circulation, while the accelerator-side pressure was maintained at approximately \(10^{-4}\) Pa. The closed-loop gas-circulation system maintained stable target conditions, while gas-transport calculations indicated that the axial pressure nonuniformity remained within approximately 1.6% under representative operating conditions. Calorimeter measurements were consistent with the thermal calculations, supporting the sensitivity correction used for beam-power determination. Beam commissioning with \(^{14}\mathrm{N}(p,γ)^{15}\mathrm{O}\) and \(^{12}\mathrm{C}(p,γ)^{13}\mathrm{N}\) at the 600 kV Cockcroft-Walton accelerator of the China Institute of Atomic Energy (CIAE) demonstrated stable operation of the gas-target and \(γ\)-ray detection systems and provided information on the influences of reaction position and beam heating. These results demonstrate the operating stability and diagnostic capability of the system for future high-current, low-energy nuclear-reaction measurements at JUNA.

astro-ph.GA

Research and simulation of analytical polarization control enabled by optical computing on an integrated photonics chip

Dynamic polarization controllers are key devices with broad applications in many fields. However, most on-chip polarization controllers still rely on traditional blind-search methods, whereas analytical optical-computing approaches remain insufficiently explored, particularly with respect to calibration and endless polarization control. With the accurate relative phase of Mach-Zehnder interferometer (MZI) being fully controllable on an integrated photonics chip, we present an analytical polarization control (APC) method using four phase shifters and optical computing, eliminating the need for the traditional inefficient blind-search procedure. The basic structures and operations of APC are clarified. The proposed calibration method and endless control method enable continuous APC while compensating for phase differences within the MZI structures. We simulate the influence of the endless control unit on polarization control and quantify the effect of the fourth phase difference on the output extinction ratio. With the fourth phase shifter, the phase difference encountered during Stokes vector measurement can be effectively compensated, and rotations around all three axes on the Poincaré sphere can be realized. These results establish a practical APC architecture based on optical computing for photonics chips. The proposed APC methods, combined with a FPGA-based hardware acceleration, will enable high speed on-chip polarization controllers.

physics.optics

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application

Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be used to derive its group-level aggregate and its individual deviation, allowing researchers to estimate both between-group and within-group associations rather than treating AI measurement as response prediction alone. We validate the framework using the 2022 China Family Panel Studies (CFPS), where occupations provide the empirical grouping structure and several job-related survey variables provide validation benchmarks. For computer use, foreign-language use, weekly hours, and management responsibilities, we compare survey measures with AI-derived measures in response-level, model-level, contextual, and boundary-condition validations. The results show that AI contextual measurement can recover much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics are available. Weekly hours provides the strongest validation case, with AI-derived measures reproducing the large negative between- and within-occupation associations with satisfaction observed in CFPS. The framework also identifies clear boundary conditions: performance deteriorates when information is restricted to occupation and basic demographics, and recovery is weaker when several related concepts are treated as simultaneously unobserved. The findings suggest that AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets.

cs.AI

BridgeFlow: Fast and Robust SE(2)-Equivariant Motion Planning with Flow Matching

In robotic motion planning, equivariance to rigid body transformations is crucial for robust spatial generalization. However, current learning-based planners face a critical dilemma: they either lack inherent equivariance, treating transformed tasks as novel scenarios, or enforce it via computationally expensive specialized architectures that bottleneck real-time inference. To break this trade-off, we propose BridgeFlow, a fast and strictly SE(2)-equivariant generative motion planning framework. Rather than relying on heavy equivariant networks, BridgeFlow achieves exact spatial equivariance via a lightweight task-centric canonicalization module, enabling generalization using standard architectures. To further accelerate inference, we pair a Brownian bridge informative prior with context-aware mini-batch optimal transport. This constructs a straightened vector field that minimizes transport costs and stabilizes training. Furthermore, environmental awareness is explicitly embedded via Classifier-Free Guidance. Evaluations in dense 2D environments and on a 7-DoF Franka manipulator demonstrate that BridgeFlow achieves up to a 15x inference speedup and a 2x higher valid trajectory rate over state-of-the-art diffusion baselines, alongside robust generalization to entirely unseen environments and arbitrary spatial transformations.

cs.RO

A Kung Fu Athlete Bot That Can Do It All Day: Highly Dynamic, Balance-Challenging Motion Dataset and Autonomous Fall-Resilient Tracking

Current humanoid motion tracking systems can execute routine and moderately dynamic behaviors, yet significant gaps remain near hardware performance limits and algorithmic robustness boundaries. Martial arts represent an extreme case of highly dynamic human motion, characterized by rapid center-of-mass shifts, complex coordination, and abrupt posture transitions. However, datasets tailored to such high-intensity scenarios remain scarce. To address this gap, we construct KungFuAthlete, a high-dynamic martial arts motion dataset derived from professional athletes' daily training videos. The dataset includes ground and jump subsets covering representative complex motion patterns. The jump subset exhibits substantially higher joint, linear, and angular velocities compared to commonly used datasets such as LAFAN1, PHUMA, and AMASS, indicating significantly increased motion intensity and complexity. Importantly, even professional athletes may fail during highly dynamic movements. Similarly, humanoid robots are prone to instability and falls under external disturbances or execution errors. Most prior work assumes motion execution remains within safe states and lacks a unified strategy for modeling unsafe states and enabling reliable autonomous recovery. We propose a novel training paradigm that enables a single policy to jointly learn high-dynamic motion tracking and fall recovery, unifying agile execution and stabilization within one framework. This framework expands robotic capability from pure motion tracking to recovery-enabled execution, promoting more robust and autonomous humanoid performance in real-world high-dynamic scenarios.

cs.RO

Semi-Device-Independent Quantum Random Number Generator Resistant to General Attacks

Quantum random number generators (QRNGs) produce true random numbers based on the inherent randomness of quantum theory, rendering them a foundational segment of quantum cryptography. Distinguished from trusted-device QRNGs whose security depends on characterized devices, semi-device-independent (semi-DI) QRNGs permit partial devices to be defective or even maliciously manipulated, which achieves a good trade-off between generation rate and security. In this paper, we propose a semi-DI QRNG that resists general attacks while accounting for finite-size effects. The protocol requires no rigorous characterization of the source and measurement devices other than limiting the energy of the emitted states, significantly reducing the demands on practical QRNG systems. Leveraging the tight Kato inequality for correlated variables, we show that our protocol generates more randomness than it consumes. Furthermore, we demonstrate the scheme on a continuous-variable system with ternary inputs of states. Heterodyne detection is employed to enable phase compensation through data postprocessing, alleviating the stringent requirement on system stability. The system operates at 100 MHz, achieving a net random number generation rate of 1.165 Mbps at 5.3x10^9 rounds. Our work offers a promising approach to achieve both the robust security and high generation rate with a simple experimental setup.

quant-ph

Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware Benchmark

While cloud-native microservice architectures have revolutionized software development, their inherent operational complexity makes failure Root Cause Analysis (RCA) a critical yet challenging task. Numerous data-driven RCA models have been proposed to address this challenge. However, we find that the benchmarks used to evaluate these models are often too simple to reflect real-world scenarios. Our preliminary study reveals that simple rule-based methods can achieve performance comparable to or even surpassing state-of-the-art (SOTA) models on four widely used public benchmarks. This finding suggests that the oversimplification of existing benchmarks might lead to an overestimation of the performance of RCA methods. To further investigate the oversimplification issue, we conduct a systematic analysis of popular public RCA benchmarks, identifying key limitations in their fault injection strategies, call graph structures, and telemetry signal patterns. Based on these insights, we propose an automated framework for generating more challenging and comprehensive benchmarks that include complex fault propagation scenarios. Our new dataset contains 1,430 validated failure cases from 9,152 fault injections, covering 25 fault types across 6 categories, dynamic workloads, and hierarchical ground-truth labels that map failures from services down to code-level causes. Crucially, to ensure the failure cases are relevant to IT operations, each case is validated to have a discernible impact on user-facing SLIs. Our re-evaluation of 11 SOTA models on this new benchmark shows that they achieve low Top@1 accuracies, averaging 0.21, with the best-performing model reaching merely 0.37, and execution times escalating from seconds to hours.

cs.SE

Continuous-variable quantum key distribution over 50.4 km fiber using integrated silicon photonic transmitter and receiver

Quantum key distribution (QKD) is the fastest-growing and relatively mature technology in the field of quantum information, enabling information-theoretically secure key distribution between two remote users. Although QKD based on off-the-shelf telecom components has been validated in both laboratory and field tests, its high cost and large volume remain major obstacles to large-scale deployment. Photonic integration, featured by its compact size and low cost, offers an effective approach to addressing the above challenges faced by QKD. Here, we implement a high-performance, integrated local local oscillator continuous-variable (CV) QKD system based on an integrated silicon photonic transmitter and receiver. By employing a high-speed silicon photonic integrated in-phase and quadrature modulator, a low-noise and high bandwidth silicon photonic integrated heterodyne detector, and digital signal processing, our CV-QKD system achieves a symbol rate of up to 1.5625 GBaud. Furthermore, the system achieves asymptotic secret key rates of 31.05 and 5.05 Mbps over 25.8 and 50.4 km standard single-mode fiber, respectively, using an 8-phase-shift keying discrete modulation. Our integrated CV-QKD system with high symbol rate and long transmission distance pays the way for the quantum secure communication network at metropolitan area.

quant-ph

CoCAViT: Compact Vision Transformer with Robust Global Coordination

In recent years, large-scale visual backbones have demonstrated remarkable capabilities in learning general-purpose features from images via extensive pre-training. Concurrently, many efficient architectures have emerged that have performance comparable to that of larger models on in-domain benchmarks. However, we observe that for smaller models, the performance drop on out-of-distribution (OOD) data is disproportionately larger, indicating a deficiency in the generalization performance of existing efficient models. To address this, we identify key architectural bottlenecks and inappropriate design choices that contribute to this issue, retaining robustness for smaller models. To restore the global field of pure window attention, we further introduce a Coordinator-patch Cross Attention (CoCA) mechanism, featuring dynamic, domain-aware global tokens that enhance local-global feature modeling and adaptively capture robust patterns across domains with minimal computational overhead. Integrating these advancements, we present CoCAViT, a novel visual backbone designed for robust real-time visual representation. Extensive experiments empirically validate our design. At a resolution of 224*224, CoCAViT-28M achieves 84.0% top-1 accuracy on ImageNet-1K, with significant gains on multiple OOD benchmarks, compared to competing models. It also attains 52.2 mAP on COCO object detection and 51.3 mIOU on ADE20K semantic segmentation, while maintaining low latency.

cs.CV

High-Performance Fully Passive Discrete-State Continuous-Variable Quantum Key Distribution With Local Local Oscillator

We propose and demonstrate a fully passive discrete-state continuous-variable quantum key distribution (CV-QKD), which can eliminate all modulator side channels on the source side, using a local local oscillator (LLO). The CV-QKD system achieves a maximum transmission length of 100 km with a repetition rate of 1 GHz using specially designed phase rotation and discretization methods, and the corresponding secret key bit rate is 127 kbps, as estimated based on the amplitude of prepared states at the transmitter, as well as the first- and second-order moments of quadratures at the receiver by employing the convex optimization without imposing any assumptions on the quantum channel. The performance of the protocol is similar to that of modulated CV LLO protocols and better than those of passive discrete-variable and CV protocols. Our protocol is expected to play an important role in the quantum metropolitan area networks and quantum access networks with high realistic security.

quant-ph

Reliable Disentanglement Multi-view Learning Against View Adversarial Attacks

Trustworthy multi-view learning has attracted extensive attention because evidence learning can provide reliable uncertainty estimation to enhance the credibility of multi-view predictions. Existing trusted multi-view learning methods implicitly assume that multi-view data is secure. However, in safety-sensitive applications such as autonomous driving and security monitoring, multi-view data often faces threats from adversarial perturbations, thereby deceiving or disrupting multi-view models. This inevitably leads to the adversarial unreliability problem (AUP) in trusted multi-view learning. To overcome this tricky problem, we propose a novel multi-view learning framework, namely Reliable Disentanglement Multi-view Learning (RDML). Specifically, we first propose evidential disentanglement learning to decompose each view into clean and adversarial parts under the guidance of corresponding evidences, which is extracted by a pretrained evidence extractor. Then, we employ the feature recalibration module to mitigate the negative impact of adversarial perturbations and extract potential informative features from them. Finally, to further ignore the irreparable adversarial interferences, a view-level evidential attention mechanism is designed. Extensive experiments on multi-view classification tasks with adversarial attacks show that RDML outperforms the state-of-the-art methods by a relatively large margin. Our code is available at https://github.com/Willy1005/2025-IJCAI-RDML.

cs.LG

Formation Maneuver Control Based on the Augmented Laplacian Method

This paper proposes a novel formation maneuver control method for both 2-D and 3-D space, which enables the formation to translate, scale, and rotate with arbitrary orientation. The core innovation is the novel design of weights in the proposed augmented Laplacian matrix. Instead of using scalars, we represent weights as matrices, which are designed based on a specified rotation axis and allow the formation to perform rotation in 3-D space. To further improve the flexibility and scalability of the formation, the rotational axis adjustment approach and dynamic agent reconfiguration method are developed, allowing formations to rotate around arbitrary axes in 3-D space and new agents to join the formation. Theoretical analysis is provided to show that the proposed approach preserves the original configuration of the formation. The proposed method maintains the advantages of the complex Laplacian-based method, including reduced neighbor requirements and no reliance on generic or convex nominal configurations, while achieving arbitrary orientation rotations via a more simplified implementation. Simulations in both 2-D and 3-D space validate the effectiveness of the proposed method.

eess.SY

Highly integrated broadband entropy source for quantum random number generators based on vacuum fluctuations

In this work, we designed and experimentally verified a highly integrated broadband entropy source for a quantum random number generator (QRNG) based on vacuum fluctuations. The core of the entropy source is a hybrid laser-and-silicon-photonics chip, which is only 6.3 $ \times $ 2.6 $ \times $ 1.5 mm$^{3}$ in size. A balanced homodyne detector based on cascaded radio-frequency amplifiers in the entropy source achieves a 3 dB bandwidth of 2.4 GHz and a common-mode rejection ratio above 25 dB. The quantum-to-classical-noise ratio is 9.51 dB at a photoelectron current of 1 mA. The noise equivalent power and equivalent transimpedance are 8.85$\,\text{pW}/\sqrt{\text{Hz}}$ , and 22.8 k$Ω$, respectively. After optimization using equalizer technology that eliminates the dependence of adjacent samples, the quantum random number generation rate reaches 67.9 Gbps under average conditional minimum entropy and 61.9 Gbps under the worst-case conditional minimum entropy. The developed hybrid chip enhances the integrability and speed of QRNG entropy sources based on vacuum fluctuations.

quant-ph

Leveraging Large Self-Supervised Time-Series Models for Transferable Diagnosis in Cross-Aircraft Type Bleed Air System

Bleed Air System (BAS) is critical for maintaining flight safety and operational efficiency, supporting functions such as cabin pressurization, air conditioning, and engine anti-icing. However, BAS malfunctions, including overpressure, low pressure, and overheating, pose significant risks such as cabin depressurization, equipment failure, or engine damage. Current diagnostic approaches face notable limitations when applied across different aircraft types, particularly for newer models that lack sufficient operational data. To address these challenges, this paper presents a self-supervised learning-based foundation model that enables the transfer of diagnostic knowledge from mature aircraft (e.g., A320, A330) to newer ones (e.g., C919). Leveraging self-supervised pretraining, the model learns universal feature representations from flight signals without requiring labeled data, making it effective in data-scarce scenarios. This model enhances both anomaly detection and baseline signal prediction, thereby improving system reliability. The paper introduces a cross-model dataset, a self-supervised learning framework for BAS diagnostics, and a novel Joint Baseline and Anomaly Detection Loss Function tailored to real-world flight data. These innovations facilitate efficient transfer of diagnostic knowledge across aircraft types, ensuring robust support for early operational stages of new models. Additionally, the paper explores the relationship between model capacity and transferability, providing a foundation for future research on large-scale flight signal models.

eess.SP

DoubleDiffusion: Combining Heat Diffusion with Denoising Diffusion for Texture Generation on 3D Meshes

This paper addresses the problem of generating textures for 3D mesh assets. Existing approaches often rely on image diffusion models to generate multi-view image observations, which are then transformed onto the mesh surface to produce a single texture. However, due to the gap between multi-view images and 3D space, such process is susceptible to arange of issues such as geometric inconsistencies, visibility occlusion, and baking artifacts. To overcome this problem, we propose a novel approach that directly generates texture on 3D meshes. Our approach leverages heat dissipation diffusion, which serves as an efficient operator that propagates features on the geometric surface of a mesh, while remaining insensitive to the specific layout of the wireframe. By integrating this technique into a generative diffusion pipeline, we significantly improve the efficiency of texture generation compared to existing texture generation methods. We term our approach DoubleDiffusion, as it combines heat dissipation diffusion with denoising diffusion to enable native generative learning on 3D mesh surfaces.

cs.CV

Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing

Effectively distinguishing the pronunciation correlations between different written texts is a significant issue in linguistic acoustics. Traditionally, such pronunciation correlations are obtained through manually designed pronunciation lexicons. In this paper, we propose a data-driven method to automatically acquire these pronunciation correlations, called automatic text pronunciation correlation (ATPC). The supervision required for this method is consistent with the supervision needed for training end-to-end automatic speech recognition (E2E-ASR) systems, i.e., speech and corresponding text annotations. First, the iteratively-trained timestamp estimator (ITSE) algorithm is employed to align the speech with their corresponding annotated text symbols. Then, a speech encoder is used to convert the speech into speech embeddings. Finally, we compare the speech embeddings distances of different text symbols to obtain ATPC. Experimental results on Mandarin show that ATPC enhances E2E-ASR performance in contextual biasing and holds promise for dialects or languages lacking artificial pronunciation lexicons.

eess.AS

Calibration of cascaded phase shifters using pairwise scan method in silicon photonics integrated chip

Cascaded phase shifters (CPSs) based on silicon photonics integrated chips play important roles in quantum information processing tasks. Owing to an increase in the scale of silicon photonics chips, the time required to calibrate various CPSs has increased. We propose a pairwise scan method for rapidly calibrating CPSs by introducing equivalent Mach Zehnder Interferometer structures and a reasonable constraint of the initial relative phase. The calibration can be nearly completed when the scanning process is finished, and only a little calculation is required. To achieve better performance, the key components, thermal optical phase shifter and 2 * 2 50/50 multimode interference coupler, were simulated and optimized to prevent thermal crosstalk and ensure a good balance. A 6-CPSs structure in a packaged silicon photonics chip under different temperature was used to verify the rapid pairwise scan method, and a fidelity of 99.97% was achieved.

quant-ph