arXiv ScienceSearch

arXiv subjects

Xu Cheng

Publications and source records attributed to Xu Cheng.

At least 19 recordsLinked to original sources

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end. We scale this architecture to 8.9B parameters and train it on 5.73T tokens from the Dolma-3 dataset, marking the largest demonstration of a latent-space language model to date. Remarkably, by consuming only 51.3% of the total training tokens, NCP-ArchPreview achieves the final pretraining loss of OLMo-3-7B. Following full pretraining, it outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, including a notable 5.99-point gain on GSM8K. Controlled experiments isolate a clear progression of performance gains stemming from both the latent architecture and the NCP objective. Furthermore, utilizing only 85% of the standard computation, NCP-ArchPreview approaches the training loss of a strictly parameter-aligned 8.9B baseline. The learned latent space remains highly valuable after the pretraining stage: updating just the 17M-parameter VQ module yields a novel, lightweight interface for domain adaptation, while a simple injection of concept representations into a DFlash2 drafter improves the mean accepted length by 4.17% with negligible overhead.

cs.CL

Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities

In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance while maintaining low computational cost. Existing methods struggle to address semantic degradation caused by missing modalities. We propose Compass, a lightweight network for robust crack segmentation under arbitrary missing modalities. Compass comprises Degradation Simulation Distillation (DSD), Needle Block, and Evidential Topology-Preserving Fusion (ETPF). DSD constructs a degradation simulation stream that mimics more severe missing conditions and performs reciprocal distillation with the original stream, decoupling complete perception from degradation adaptation. Within DSD, Feature-Aware Prototype Transmitter (FAPT) performs modality agnostic prototype-guided feature completion to maintain semantic integrity under incomplete modality conditions. As a lightweight backbone, Needle injects crack-direction cues into WKV modulation and combines connectivity-aware gating with anisotropic context probing for structure-aware modeling. ETPF fuses multimodal features via Dempster-Shafer evidential combination with uncertainty-gated decoding, preserving crack topology while suppressing unreliable features. Experiments on three datasets demonstrate state-of-the-art (SOTA) performance under diverse missing modality scenarios. Even with 90\% depth modality missing on CrackDepth, Compass achieves F1 of 0.8216 and mIoU of 0.8434 with only 2.58M parameters. The code is available at https://github.com/Karl1109/Compass.

cs.CV

LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data. However, existing SF-UniDA methods rely on inefficient techniques such as threshold tuning and clustering. Foundation models (FMs), known for their generalization and zero-shot capabilities, remain underexplored in SF-UniDA. In this paper, we propose a framework that leverages foundation models (LFM) for SF-UniDA. We use a vision-language model (VLM) to compute similarities between target samples and text labels, including those for unknown classes generated by prompting a large language model. The label shift type is determined by analyzing the coefficient of variation of a similarity-based sample-level score. Unknown samples are identified using a binary Gaussian mixture model fitted to another similarity-based metric. Under a consensus strategy, the pseudo-labels generated by the VLM are refined by the target model initialized with the pre-trained source model, integrating knowledge from both the source domain and foundation models. Finally, these refined pseudo-labels are used to train the target model. Extensive experiments across all possible label shifts and multiple benchmarks demonstrate the effectiveness and superiority of our proposed LFM framework. Our code is available at https://github.com/iamjingli/LFM.

cs.CV

Noise-Robust Box-Supervised Infrared Small Target Detection via Physics-Inspired Soft Label Optimization

Infrared small target detection (IRSTD) commonly relies on pixel-level mask supervision. Such annotations, however, are costly and inherently uncertain because infrared targets have blurred boundaries and weak textures. We formulate box-supervised IRSTD as a problem distinct from generic box-to-mask segmentation and point-supervised IRSTD. Its central challenge is to construct stable pixel-level soft supervision from highly contaminated boxes. To this end, we propose Hotspot-Anchored Label Optimization (HALO). HALO localizes a radiometric anchor inside each box under local background-statistics constraints, then synthesizes a Physically Anchored Gaussian (PAG) soft label around the anchor. This turns noisy box supervision into continuous, pixel-level soft labels. The entire process is performed offline before training, remains decoupled from the detector backbone, and requires no online label updates. Experiments on public datasets show that HALO is competitive with representative box-supervised methods under standard tight boxes. Under looser or shifted box annotations that better approximate real scenarios, HALO is substantially more robust while remaining consistent across backbones. We further introduce a contamination-aware operating-regime analysis to characterize the effective boundary of this class of methods and reveal how intrinsic signal-to-clutter ratio relates to performance.

cs.CV

Temporal Dynamical Quantum Phase Transition in Dicke Model with Trapped Ions

Temporal non-analyticities in the rate function of the Loschmidt echo manifests a class of dynamical quantum phase transitions (DQPTs) that has emerged as a powerful framework for understanding far-from-equilibrium many-body dynamics. While such DQPT has been extensively studied theoretically in spin-boson systems such as the Dicke model, their experimental observation remains elusive. In particular, the dynamics of DQPT in asymmetric spin subspaces and under the influence of spin dissipation are largely unexplored. Here, we report an experimental study of temporal DQPT in a generalized Dicke model using a trapped-ion quantum simulator. By coupling a linear chain of $\rm{^{40}Ca^{+}}$ ions to a collective center-of-mass motional mode, we probe the quench dynamics starting from both symmetric and asymmetric initial states. We extract the rate function and identify temporal turn-around points that are in quantitative agreement with theoretical predictions. Additionally, we investigate the impact of spin dissipation on these dynamics. Our results establish an experimental platform for probing complex many-body out-of-equilibrium phenomena and advance the development of hybrid oscillator-spin quantum simulators.

quant-ph

Eigenvalue Estimates for Schr\"odinger Operators on Ricci Shrinkers

Let $(M, g, f, \tau)$ be a complete Ricci shrinker satisfying $\textrm{Ric}+\nabla^2f=\frac{g}{2\tau}$ and let $R$ denote its scalar curvature. For a confined function $V$ on $M$, we obtain a lower bound for the lowest eigenvalue of the Schr\"odinger operator $-\Delta+\frac{R}{4}+V$, expressed in terms of an integral quantity involving $V$ and the shrinker entropy, and the equality case is characterized by the potential functions. We further generalize this estimate to complete Riemannian manifolds via Perelman's $\mu$-functional. We also study the drifted Schr\"odinger operator $-\Delta_f+V$ on smooth metric measure spaces. In particular, on Ricci shrinkers, we derive a lower bound for its lowest eigenvalue, with equality if and only if $V$ is affine.

math.DG

SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation

Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face significant bottlenecks in balancing crack topology modeling with computational efficiency, often failing to reconcile high segmentation quality with low resource demands. To address these limitations, we propose the Ultra-Compact Structure-Calibrated Vision RWKV (SCRWKV), a network that achieves high-precision modeling via a novel Structure-Field Encoder (SFE) backbone while maintaining linear complexity. The SFE integrates the Adaptive Multi-scale Cascaded Modulator (AMCM) to enhance texture representation and utilizes the Structure-Calibrated Insight Unit (SCIU) as its core engine. Specifically, the SCIU employs the Geometry-guided Bidirectional Structure Transformation (GBST) to capture topological correlations and integrates the Dynamic Self-Calibrating Decay (DSCD) into Dy-WKV to suppress noise propagation. Furthermore, we introduce a lightweight Cross-Scale Harmonic Fusion (CSHF) decoder to achieve precise feature aggregation. Systematic evaluations on multiple benchmarks characterized by complex textures and severe interference demonstrate that SCRWKV, with only 1.22M parameters, significantly outperforms SOTA methods. Achieving an F1 score of 0.8428 and mIoU of 0.8512 on the TUT dataset, the model confirms its robust potential for efficient real-world deployment. The code is available at https://github.com/zhxhzy/SCRWKV.

cs.CV

When W4A4 Breaks Camouflaged Object Detection: Token-Group Dual-Constraint Activation Quantization

Camouflaged object detection (COD) segments objects that intentionally blend with the background, so predictions depend on subtle texture and boundary cues. COD is often needed under tight on-device memory and latency budgets, making low-bit inference highly desirable. However, COD is unusually hard to quantize aggressively. We study post-training W4A4 quantization of Transformer-based COD and find a task-specific cliff: heavy-tailed background tokens dominate a shared activation range, inflating the step size and pushing weak-but-structured boundary cues into the zero bin. This exposes a token-local bottleneck -- remove cross-token range domination and bound the zero-bin mass under 4-bit activations. To address this, we introduce COD-TDQ, a COD-aware Token-group Dual-constraint activation Quantization method. COD-TDQ addresses this token-local bottleneck with two coupled steps: Direct-Sum Token-Group (DSTG) assigns token-group scales to suppress cross-token range domination, and Dual-Constraint Range Projection (DCRP) projects each token-group clip range to keep the step-to-dispersion ratio and the zero-bin mass bounded. Across four COD benchmarks and two baseline models (CFRN and ESCNet), COD-TDQ consistently achieves an $S_\alpha$ score more than 0.12 higher than that of the state-of-the-art quantization method without retraining. The code is available at https://github.com/MCG-NKU/nku-model-compre.

cs.CV

NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results

This paper presents an overview of the NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models. This challenge utilizes a new short-form UGC (S-UGC) video restoration benchmark, termed KwaiVIR, which is contributed by USTC and Kuaishou Technology. It contains both synthetically distorted videos and real-world short-form UGC videos in the wild. For this edition, the released data include 200 synthetic training videos, 48 wild training videos, 11 validation videos, and 20 testing videos. The primary goal of this challenge is to establish a strong and practical benchmark for restoring short-form UGC videos under complex real-world degradations, especially in the emerging paradigm of generative-model-based S-UGC video restoration. This challenge has two tracks: (i) the primary track is a subjective track, where the evaluation is based on a user study; (ii) the second track is an objective track. These two tracks enable a comprehensive assessment of restoration quality. In total, 95 teams have registered for this competition. And 12 teams submitted valid final solutions and fact sheets for the testing phase. The submitted methods achieved strong performance on the KwaiVIR benchmark, demonstrating encouraging progress in short-form UGC video restoration in the wild.

cs.CV

TIEG-Youpu Solution for NeurIPS 2022 WikiKG90Mv2-LSC

WikiKG90Mv2 in NeurIPS 2022 is a large encyclopedic knowledge graph. Embedding knowledge graphs into continuous vector spaces is important for many practical applications, such as knowledge acquisition, question answering, and recommendation systems. Compared to existing knowledge graphs, WikiKG90Mv2 is a large scale knowledge graph, which is composed of more than 90 millions of entities. Both efficiency and accuracy should be considered when building graph embedding models for knowledge graph at scale. To this end, we follow the retrieve then re-rank pipeline, and make novel modifications in both retrieval and re-ranking stage. Specifically, we propose a priority infilling retrieval model to obtain candidates that are structurally and semantically similar. Then we propose an ensemble based re-ranking model with neighbor enhanced representations to produce final link prediction results among retrieved candidates. Experimental results show that our proposed method outperforms existing baseline methods and improves MRR of validation set from 0.2342 to 0.2839.

cs.CL

MVRD-Bench: Multi-View Learning and Benchmarking for Dynamic Remote Photoplethysmography under Occlusion

Remote photoplethysmography (rPPG) is a non-contact technique that estimates physiological signals by analyzing subtle skin color changes in facial videos. Existing rPPG methods often encounter performance degradation under facial motion and occlusion scenarios due to their reliance on static and single-view facial videos. Thus, this work focuses on tackling the motion-induced occlusion problem for rPPG measurement in unconstrained multi-view facial videos. Specifically, we introduce a Multi-View rPPG Dataset (MVRD), a high-quality benchmark dataset featuring synchronized facial videos from three viewpoints under stationary, speaking, and head movement scenarios to better match real-world conditions. We also propose MVRD-rPPG, a unified multi-view rPPG learning framework that fuses complementary visual cues to maintain robust facial skin coverage, especially under motion conditions. Our method integrates an Adaptive Temporal Optical Compensation (ATOC) module for motion artifact suppression, a Rhythm-Visual Dual-Stream Network to disentangle rhythmic and appearance-related features, and a Multi-View Correlation-Aware Attention (MVCA) for adaptive view-wise signal aggregation. Furthermore, we introduce a Correlation Frequency Adversarial (CFA) learning strategy, which jointly enforces temporal accuracy, spectral consistency, and perceptual realism in the predicted signals. Extensive experiments and ablation studies on the MVRD dataset demonstrate the superiority of our approach. In the MVRD movement scenario, MVRD-rPPG achieves an MAE of 0.90 and a Pearson correlation coefficient (R) of 0.99. The source code and dataset will be made available.

cs.CV

Adaptive Dual-Teacher Distillation with Subnetwork Rectification for Bridging Semantic Gaps in Black-Box Domain Adaptation

Assuming that neither source data nor source model parameters are accessible, black-box domain adaptation (BBDA) represents a highly practical yet challenging setting, where transferable knowledge is limited to the predictions of a black-box source model. Existing approaches exploit such knowledge via pseudo-label refinement or by leveraging vision-language models (ViLs), but they often fail to reconcile the inherent discrepancy between task-specific knowledge from black-box models and language-aligned semantic priors of ViLs, resulting in suboptimal integration and degraded adaptation performance. To address this challenge, we propose adaptive Dual-Teacher Distillation with Subnetwork Rectification (DDSR), a framework that explicitly reconciles these complementary yet inconsistent knowledge sources. DDSR employs an adaptive prediction fusion strategy to integrate predictions from the black-box source model and a ViL, generating reliable pseudo-labels for the target domain. A subnetwork-based regularization mechanism mitigates overfitting to noisy supervision by enforcing output consistency and gradient divergency. Furthermore, progressively improved target predictions iteratively refine both pseudo-labels and ViL prompts, enhancing semantic alignment. Finally, class-wise prototypes are used to further optimize target predictions via self-training. Extensive experiments on multiple benchmark datasets demonstrate that DDSR consistently outperforms state-of-the-art methods, including those with access to source data or source model parameters.

cs.CV

Non-Abelian Aharonov-Bohm Caging in Synthetic Dimensions with a Trapped Ion

Aharonov-Bohm (AB) caging is a complete localization phenomenon in two-dimensional lattices due to destructive interference induced by the background gauge fields. However, current investigations of AB caging are mostly restricted to the Abelian gauge field case, and the observation of AB caging under non-Abelian gauge fields in a quantum system still remains elusive. Here, we report experimental realization of tunable synthetic non-Abelian SU(2) gauge fields in a rhombic lattice, engineered within the synthetic dimensions of a vibrating trapped ion with multiple levels. We realize AB caging under both Abelian and non-Abelian gauge fields and systematically investigate the distinctive transport properties of the non-Abelian case. In particular, we observe typical emergent quantum dynamics unique to non-Abelian AB caging, including initial-state-dependent dynamics, second-order effects, and asymmetric caging behavior. These observations demonstrate the trapped ion system as a powerful platform for simulating emergent phenomena in high-dimensional quantum systems with exotic synthetic gauge fields.

quant-ph

Divergence Identity for the scalar curvature and Rigidity of Codazzi Tensors

We introduce a local vector field on an $n$-dimensional Riemannian manifold, defined as the sum of the covariant derivatives of a local orthonormal frame, and derive an explicit identity for its divergence, decomposed into a scalar curvature term and an auxiliary term involving connection coefficients. This result is applied to rigidity problems for Codazzi symmetric tensors. In particular, we give a new proof of a Tang-Yan theorem, which states that on a closed $n$-dimensional manifold with nonnegative scalar curvature, a smooth Codazzi symmetric tensor whose trace invariants up to order $n-1$ are constant must have constant eigenvalues. We also obtain further rigidity results under assumptions on elementary symmetric functions of the eigenvalues, with applications to the isoparametric rigidity of closed hypersurfaces in the unit sphere.

math.DG

Neural Factorization-based Bearing Fault Diagnosis

This paper studies the key problems of bearing fault diagnosis of high-speed train. As the core component of the train operation system, the health of bearings is directly related to the safety of train operation. The traditional diagnostic methods are facing the challenge of insufficient diagnostic accuracy under complex conditions. To solve these problems, we propose a novel Neural Factorization-based Classification (NFC) framework for bearing fault diagnosis. It is built on two core idea: 1) Embedding vibration time series into multiple mode-wise latent feature vectors to capture diverse fault-related patterns; 2) Leveraging neural factorization principles to fuse these vectors into a unified vibration representation. This design enables effective mining of complex latent fault characteristics from raw time-series data. We further instantiate the framework with two models CP-NFC and Tucker-NFC based on CP and Tucker fusion schemes, respectively. Experimental results show that both models achieve superior diagnostic performance compared with traditional machine learning methods. The comparative analysis provides valuable empirical evidence and practical guidance for selecting effective diagnostic strategies in high-speed train bearing monitoring.

cs.LG

Radar-APLANC: Unsupervised Radar-based Heartbeat Sensing via Augmented Pseudo-Label and Noise Contrast

Frequency Modulated Continuous Wave (FMCW) radars can measure subtle chest wall oscillations to enable non-contact heartbeat sensing. However, traditional radar-based heartbeat sensing methods face performance degradation due to noise. Learning-based radar methods achieve better noise robustness but require costly labeled signals for supervised training. To overcome these limitations, we propose the first unsupervised framework for radar-based heartbeat sensing via Augmented Pseudo-Label and Noise Contrast (Radar-APLANC). We propose to use both the heartbeat range and noise range within the radar range matrix to construct the positive and negative samples, respectively, for improved noise robustness. Our Noise-Contrastive Triplet (NCT) loss only utilizes positive samples, negative samples, and pseudo-label signals generated by the traditional radar method, thereby avoiding dependence on expensive ground-truth physiological signals. We further design a pseudo-label augmentation approach featuring adaptive noise-aware label selection to improve pseudo-label signal quality. Extensive experiments on the Equipleth dataset and our collected radar dataset demonstrate that our unsupervised method achieves performance comparable to state-of-the-art supervised methods. Our code, dataset, and supplementary materials can be accessed from https://github.com/RadarHRSensing/Radar-APLANC.

cs.CV

Quantum Simulation of Oscillatory Unruh Effect with Superposed Trajectories

The Unruh effect predicts an astonishing phenomenon that an accelerated detector would detect counts despite being in a quantum field vacuum in the rest frame. Since the required detector acceleration for its direct observation is prohibitively large, recent analog studies on quantum simulation platforms help to reveal various properties of the Unruh effect and explore the not-yet-understood physics of quantum gravity. To further reveal the quantum aspect of the Unruh effect, analogous experimental exploration of the correlation between the detector and the field, and the consequences for coherent quantum trajectories of the detector without classical counterparts, are essential steps but are currently missing. Here, we utilize a laser-controlled trapped ion to experimentally simulate an oscillating detector coupled with a cavity field. We observe joint excitation of both the detector and the field in the detector's frame, coincide with the coordinated dynamics predicted by the Unruh effect. Particularly, we simulate the detector moving in single and superposed quantum trajectories, where the latter case shows coherent interference of excitation. Our demonstration reveals properties of quantum coherent superposition of accelerating trajectories associated with quantum gravity theories that have no classical counterparts, and may offer a new avenue to investigate phenomena in quantum field theory and quantum gravity. We also show how a generalization of the method and results in this work may be beneficial for direct observation of the Unruh effect.

quant-ph

N\'eel-Vector-Orientation Induced Direction-Robust Spin Filtering in Two-Dimensional Altermagnets

Whether an antiferromagnet can host direction-robust spin-polarized transport without a conventional spin-selective band gap remains a central challenge in antiferromagnetic spintronics. Here we establish a gapless, direction-robust spin-filtering mechanism in a compensated two-dimensional altermagnetic Weyl semimetal that requires neither a spin-selective band gap nor a large velocity contrast between spin projections. Using Janus monolayer Ta$_2$TeSeO as a realistic platform, we combine symmetry analysis with first-principles calculations, full-Brillouin-zone Wannier interpolation, and semiclassical transport. Rotating the N\'eel vector removes a unitary-mirror constraint and shifts one Weyl-cone pair away from its parent high-symmetry line. For an in-plane N\'eel vector, the residual $C_{2z}\mathcal T$ symmetry forbids the independent $\sigma_y$ mass that would open a local gap, allowing the reconstructed cones to shift in momentum while remaining gapless. Breaking unitary $C_{2z}$ simultaneously lifts the energy equivalence of the remaining mirror-pinned Weyl cones. The resulting coexistence of a metallic spin-projected manifold and a low-DOS Weyl-derived manifold produces a predominantly DOS-driven conductance imbalance. At charge neutrality and 20~K, the longitudinal conductivity polarization for $\mathbf n\parallel x$ remains positive for every in-plane current direction and ranges from $76.4\%$ to $82.0\%$. The degenerate in-plane magnetic anisotropy facilitates reversible switching between symmetry-related spin-filtering states using strain or weak anisotropic fields. This N\'eel-vector-driven symmetry mechanism provides a general route to direction-robust gapless spin filtering in compensated altermagnets.

cond-mat.mes-hall