arXiv ScienceSearch

arXiv subjects

Zhe Chen

Publications and source records attributed to Zhe Chen.

At least 19 recordsLinked to original sources

One- and two-dimensional cluster states for topological phase simulation and measurement-based quantum computation

Quantum entanglement is a fundamental resource for quantum information processing and serves as a critical benchmark for quantum hardware performance. Cluster states are a special class of entangled states that serve as universal resources for measurement-based quantum computation and possess an intrinsic symmetry-protected topological order, which confers robustness against symmetry-respecting noise. Here we report the scalable preparation and verification of genuine multipartite cluster states on the 105-qubit Zuchongzhi 3.1 superconducting processor. We achieve one-dimensional cluster states of up to 95 qubits and two-dimensional cluster states of up to 72 qubits. The symmetry-protected topological cluster states exhibit input-state-dependent robustness under symmetry-breaking perturbations due to an operational parity structure that enhances the performance of measurement-based quantum computation. Furthermore, we use our two-dimensional cluster states to implement the Deutsch-Jozsa algorithm within the measurement-based quantum computation framework, achieving higher output-state fidelity compared with traditional circuit-based models and a query efficiency advantage over classical approaches. Our work establishes a scalable platform that combines large-scale entanglement generation, symmetry-protected topological order and practical quantum algorithms to enable robust, fault-tolerant measurement-based quantum computation.

quant-ph

Enhanced Superconductivity in Multilayer FeSe Films by Simplified Molecular Beam Epitaxy

Multi-unit-cell (UC) \b{eta}-FeSe films grown on SrTiO3(100) continue to attract attention because of the significant enhancement in the superconducting transition temperature (Tc) compared to that in bulk FeSe. In prior reports of molecular beam epitaxy (MBE)-grown \b{eta}-FeSe/SrTiO3(100), elaborate growth protocols have been used to achieve enhanced Tc, leading to a general belief that careful pre-treatment of the SrTiO3 substrate and post-growth annealing in ultrahigh vacuum (UHV) are essential. Here, we report a greatly simplified protocol for the MBE growth of superconducting multi-UC \b{eta}-FeSe films on SrTiO3(100), eliminating the need for careful substrate pre-treatment and post-growth UHV annealing while still achieving an enhanced Tc. With appropriate capping, epitaxial films with 14 UC thickness exhibit a zero-resistance transition temperature Tc ~ 20 K in ex situ electrical transport measurements. The MBE optimization process is guided by the growth-parameter dependencies of film morphology and structural properties, as characterized by reflection high-energy electron diffraction, X-ray diffraction, atomic force microscopy, and scanning transmission electron microscopy.

cond-mat.supr-con

Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation

Benefiting from ultra-long behavior sequence modeling, existing recommender systems bring users a better experience via simultaneously considering their long-term and short-term interests. Nevertheless, extended sequence lengths introduce substantial burdens on training efficiency and serving throughput. Prior approaches typically utilize search-based or cluster-based compression on ultra-long sequences at the cost of fine-grained information, or rely on various lightweight target attention structures incapable of sufficient sequential feature extraction. In this paper, we balance the effectiveness and efficiency for ultra-long sequence modeling via full transformer modeling accompanied with a two-stage knowledge distillation framework. First, both teacher and student models take the full attention mechanism rather than pure target-sequence attention for effective sequence scaling. For student models, we propose several simple yet well-motivated token merge approaches, significantly compressing the sequence length while maintaining an acceptable performance. Then, a one-time teacher is heavily trained with full sequence tokens, further boosting the performance of student models via knowledge distillation. The proposed paradigm named TM20K has been successfully deployed in ByteDance's e-commerce advertising recommender system that extends the e-commerce sequence length to 20K, delivering substantial improvements in key business metrics (e.g., ADSS +1.036\%) while keeping the training and serving cost nearly the same as the online state-of-the-art model (e.g., serving latency only +5.6\%).

cs.IR

SieveIVF: Threshold-Aware IVF Execution for Large-Scale Training Data Deduplication

Embedding-based training data deduplication retrieves candidate duplicate edges above an application similarity threshold, but fixed-probe inverted-file (IVF) search ignores this predicate when giving every query the same partition budget. Across four Hunyuan workloads, qualifying neighbors appear early despite sharply varying search depths. We present SieveIVF, a threshold-aware IVF executor that stops after $W$ consecutive searches find no qualifying candidate. The systems challenge is to preserve partition-major batching when each query's remaining work depends on prior results. Continuous batching groups ready queries by partition. A lookahead scheduler layers on top, exposing only work committed by the stopping rule to increase concurrency without changing stopping decisions or returned results. We implement SieveIVF in Lance. At $W=8$, SieveIVF is $4.1$--$7.6\times$ faster than fixed-probe IVF on four 10M Hunyuan workloads and $6.1$--$8.4\times$ faster on two public 100M workloads under the same index and search parameters, with pooled filtered top-10 recall losses of $0.03$--$1.13$ percentage points on Hunyuan and $1.43$--$2.29$ percentage points on the public workloads. These results show how an application predicate can guide IVF work allocation without changing the index or bounded top-$k$ interface.

cs.DB

TEngineDB-V: An OLAP-Native Vector Search System for Large-$k$ Workloads at Tencent

Vector search systems are essential infrastructure for modern data-driven applications. Large-$k$ analytical vector search, which retrieves $k=10^3$--$10^5$ results for analytics (e.g., aggregation, filtering, joins), is increasingly important for emerging workloads, including LLM data management and advertising analysis at Tencent. Existing systems remain inadequate: specialized vector databases often cap $k$ (e.g., $k \leq 10^4$) to satisfy tail-latency constraints and offer limited analytical support, while OLAP systems typically embed per-segment vector indexes as black boxes, causing severe read/compute amplification and preventing native query optimization. This paper presents TEngineDB-V, an OLAP-native vector search system for large-$k$ workloads. TEngineDB-V makes vector search a first-class analytical primitive in Tencent's OLAP engine through a global segment-decoupled index materialized as relational tables, eliminating scatter-gather execution, reducing amplification, and enabling native storage optimizations. It decomposes IVFPQ-based search into relational operators, integrates OLAP optimizations, and introduces DPPQ, which combines direction-aware quantization with hierarchical residual refinement to improve recall while preserving relational efficiency. TEngineDB-V further incorporates index-aware query rewriting and a distributed-aware cost model for efficient distributed execution. Experiments show that TEngineDB-V achieves up to a $145\times$ speedup over competitive systems such as StarRocks, and up to a $52\times$ improvement in 10-billion-scale production deployments.

cs.DB

Randomization inference for stepped-wedge designs with noncompliance with application to a palliative care pragmatic trial

While palliative care is increasingly commonly delivered to hospitalized patients with serious illnesses, few studies have estimated its causal effects. Courtright et al. (2016) adopted a stepped-wedge cluster-randomized design to assess the effect of palliative care on a patient-centered outcome. The randomized intervention was a nudge to administer palliative care but did not guarantee receipt of palliative care, resulting in noncompliance. A subsequent analysis using methods suited for standard trial designs produced statistically anomalous results, as an intention-to-treat analysis found no effect while an instrumental variable analysis did (Courtright et al. 2024). This highlights the need for a more principled approach to address noncompliance in stepped-wedge designs. We provide a formal causal inference framework for the stepped-wedge design with noncompliance by introducing a relevant causal estimand and corresponding estimators and inferential procedures. Through numerical studies, we compare an array of estimators and provide practical guidance in choosing an analysis method. Finally, we apply our recommended methods to reanalyze the palliative care pragmatic trial, producing point estimates suggesting a larger effect than the original analysis, but intervals that did not reach statistical significance.

stat.ME

ViDS: Video Diffusion Shader using 3D Face Tracking

We introduce ViDS, a Video Diffusion Shader that leverages 3D face tracking for expressive and identity-preserving portrait animation. We first reconstruct the identity-specific 3DMM mesh from the reference image, and then animate it using expression and pose parameters from a driving video. Leveraging dense geometric cues from 3DMM normal maps, we employ a video diffusion model as a neural shader to synthesize lifelike portrait animations while preserving the appearance and identity of the reference image. We find that more accurate 3DMM tracking enables finer-grained expression control. We also introduce an autoregressive diffusion sampling process that extends generation beyond the model's native window while reducing discontinuities between adjacent clips. Compared with prior diffusion-based approaches for portrait animation that rely on landmark-based conditioning or implicit motion latents, our method achieves more detailed and consistent expression and pose control while faithfully preserving identity and appearance. Detailed ablation studies validate the effectiveness of our design choices. Project page: https://fusheng-ji.github.io/ViDS/

cs.CV

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friendly and comprehensive framework for researchers and developers to evaluate existing multi-modality models and publish \textbf{reproducible} evaluation results. In VLMEvalKit, we implement over 450+ large multi-modality model configurations, including both proprietary APIs and open-source models, and support 330+ benchmarks across diverse multi-modal benchmarks. By implementing a single interface, new models can be easily added to the toolkit, while the toolkit automatically handles the remaining workloads, including data preparation, distributed inference, prediction post-processing, and metric calculation. VLMEvalKit has also evolved to a broader evaluation suite spanning video/audio, document understanding, GUI grounding, spatial reasoning, safety, scientific reasoning, and multi-turn dialogue. Based on the evaluation results obtained with the toolkit, we host the OpenVLM Leaderboard, a comprehensive leaderboard to track the progress of multi-modality learning research. The toolkit is released on https://github.com/open-compass/VLMEvalKit and is actively maintained.

cs.CV

Surface code logical operations on a superconducting quantum processor

Fault-tolerant quantum computation requires logical operations that manipulate encoded information while preserving quantum error-correction protection. In planar surface-code architectures, code deformation and lattice surgery provide a local, measurement-based route to such operations. Here we experimentally realize key elements of patch-based surface-code logical processing on a 107-qubit superconducting quantum processor. We first implement a reusable primitive layer comprising merge and split, patch expansion and shrinkage, and deformations mediated by domain walls and twist defects. We then compose these primitives to realize logical state routing, the logical controlled-NOT gate, and the single-qubit Hadamard and phase gates, which together form a Clifford-generating set. All operations are implemented on distance-three rotated surface-code patches with multi-round syndrome extraction and neural-network decoding, without post-selection. Our results advance superconducting surface-code experiments from protected logical memory to active, patch-based fault-tolerant logical operations.

quant-ph

LearniBridge: Learnable Calibration of Feature Caching for Diffusion Models Acceleration

Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs. Feature caching accelerates inference by reusing intermediate representations. Existing methods rely on historical features for implementation simplicity, yet suffer from severe error accumulation at high acceleration ratios. To address this limitation, we investigate the nature of the requisite feature correction. We demonstrate that the optimal calibration update is characterized by a shared low-rank subspace across diverse prompts. Guided by this structural insight, we propose LearniBridge, a learnable calibration mechanism for feature caching that bridges multiple timesteps through lightweight LoRA updates. This mechanism enables effective calibration requiring only 3-5 training samples. Extensive experiments on image and video generation show that LearniBridge achieves up to $5.87\times$, $5.75\times$, and $4.10\times$ acceleration on FLUX, HunyuanVideo, and WAN2.1, respectively. On WAN2.1, it improves VBench by 1.28% over the previous SOTA at $4.10\times$ acceleration. Our code is available at https://github.com/Iiiiiiirene/LearniBridge.

cs.CV

SCAN-Planner: Spatial Collision-Aware Local Planning for Route-Guided Long-Range Quadruped Navigation

Quadruped robots are increasingly expected to navigate through narrow passages, cluttered indoor scenes, and large-scale 3D unstructured environments. Existing local planners commonly approximate the robot using isotropic geometric inflation or rely on planar and elevation-map representations, leading to conservative motion in tight spaces and limited reasoning about overhanging structures. This letter presents SCAN-Planner, a spatial collision-aware local planning framework for long-range quadruped navigation. A yaw-aware twin-cylinder footprint is used to model the elongated robot body, enabling whole-body collision evaluation through sparse queries in an inflated 3D occupancy map. We further introduce a projected A* search that generates collision-free guidance on an interpolated ground-following surface, with z-gradient suppression to avoid obstacles horizontally while maintaining vertical stability. For large-scale deployment, a robot-centric sliding map with boundary fallback provides high-resolution local collision checking and recovery from local dead ends. Simulation and real-world experiments demonstrate that SCAN-Planner generates safe, smooth, and efficient trajectories in dense clutter, 3D unstructured scenes, stair traversal, and long-range navigation tasks.

cs.RO

Conflict-Aware Federated Fine-Tuning of Large Language Models with Mixture-of-Experts

The continuous scaling of large language models (LLMs) incurs prohibitive computational costs, making Mixture-of-Experts (MoE) a scalable alternative for efficient fine-tuning via sparse activation. While federated learning (FL) emerges as the paradigm for privacy-preserving collaborative optimization, integrating MoE into FL under data heterogeneity may trigger conflicting expert optimizations. Client-specific data distributions force same-indexed experts to optimize under inconsistent or even conflicting feature-label correlations. This mismatch induces destructive interference during aggregation, thus destabilizing the optimization trajectory and degrading model performance. To address this issue, we propose FC-MoE, a federated conflict-aware framework for MoE fine-tuning. It employs an importance aware weighting scheme to prioritize reliable local updates and utilizes gradient consensus projection to suppress conflicting updates, ensuring a stable global optimization path. Moreover, a local knowledge retention mechanism further preserves specialized client expertise by re-anchoring domain-specific residuals. Extensive experiments demonstrate that FC-MoE accelerates convergence and enhances both global and local model performance in non-IID federated environments.

cs.LG

Agon: A Semi-Supervised Framework for Robust Satellite Interference Detection

The rapid expansion of non-geostationary orbit (NGSO) satellites alongside existing geostationary orbit (GSO) systems has intensified spectrum congestion and inter-system interference, placing stringent demands on real-time interference management to sustain reliable coexistence in next-generation communication networks. While existing machine learning (ML)-based reconstruction models have made strides, they remain constrained to an area under the curve (AUC) of 0.83 due to fixed thresholds, causing unacceptable false alarm rates that undermine critical link reliability. Additionally, their decoupled training paradigm neglects cross-domain dependencies, limiting time and frequency-domain AUCs to 0.83 and 0.71, respectively. To address these limitations, this paper introduces a semi-supervised satellite interference detection framework named Agon, employing a novel two-stage hybrid learning paradigm. Agon integrates masked autoencoder (MAE) pre-training of a dual attention transformer (DAT) with multi-task fine-tuning to optimize a direct binary classifier, effectively eliminating unstable thresholds. Furthermore, it incorporates high-order statistics (HOS)-augmented attention and wavelet regularization to bolster noise robustness and structural fidelity. Extensive validation on public NGSO-GSO dataset and a high-fidelity NGSO-NGSO dataset demonstrates that Agon achieves state-of-the-art (SOTA) detection performance, with a 25.3% improvement in AUC. Moreover, the multi-task learning (MTL) framework facilitates accurate modulation classification with accuracies exceeding 90%, while simultaneously maintaining optimal detection performance across diverse scenarios characterized by varying off-axis angles and interference-to-noise ratios (INRs).

cs.NI

Aidos: A Hybrid Optimization Algorithm for Beam Hopping Scheduling in NGSO Mega-Constellations

With the rapid proliferation of non-geostationary orbit (NGSO) mega-constellations, beam hopping (BH) has become indispensable for resource scheduling in multi-satellite, multi-coverage scenarios. By dynamically adjusting spot beam power and pointing within each time slot, BH enables highly efficient spectrum utilization. A principal engineering challenge is the real-time generation of beam hopping time plans (BHTP). Traditional algorithms, such as the round-robin strategy, distribute beams evenly across all service cells in a round-robin fashion. However, real traffic follows a long-tail distribution; the most active 10% of hotspot cells generate more than 50% of the aggregate demand, making uniform allocation inadequate. To address this issue, existing frameworks adopt a genetic algorithm (GA), whose throughput is approximately 80.7% higher than the traditional baseline. Operational satellite footprints encompass more than 1,000 service cells. The GA requires 67.8 s to generate a BHTP for 1,127 cells. With a 550 km LEO satellite providing only a 300 s visibility window, multiple online recomputations are impractical. State-of-the-art algorithms, such as multi-agent deep reinforcement learning (MADRL), fail to converge once the cell count exceeds 200. To overcome these challenges, we propose a novel BH scheduling algorithm Aidos. The algorithm integrates traffic-aware random-key encoding into a multi-objective metaheuristic search, and then applies a sliding-window Beta resampling strategy during adaptive distribution evolution, to improve both the search efficiency and the solution quality of the BHTP. Experiments demonstrate that Aidos improves throughput by 79.2% and reduces latency by 99.45%. Its average computation time is 9.3 s, enabling online replanning within a 300 s satellite overpass window.

cs.NI

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a growing concern for fairness and information integrity. Research on generative engine optimization (GEO) has produced many manipulation methods, but each is evaluated on its own dataset with its own metrics, so their relative strength and detectability stay unclear. We present GEO-Bench, a benchmark that evaluates GEO ranking-manipulation attacks under one protocol. It unifies black-box prompt-based attacks (TAP, Zero-Shot), white-box gradient-based attacks (STS, RAF, StealthRank), and ten white-hat C-SEO strategies. We score every method on five datasets against a fixed open-weight ranker (Llama-3.1-8B-Instruct), using metrics for both effectiveness (NRG, Success@α, Promote@α) and stealth (keyword violation rate, perplexity ratio). Our evaluation shows that effectiveness and stealth trade off across adversarial attacks, that black-box content rewriting matches or exceeds gradient-based attacks on rank promotion while producing more fluent text and can evade both keyword- and perplexity-based detection on some domains, and that the access model does not predict attack strength. By standardizing datasets, attack implementations, and metrics, GEO-Bench enables the first direct comparison across these attack paradigms and supports the development of detection methods.

cs.CR

Rec-Distill: An Industrial Distillation Pipeline for Large-Scale Recommendation Models

Large recommendation models have demonstrated substantial potential gains under scaling laws, yet these gains are difficult to realize in industrial recommendation systems because real-world deployment requires lightweight models with strict serving efficiency and latency guarantees. This creates a fundamental gap between offline model scaling and online deployment. In this work, we present Rec-Distill, an industrial distillation pipeline that transfers the performance gains of large-scale recommendation modeling to efficient serving models. Rec-Distill combines large-teacher scaling with student-side transfer optimization through decoupled training, black-box distillation, debiasing mechanism, and a hybrid batch-streaming pipeline for dynamic recommendation environments. Across multiple recommendation and advertising scenarios on real-world platforms, our framework scales teacher models up to 24B dense parameters and 20K behavior sequence length, while enabling lightweight students to recover a substantial portion of teacher gains, with distillation transferability exceeding 60% in the best setting. Extensive offline and online experiments further show that these transferred gains consistently translate into measurable business improvements under industrial constraints. These results demonstrate that Rec-Distill provides a practical framework for distilling large-scale recommendation models into deployable, cost-efficient serving systems, while also establishing a reliable path toward scaling recommendation models to even larger regimes in the future.

cs.IR

Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-world deployment faces a fundamental device--cloud dilemma: on-device models are efficient but often brittle, while cloud models are stronger but costly in computation. State-of-the-art LLM device--cloud routers usually make coarse task-level decisions, which cannot adapt to the changing difficulty of multi-step agent interactions. To address this issue, we present Hera, a step-level device--cloud LLM agent coordinator for long-horizon tasks achieving a strong performance--cost Pareto frontier. Hera adopts a novel two-stage training paradigm: (1) imitation learning for cold-start, followed by (2) reinforcement learning that jointly optimizes task success and cloud usage efficiency. The first stage casts step-level routing as a supervised classification problem: the device agent is replayed on cloud trajectories, with each state labeled by the agreement between device and cloud actions. In the second stage, we perform cost-aware reinforcement learning by grouping identical states across trajectories and updating Hera with labels favoring higher expected return and fewer future cloud calls. We evaluate Hera on ALFWorld, WebShop, and AppWorld, where it consistently outperforms prior methods, achieving 92.5% of the cloud-only success rate with cloud use in only 46.3% of steps.

cs.AI

TravExplorer: Cross-Floor Embodied Exploration via Traversability-Aware 3-D Planning

Zero-shot Object Navigation (ZSON) has shown promise for open-vocabulary target search in unseen environments, yet most existing systems remain tied to planar representations and single-floor assumptions. These assumptions become inadequate in real buildings, where navigation involves floors, stairs, landings, and vertically overlapping spaces. This article presents TravExplorer, a cross-floor embodied exploration framework that couples zero-shot semantic guidance with traversability-aware 3-D planning. TravExplorer maintains a unified volumetric map that distinguishes occupied structures from robot-reachable support surfaces and extracts traversable frontiers from connected support surfaces, including floors, stairs, and landings. A FOV-aware active perception strategy further resolves incomplete observations during cross-floor traversal. To reduce semantic-reasoning latency, a lightweight guidance module aligns a probabilistic instance map from online open-vocabulary segmentation with a spatial value map from fast image-to-text matching. Based on these geometric and semantic memories, a hierarchical planner performs target-aware frontier touring over object hypotheses, traversable frontiers, and stair landmarks, and generates executable cross-floor motions through foothold-guided 3-D search and vertically constrained local trajectory optimization. Experiments over 4,195 simulated episodes on HM3D and MP3D demonstrate consistent advantages over representative ObjectNav baselines. Fifty real-world trials on a Unitree Go2 further validate open-vocabulary target search across single-floor and cross-floor indoor environments without prior maps or human intervention. The code will be released at https://github.com/wuyi2121/TravExplorer.

cs.RO