arXiv ScienceSearch

arXiv subjects

Zhenyu Li

Publications and source records attributed to Zhenyu Li.

At least 19 recordsLinked to original sources

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typically separate geometry reconstruction, part reasoning, and articulation estimation into different stages. This separation can weaken consistency between shape, active parts, and motion, while also incurring substantial inference cost. We introduce Artic-O, an end-to-end, feed-forward framework for articulated object reconstruction via latent geometry learning. Instead of fitting geometry in image or view space, Artic-O maps sparse multi-state observations into a pretrained latent geometry space, where a frozen flow-matching decoder provides a complete-shape prior for recovering visible and occluded structures. To connect geometry with articulation, Artic-O fuses visual tokens, geometry latents, and point-wise decoder features in an image-grounded part-reasoning module for active-part segmentation and articulation prediction. We further train the model with a geometry-to-articulation curriculum and a decoupled two-pass strategy to balance reconstruction and part-level supervision. On PartNet-Mobility, Artic-O achieves strong reconstruction quality while being substantially more efficient than LARM, a strong prior method. It reduces Chamfer Distance, improves F-score, and achieves comparable or better articulation accuracy across most joint metrics, while reducing inference time from 9 minutes to about 0.3 seconds per object.

cs.CV

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows

Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets, memories, and handoff notes, from which downstream components act. For action-constraining state, topical retention is insufficient: an artifact may mention an unresolved condition while changing it from a requirement that must be resolved before execution into information that may merely inform the next action. We study this action-binding role as operational state preservation. Safety blockers provide a controlled instance because each source state has an explicit prerequisite, authority, fallback, and execution consequence. We condition on correct upstream identification, vary the handoff transformation, and evaluate an executor restricted to the resulting artifact. Across 1,296 controlled synthetic episodes, direct-handoff controls preserve every blocker, whereas compression, plan assimilation, convergence, ownership deferral, and precedent substitution repeatedly turn binding state into caveats or non-binding considerations. Normal handoff compression produces 100.0% deactivation and 54.2% forbidden action. Restoring all four state fields raises preservation to 100.0% and reduces forbidden action to 0.0%. Fixed-artifact interventions further separate preservation from containment: downstream verification eliminates forbidden action while artifact deactivation remains 95.3%. These results identify a state-transmission failure between information extraction and action. Handoff transformations can retain state content while weakening its constraints on downstream action. Semantic availability does not guarantee operational preservation.

cs.AI

Perturbative Variational Quantum Eigensolver via Reduced Density Matrices

Current noisy intermediate-scale quantum (NISQ) devices remain limited in their ability to perform accurate quantum chemistry simulations due to restricted numbers of high-fidelity qubits and short coherence times. To overcome these challenges, we introduce a reduced density matrix (RDM)-based perturbative variational quantum eigensolver (VQE) framework that augments active-space VQE with perturbation theory to recover electron correlation beyond the active space without increasing the qubit count or variational circuit depth. We formulate a fully coupled approach (VQE-PTs) and a diagonal approximation (VQE-PT). The former retains couplings among orthonormalized perturbers, whereas the latter neglects these couplings to simplify the classical post-processing. Numerical simulations of HF, N$_2$, and F$_2$ show that VQE-PTs provides a robust formulation across different molecular systems, while VQE-PT offers an efficient approximation. We further experimentally implement VQE-PT on the Quafu superconducting quantum processor for F$_2$, achieving a mean absolute error of 1.2 millihartree along the potential energy surface after error mitigation. These results demonstrate perturbative VQE as a practical framework for incorporating dynamic correlation in quantum chemistry simulations.

quant-ph

TurboRetry: Mitigating Large-Scale QUIC Handshake Floods with Off-the-Shelf DPU Offloading

The modern transport protocol QUIC is designed to enhance network performance and security, but it remains vulnerable to handshake flooding attacks. Such attacks exhaust CPU resources by forcing the server to perform expensive cryptographic operations via a large number of handshaking requests. QUIC provides a built-in defense mechanism, the Retry mechanism, to mitigate these attacks. However, our experiments reveal that it can still become a performance bottleneck under large-scale QUIC handshake floods due to substantial computational overhead. In this paper, we design and implement TurboRetry, a split design, that offloads the Retry mechanism onto DPUs to efficiently mitigate QUIC handshake floods. TurboRetry partitions the tasks of the Retry into two categories, and then assigns them to the DPUs and the host, respectively. To preserve QUIC semantics and reduce the coordination overhead, TurboRetry designs an extended Retry token format and an efficient cooperation scheme. In addition, TurboRetry offloads the connection authorization task to the on-path DPA to further improve both performance and security. Our evaluation shows that TurboRetry outperforms the host-side implementation by a wide margin, improving throughput by 10-20$\times$.

cs.CR

Embedded quantum computing for many-body surface reaction

Predictive simulations of catalytic interfaces require correlated electronic-structure treatments that describe localized chemical transformations while retaining the influence of the extended metallic environment. We introduce QC-DFET, a quantum-computing density-functional embedding framework that maps surface-reaction active spaces to compact, environment-aware qubit Hamiltonians. A reaction-consistent active-space protocol preserves orbital continuity along reaction coordinates, while quantum-selected configuration interaction based on measurements from the Zuchongzhi superconducting quantum processor and strongly contracted perturbation theory capture static and dynamic correlation. On Cu(111), QC-DFET treats active spaces up to 28 qubits and is validated through a hierarchy of experimentally constrained surface-chemistry challenges. H2 dissociation/desorption tests balanced bond breaking and recombination barriers, CO adsorption tests site selectivity and metal-adsorbate bonding, and formate hydrogenation tests competing hydrogenation branches with different kinetic and thermodynamic signatures. Across these cases, QC-DFET reproduces bidirectional H2 barriers, recovers the observed top-site preference and adsorption strength of CO, and reconciles the experimentally benchmarked H2COO* reverse barrier with the lower forward barrier to HCOOH*. These results establish embedded quantum computing as a practical route to correlated surface-reaction energetics.

quant-ph

A High-Performance Pauli-Algebra Framework for Large-Scale Quantum Simulations

Efficient manipulation of Pauli-algebraic objects is a key bottleneck in the classical emulation and benchmarking of quantum algorithms for chemistry and many-body physics. This bottleneck appears in Hamiltonian construction, variational ansatz preparation, expectation-value and gradient evaluation, and real-time propagation, all of which require repeated Pauli-algebra operations. Here, we present a high-performance Pauli-algebra framework tailored to quantum many-body and quantum-chemical simulations. The framework combines compact binary symplectic encoding, canonical coefficient reduction, and grouped sparse operator representations that exploit shared bit-flip patterns among Pauli strings. The resulting Julia/C\texttt{++} implementation accelerates Pauli multiplication, Hamiltonian construction, and operator--state multiplication in sparse and symmetry-adapted many-electron spaces. Benchmarks demonstrate efficient Hamiltonian construction, large-active-space VQE and ADAPT-VQE calculations, and real-time variational dynamics on modern multicore CPU and GPU architectures. These results show that structure-aware Pauli-algebra engines provide a scalable classical backend for developing and benchmarking quantum algorithms in quantum chemistry and many-body simulation.

quant-ph

False Positives Raised by Quantum Readout Error Mitigation

Quantum readout error mitigation is essential for noisy intermediate-scale quantum devices to achieve reliable data. The conventional approaches, conflating initialization errors with measurement errors, not only suppress the influence of measurement errors, but also strengthen that of initialization errors, which is a systematic bias grows exponentially with the qubit number. Here, we have proved that this effect causes severe fidelity overestimation for all stabilizer states and might lead to false positives in large-scale entangled state characterization. Similarly, the results from algorithms like the variational quantum eigensolver and time evolution also deviate negatively, and cover up other errors in the quantum circuit. These findings highlight the critical need for rigorous benchmarking and careful management of initialization errors. Consequently, we establish an upper bound for the tolerable initialization error rate to ensure effective error mitigation at a given system scale.

quant-ph

CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging

Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond functional correctness. While rubric-based LLM judges improve interpretability by decomposing evaluation into explicit criteria, most existing pipelines remain pointwise: they score each response independently and derive preferences by comparing aggregated scores. We show that this design is poorly matched to pairwise code preference prediction and can underperform a strong monolithic judge. We propose CriterAlign, a criterion-centric framework that adapts rubric-based judging to pairwise preference evaluation through direct criterion-level pairwise judgments, tie-driven criterion refinement, swap-consistency filtering, and final pairwise synthesis. We further introduce Human-Preference-Aligned Guidance (HPAG), synthesized offline from training examples by extracting recurring rationale gaps between human preferences and monolithic judge predictions, and injected into the criterion generator, criterion judge, and final judge. On BigCodeReward, CriterAlign improves a Qwen2.5-VL-32B monolithic judge from 60.4% to 66.3% accuracy, with ablations confirming the contributions of pairwise criterion design and HPAG.

cs.SE

Physically Motivated Ansatz for Open Fermionic Systems on Quantum Computer

Determining non-equilibrium steady states (NESS) of open fermionic systems is a fundamental problem akin to finding ground states of closed systems. To address this, variational quantum algorithms can be used to solve the Lindblad master equation, much like the Schrödinger equation, yet ansatz design for NESS remains challenging. Existing approaches rely mostly on hardware-efficient ansätze (HEA), which suffer from the barren plateau problem. Here, we introduce a physically motivated ansatz named NE-UCC. Numerical simulations demonstrate that NE-UCC reliably converges to the steady state even in strongly correlated regimes far from equilibrium, reducing the infidelity by up to ten orders of magnitude compared to HEA. Furthermore, NE-UCC facilitates the exploration of excited eigenmodes with specific symmetries.

quant-ph

Improved time-translationally invariant tensor network influence functional method for Anderson impurity problems

The Anderson impurity model (AIM) is of fundamental importance in condensed matter physics for studying strongly correlated phenomena. However, accurately simulating its long-time dynamics still remains a significant numerical challenge. A class of recently developed numerical approaches represents the Feynman-Vernon influence functional (IF), which encodes all the bath effects on the impurity, as a matrix product state (MPS) in the temporal domain. The computational cost of this approach is largely determined by the bond dimension $χ$ of the temporal MPS. In this work, we propose an efficient and accurate method that, when the hybridization function in the IF can be approximated as a sum of $n$ exponential functions, systematically constructs the IF as an MPS by multiplying $O(n)$ small MPSs, each with bond dimension $2$. Our method yields a worst case scaling of $χ$ as $2^{8n}$ and $2^{2n}$ for real- and imaginary-time evolution respectively. We demonstrate the performance of our method for two commonly used bath spectral functions, and show that the required bond dimensions are significantly smaller than the worst case.

cond-mat.str-el

Quantum key distribution without authentication and information leakage

Quantum key distribution (QKD) is the most extensively researched quantum cryptographic protocol, leveraging quantum phenomena to enable information-theoretically secure key establishment. Conventional QKD systems rely on authenticated classical post-processing channels to prevent impersonation attacks. However, QKD inherently lacks authentication capabilities, necessitating an external authentication mechanism. Furthermore, the classical post-processing steps introduce information leakage, exposing QKD to additional attack strategies and diminishing the final key rate. In this study, we introduce a novel QKD protocol that eliminates the need for separate authentication, prevents information leakage, and significantly improves the achievable key rate. By removing public classical channels entirely, our design attains perfect secrecy with the use of two reusable pre-shared keys.

quant-ph

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation

Subject-driven image generation faces an "Identity-Diversity Paradox", where strong identity preservation often leads to rigid and low-diversity outputs. We propose a post-training framework called DivRL that jointly optimizes identity consistency and structural diversity simultaneously by leveraging disentangled visual features from a robust similarity model. Specifically, we introduce a Negative Self-Similarity Measure (nSSM) to quantify structural diversity, and Visual Semantic Matching (VSM) to evaluate identity consistency. We propose an "Explore-and-Suppress" strategy that treats VSM as a gated constraint: the model freely explores structurally diverse configurations, and only samples that violate the identity threshold are penalized via a quadratic hinge loss. This converts identity preservation from a competing objective into a feasibility constraint, allowing nSSM and VSM to improve jointly. Experiments demonstrate that our method effectively pushes the model to generate both consistent and diverse images and improves structural diversity while maintaining comparable identity consistency through a gated optimization formulation.

cs.CV

LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning

We present Layered Ray Intersections (LaRI), a fully supervised method for occluded geometry reasoning from a single image. Unlike conventional depth estimation, which is limited to visible surfaces, LaRI predicts multiple surfaces intersected by the camera rays using layered point maps. Compared to the existing approaches that leverage neural implicit representations or iterative refinement, LaRI achieves complete scene reconstruction in one feed-forward pass, enabling efficient and view-aligned geometric reasoning to underpin both object-level and scene-level tasks. We further propose to predict the ray stopping index, which identifies valid intersecting pixels and layers from LaRI's output. To better underpin and evaluate this task, we build an annotation pipeline using rendering engines, construct annotations for five public datasets, including synthetic and real-world data covering 3D objects and scenes. As a generic method, LaRI's performance is validated in object-level and scene-level reconstruction tasks.

cs.CV

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visual context. Two bottlenecks limit current systems. First, existing tool-use harnesses treat images returned by search, browsing, or transformation as transient outputs, so intermediate visual evidence cannot be re-consumed by later tools. Second, training data is usually built by fixed curation recipes that cannot track the target agent's evolving capability. To address these challenges, we first introduce a visual-native agent harness centered on an image bank reference protocol, which registers every tool-returned image as an addressable reference and makes intermediate visual evidence reusable by later tools. On top of this harness, On-policy Data Evolution (ODE) runs a closed-loop data generator that refines itself across rounds from rollouts of the policy being trained. This per-round refinement makes each round's data target what the current policy still needs to learn. The same framework supports both diverse supervised fine-tuning data and policy-aware reinforcement learning data curation, covering the full training lifecycle of the target agent. Across 8 multimodal deep search benchmarks, ODE improves the Qwen3-VL-8B agent from 24.9% to 39.0% on average, surpassing Gemini-2.5 Pro in standard agent-workflow setting (37.9%). At 30B, ODE raises the average score from 30.6% to 41.5%. Further analyses validate the effectiveness of image-bank reuse, especially on complex tasks requiring iterative visual refinement, while rollout-feedback evolution yields more grounded SFT traces and better policy-matched RL tasks than static synthesis.

cs.CL

Adversarial Attacks on Robot Localization Systems via Deep Feature Perturbation

Robot localization systems are critical for autonomous navigation and safety. Adversarial perturbations can mislead these systems, resulting in mislocalization, navigation errors, or unsafe interactions, especially in mission-critical scenarios. This paper investigates the vulnerability of deep learning based localization pipelines to adversarial attacks. We propose a novel framework for generating adversarial queries that specifically target Product Quantization (PQ) in visual localization systems. Our method employs a Lightweight Product Quantization Network (LPQN) to perturb query feature encodings, misleading the retrieval process by returning semantically irrelevant database entries. Adversarial queries are generated via a two-phase procedure: a forward pass that perturbs feature distributions and a backward pass that refines the perturbation through optimization. The lightweight design of LPQN allows the creation of subtle yet highly effective perturbations with minimal computational overhead. Extensive experiments in both controlled and real-world robotic environments demonstrate that our approach substantially degrades PQN performance, exposing critical vulnerabilities in practical applications.

cs.CV

RIS-Assisted Survivable Backhaul Recovery in Small-Cell Systems

The increasing densification of small-cell networks substantially expands cable-based backhaul infrastructure, creating heightened vulnerability to cable link failures. This paper proposes a reconfigurable intelligent surface (RIS)-assisted backup framework that exploits a key insight: during backhaul cable failures, base station (BS) radio components remain functional, enabling wireless backhaul traffic redistribution. Our framework maintains network connectivity by redistributing disconnected BS backhaul traffic to neighboring BSs through RIS-assisted wireless links. To maximize survivability across varying traffic conditions, we formulate a joint optimization problem that maximizes total resolvable backhaul traffic by jointly deciding BS selection, RIS phase shifts, and precoding vectors. The inherent non-convexity arising from coupling and quadratic fractional term is addressed through an alternating optimization algorithm that iteratively solves tractable convex subproblems via quadratic transformation. Comprehensive numerical evaluations demonstrate that the proposed RIS-enhanced framework significantly improves survivability from 58% to 72% under challenging high-intensity hotspot traffic conditions. Moreover, RIS provides the greatest gains for antenna-constrained systems by extending coverage to access more spare capacity of the distant BSs as well as enhancing the signal strength. Consequently, high survivability is achieved even with only two antennas per BS under moderate traffic intensity.

cs.IT

Transformer refined quantum sampling for strongly correlated electronic structure

Although quantum computing offers a promising solution for strongly correlated system simulation, existing algorithms face significant bottlenecks on current noisy intermediate-scale quantum (NISQ) devices. Here, we introduce QiankunNet-QSCI, a hybrid quantum-classical framework that addresses this challenge by combining efficient quantum-sampling with a transformer neural network. An efficient unitary selected configuration Interaction (USCI) ansatz especially designed for quantum sampling is proposed to identify the most chemically significant electronic configurations on the Zuchongzhi 3.1 quantum processor. Subsequently, the transformer model QiankunNet learns from these sparse yet critical quantum data to infer and reconstruct the complete electronic wavefunction with high fidelity. Simulation of the challenging 40-qubit [2Fe-2S] ferredoxin active center achieves chemical accuracy. Simulation of the nitrogenase P-cluster in a 114-electron 73-orbital active space also reaches 12 milli-Hartree-level agreement with the best density matrix renormalization group (DMRG) result. QiankunNet-QSCI thus offers a practical route to accurate quantum-assisted electronic structure calculations on current devices.

quant-ph

DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation

Generating high-quality stereo videos requires consistent depth perception and temporal coherence across frames. Despite advances in image and video synthesis using diffusion models, producing high-quality stereo videos remains a challenging task due to the difficulty of maintaining consistent temporal and spatial coherence between left and right views. We introduce DissolveStereo, a novel framework for zero-shot stereo video generation that leverages video diffusion priors without requiring paired training data. Our key innovations include a noisy restart strategy to initialize stereo-aware latent representations and an iterative refinement process that progressively harmonizes the latent space, addressing issues like temporal flickering and view inconsistencies. Importantly, we propose the use of dissolved depth maps to streamline latent space operations by reducing high-frequency depth information. Our comprehensive evaluations, including quantitative metrics and user studies, demonstrate that DissolveStereo produces high-quality stereo videos with enhanced depth consistency and temporal smoothness. In terms of epipolar consistency, our method achieves an 11.7% improvement in MEt3R score over the current state-of-the-art. Furthermore, user studies indicate strong perceptual gains over the previous arts, with an 8.0% higher perceived frame quality and 10.9% higher perceived temporal coherence. Our code is in https://github.com/shijianjian/DissolveStereo.

cs.CV