arXiv ScienceSearch

arXiv subjects

Kai Yan

Publications and source records attributed to Kai Yan.

At least 19 recordsLinked to original sources

Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs

Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room structure, associate objects across incomplete observations, and recover a globally consistent spatial configuration. Previous methods mainly focus on 3D scene generation with text input or require continuous visual inputs with additional priors, \ e.g., human-annotated masks or accurate 3D layouts, which makes these methods labor demanding and hard to apply in general cases. We present \textsc{Scenix}, a sparse-view 3D scene reconstruction framework via executable scene programs, a structured representation that can be directly instantiated into editable 3D scenes. Given sparse views, \textsc{Scenix} predicts executable scene programs through perception-grounded asset instantiation and closed-loop spatial refinement. % We present \method, a framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement. To support this task, we construct \dataset, a dataset of approximately 110,000 synthetic and real indoor scenes with multiview imagery, room structures, object-centric descriptions, and metric spatial annotations. We further introduce observation-consistent supervision that aligns each target scene with the visual evidence available in its input views. Experiments on held-out \textsc{XScene} scenes, real indoor images, and out-of-distribution SpatialGen cases evaluate structured scene prediction, object grounding, and spatial refinement.

cs.CV

On a two-dimensional Camassa-Holm-Zakharov-Kuznetsov equation

This paper is devoted to a new two-dimensional nonlinear dispersive wave model named as the Camassa-Holm-Zakharov-Kuznetsov (CH-ZK) equation which combines the nonlinear structure of the Camassa-Holm equation with the transverse Laplacian dispersion of the Zakharov-Kuznetsov equation. We first establish the local well-posedness of its Cauchy problem in a suitable Sobolev space and derive a blow-up criterion for strong solutions. Then the finite-time blow-up strong solutions for the CH-ZK equation have been constructed. Moreover, we prove a unique continuation property for the solutions to the CH-ZK equation. Finally, we investigate the existence of both peaked and smooth solitary waves, and obtain a rigidity theorem for the traveling wave solutions according to the magnitude of wave speed.

math.AP

Sharp well-posedness and ill-posedness of the Camassa-Holm equation in critical Triebel-Lizorkin spaces

This paper is devoted to the sharp well-posedness and ill-posedness of the Cauchy problem for the Camassa-Holm (CH) equation in critical Triebel-Lizorkin spaces $F^{1+\frac{1}{p}}_{p,q}(\mathbb{R})$ with $(p,q)\in[1,\infty)\times[1,\infty]$ or $p=q=\infty$. On the one hand, we establish the local well-posedness in the sense of Hadamard in $F^2_{1,q}(\mathbb{R})$ for $1\leq q<\infty$ via Lagrangian coordinate transformation. On the other hand, by means of smooth atomic decomposition, strong ill-posedness is then proved in $F^{1+\frac{1}{p}}_{p,q}(\mathbb{R})$ with $(p,q)\in(1,\infty)\times[1,\infty]$ or $p=q=\infty$ in the sense of norm inflation, which in particular yields the ill-posedness of CH in critical Sobolev spaces $W^{1+\frac{1}{p},p}(\mathbb{R})$ with $1<p<\infty$, and provides a new perspective on the ill-posedness of CH in $H^{\frac{3}{2}}(\mathbb{R})$.

math.AP

On a rod-Kadomtsev-Petviashvili shallow water equation in two dimensions

In this paper, we derive a new two-dimensional rod-Kadomtsev-Petviashvili (rod-KP) equation from the incompressible and irrotational three-dimensional Euler equation under the shallow water scaling. We establish the local well-posedness of the Cauchy problem in a suitable Sobolev space and derive the blow-up criterion for strong solutions via the energy method. Then, by exploring two appropriate conservation laws of the rod-KP equation, we are able to construct its global strong solution when the physical dimensionless parameter $\sigma=0$, and on the other hand produce the finite-time blow up solutions under some certain conditions when $\sigma \neq 0$. Furthermore, we present a uniqueness continuation property for the solutions. Finally, we investigate the existence of traveling-wave solutions in order to highlight the influence of weak transverse effects on wave stability, and we also exhibit the symmetry of solitary waves in the propagation direction.

math.AP

Transformer refined quantum sampling for strongly correlated electronic structure

Although quantum computing offers a promising solution for strongly correlated system simulation, existing algorithms face significant bottlenecks on current noisy intermediate-scale quantum (NISQ) devices. Here, we introduce QiankunNet-QSCI, a hybrid quantum-classical framework that addresses this challenge by combining efficient quantum-sampling with a transformer neural network. An efficient unitary selected configuration Interaction (USCI) ansatz especially designed for quantum sampling is proposed to identify the most chemically significant electronic configurations on the Zuchongzhi 3.1 quantum processor. Subsequently, the transformer model QiankunNet learns from these sparse yet critical quantum data to infer and reconstruct the complete electronic wavefunction with high fidelity. Simulation of the challenging 40-qubit [2Fe-2S] ferredoxin active center achieves chemical accuracy. Simulation of the nitrogenase P-cluster in a 114-electron 73-orbital active space also reaches 12 milli-Hartree-level agreement with the best density matrix renormalization group (DMRG) result. QiankunNet-QSCI thus offers a practical route to accurate quantum-assisted electronic structure calculations on current devices.

quant-ph

Real-Time Neural Hair G-Buffer Anti-Aliasing

We propose a lightweight real-time method for reconstructing strand-based hair G-Buffers from severely undersampled rasterized inputs. Our pipeline first applies neural spatial reconstruction and temporal accumulation to recover hair coverage, i.e., fractional hair visibility within a pixel, and tangent. It then uses a tangent-guided reconstruction step to complete the position, which is subsequently used for physically based deferred hair shading. We evaluate our method across a diverse set of hairstyles, including straight, wavy, afro, and ponytail styles, under both static and dynamic scenarios. Our method achieves higher hair reconstruction quality than general industrial neural reconstruction solutions such as DLSS and FSR.

cs.GR

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance

Reinforcement Learning with Verifiable Rewards (RLVR) has achieved great success in developing Large Language Models (LLMs) with chain-of-thought rollouts for many tasks such as math and coding. Nevertheless, RLVR struggles with sample efficiency on difficult problems where correct rollouts are hard to generate. Prior works propose to address this issue via demonstration-guided RLVR, i.e., to conduct Supervised FineTuning (SFT) when RL fails; however, SFT often requires a lot of data, which can be expensive to acquire. In this paper, we propose FEST, a FEw-ShoT demonstration-guided RLVR algorithm. It attains compelling results with only 128 demonstrations randomly selected from an SFT dataset. We find that three components are vital for the success: supervised signal, on-policy signal, and decaying weights on the few-shot SFT dataset to prevent overfitting from multiple-epoch training. On several benchmarks, FEST outperforms baselines with magnitudes less SFT data, even matching their performance with full dataset.

cs.LG

Commutator estimates and their applications to the transport-type equations

In this paper, we derive new commutator estimates in the Triebel-Lizorkin spaces by employing Bony's para-product decomposition, the Nikol'skij representation, and the Fefferman-Stein vector-valued maximal function. These estimates are then applied to develop a general theory for transport equations. Although analogous results are already available in the setting of Besov spaces, the methods developed there do not carry over directly to the Triebel-Lizorkin case. Our approach works for Triebel-Lizorkin spaces and, as a byproduct, also yields the corresponding results in Besov spaces. All proofs are presented in a unified manner that applies to both scales of function spaces, thereby extending and sharpening previous results on transport equations in these frameworks. Furthermore, the general theory we obtain is widely applicable to evolution equations, including incompressible and compressible ideal fluid flows, shallow water waves, and related models. As an illustration, we consider the two-component Euler-Poincar\'e system. Using the theoretical framework developed herein, we establish its local well-posedness and a blow-up criterion in both sub-critical and critical Triebel-Lizorkin spaces.

math.AP

Autonomous Laparoscope Control through Unified Mechanics-Based Representation of Multimodal Intraoperative Information

Laparoscope-holding robots can provide surgeons with a stable laparoscopic field of view (FOV) and reduce the burden on human assistants. To maintain an ideal intraoperative FOV, the robot must continuously adjust the laparoscope pose according to intraoperative information. However, intraoperative multimodal signals, such as position, force/torque, and images, differ markedly in physical meaning and units, making it difficult to build a unified representation and to generate control commands that can be used directly for laparoscope control. To address this issue, we propose a laparoscope-holding robot control method based on unified mechanics modeling of multimodal information. First, we design mapping strategies for multiple intraoperative sources, including position, force/torque, and images, and unify them into an equivalent-wrench representation in the operational space. Then, using a task-priority scheme, we inject the wrenches into the task space and the null space, respectively, and synthesize laparoscope control commands via task-priority projection, thereby achieving consistent representation and coordinated fusion of multimodal information within a single framework. Finally, taking the intraoperative remote center of motion (RCM) position, force/torque sensor readings, and laparoscopic images as examples, we construct an RCM-constraint wrench to enforce the RCM geometric constraint and reduce the contact force at the trocar site, a laparoscope-manipulation wrench to enable compliant dragging, and an instrument-tracking wrench to achieve autonomous visual tracking of the instruments. Experiments on a surgical phantom and in vivo porcine trials demonstrate that the proposed method supports multi-task operation, including compliant laparoscope manipulation and autonomous instrument tracking, while maintaining the RCM constraint and reducing sustained trocar-site loading.

cs.RO

XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios

The acquisition of high-quality, action-aligned demonstration data remains a fundamental bottleneck in scaling foundation models for dexterous robot manipulation. Although robot-free human demonstrations (e.g., the UMI paradigm) offer a scalable alternative to traditional teleoperation, current systems are constrained by sub-optimal hardware ergonomics, open-loop workflows, and a lack of systematic data-mixing strategies. To address these limitations, we present XRZero-G0, a hardware-software co-designed system for embodied data collection and policy learning. The system features an ergonomic, virtual reality interface equipped with a top-view camera and dual specialized grippers to directly improve collection efficiency. To ensure dataset reliability, we propose a closed-loop collection, inspection, training, and evaluation pipeline for non-proprioceptive data. This workflow achieves an 85% data validity rate and establishes a transparent mechanism for quality control. Furthermore, we investigate the empirical scaling behaviors and optimal mixing ratios of robot-free data. Extensive experiments indicate that combining a minimal volume of real-robot data with large-scale robot-free data (e.g., a 10:1 ratio) achieves performance comparable to exclusively real-robot datasets, while reducing acquisition costs by a factor of twenty. Utilizing XRZero-G0, we construct a 2,000-hour robot-free dataset that enables zero-shot cross-embodiment transfer to a target physical robot, demonstrating a highly scalable methodology for generalized real-world manipulation.Our project repository: https://github.com/X-Square-Robot/XRZero-G0

cs.RO

Energy Correlators from Star Integrals via Mellin Space

We explore the Mellin space representation for the collinear limit of $N$-point energy correlators in ${\cal N}=4$ super-Yang-Mills theory. We show that these correlators can be written as integro-differential operators acting on star integrals: one-loop $n$-gons in $n$ dimensions. For the three-point energy correlator, we obtain the Mellin representation, use it to relate the correlator to the massive box integral, and show how to solve this relation to match with the expected result. For the four-point energy correlator, we obtain the Mellin representation and use it to write the correlator to a sum of various box and hexagon integrals in special kinematics. Our results provide a systematic method to relate higher-point energy correlators in the collinear limit to star integrals, which are known exactly.

hep-th

Two-Loop Spacelike Splitting Amplitudes in Full-Color QCD

The study of QCD scattering amplitudes in the collinear regime provides crucial insight into the factorization properties of hadronic cross sections. In this paper, we present the first complete results for two-loop spacelike splitting amplitudes in full-color QCD, in all partonic channels and helicity configurations. We confirm the universality of a class of contributions already found in N=4 super Yang--Mills (sYM) theory, and identify previously unknown sources of collinear factorization-violating (CFV) effects. Consistent with recent observations in N=4 sYM, all CFV contributions cancel in color-summed squared amplitudes, implying the universality of single-parton collinear factorization for jet cross sections at third order in QCD.

hep-ph

Non-Markovian Cosmic-Ray Pitch-Angle Transport from Mirror Interactions

Cosmic-ray pitch-angle transport in magnetohydrodynamic (MHD) turbulence is governed by the interplay between magnetic mirroring and gyroresonant scattering. We develop a guiding-center (GC) Langevin model with explicit mirror drift and gyroresonant diffusion to describe the pitch angle evolution. This model accurately captures our test-particle simulation results in three-dimensional MHD turbulence, driven both solenoidally and compressively. We find that magnetic mirroring can drive anomalous pitch-angle diffusion at large pitch angles (including $90^\circ$) with non-Markovian memory effects, which arises from trapping of particles in magnetic wells. Gyroresonant scattering controls the escape rate from these wells. Across $M_{\rm A}$, large-pitch-angle particles are jointly regulated by mirror trapping and gyroresonant escape, exhibiting a transition from anomalous to normal diffusive pitch-angle transport as scattering strengthens, whereas small-pitch-angle particles remain gyroresonance-dominated and diffusive throughout. The pitch angle transport is found to be dominated by the compressible perturbations with marginal influence from Alfv\'en modes. In compressible turbulence with realistic damping accounted for, transit time damping (TTD) treatment fully recovers mirror interactions.

astro-ph.HE

SAP: Segment Any 4K Panorama

Promptable instance segmentation is widely adopted in embodied and AR systems, yet the performance of foundation models trained on perspective imagery often degrades on 360{\deg} panoramas. In this paper, we introduce Segment Any 4K Panorama (SAP), a foundation model for 4K high-resolution panoramic instance-level segmentation. We reformulate panoramic segmentation as fixed-trajectory perspective video segmentation, decomposing a panorama into overlapping perspective patches sampled along a continuous spherical traversal. This memory-aligned reformulation preserves native 4K resolution while restoring the smooth viewpoint transitions required for stable cross-view propagation. To enable large-scale supervision, we synthesize 183,440 4K-resolution panoramic images with instance segmentation labels using the InfiniGen engine. Trained under this trajectory-aligned paradigm, SAP generalizes effectively to real-world 360{\deg} images, achieving +17.2 zero-shot mIoU gain over vanilla SAM2 of different sizes on real-world 4K panorama benchmark.

cs.CV

Latent Wasserstein Adversarial Imitation Learning

Imitation Learning (IL) enables agents to mimic expert behavior by learning from demonstrations. However, traditional IL methods require large amounts of medium-to-high-quality demonstrations as well as actions of expert demonstrations, both of which are often unavailable. To reduce this need, we propose Latent Wasserstein Adversarial Imitation Learning (LWAIL), a novel adversarial imitation learning framework that focuses on state-only distribution matching. It benefits from the Wasserstein distance computed in a dynamics-aware latent space. This dynamics-aware latent space differs from prior work and is obtained via a pre-training stage, where we train the Intention Conditioned Value Function (ICVF) to capture a dynamics-aware structure of the state space using a small set of randomly generated state-only data. We show that this enhances the policy's understanding of state transitions, enabling the learning process to use only one or a few state-only expert episodes to achieve expert-level performance. Through experiments on multiple MuJoCo environments, we demonstrate that our method outperforms prior Wasserstein-based IL methods and prior adversarial IL methods, achieving better results across various tasks.

cs.LG

Efficient photocatalytic CO2 Reduction to C2+ Products with Pt1-xPdxSn4 Dirac Nodal Arc Semimetal

The photochemical CO2 reduction reaction represents a zero-carbon pathway for converting CO2 into value-added chemicals, yet its industrial implementation has been constrained by low selectivity and product diversity. Dirac nodal arc semimetals characterized by ultrahigh carrier mobility with over 25000 cm2 V-1 s-1 offer a promising platform to search for efficient catalysts for CO2 conversion. Herein, we demonstrate that strategic Pt incorporation into PdSn4 optimizes the electronic structure and carrier dynamics of this Dirac semimetal. Experimental and theoretical analyses reveal that the resulting Pd-Sn-Pt local electronic structure redistributes charge density around Pd and Pt atoms, which facilitates C-C coupling via *OC-COH and *OC-CHOH intermediates and enhances carrier mobility by 40% versus the pristine PdSn4 single crystal. The optimized Pd0.4Pt0.6Sn4 single crystal achieves C2H4 with formation rate of 0.000328 mol g-1 h-1, product selectivity of 73.1% and electron-based selectivity of 89%. This work establishes electronic-structure-tunable Dirac semimetals as a new paradigm for multi-carbon photochemical CO2 reduction, providing a design strategy for next-generation photocatalysts.

cond-mat.mtrl-sci

Transport equation theory in the Triebel-Lizorkin spaces and its applications to the ideal fluid flows

In this paper, we develop a general theory for the transport equation within the framework of Triebel-Lizorkin spaces. We first derive commutator estimates in these spaces, dispensing with the conventional divergence-free condition, via the Bony paraproduct decomposition and vector-valued maximal function inequalities. Building on these estimates and combining the method of characteristics with a compactness argument, we then obtain the new a priori estimates and prove local well-posedness for the transport equation in Triebel-Lizorkin spaces. The resulting theory is applicable to a wide range of evolution equations, including models for incompressible and compressible ideal fluid flows, shallow water waves, among others. As an illustration, we consider the incompressible ideal magnetohydrodynamics (MHD) system. Employing the general transport theory developed here yields a complete local well-posedness result in the sense of Hadamard, covering both sub-critical and critical regularity regimes, and provides corresponding blow-up criteria for the ideal MHD equations in Triebel-Lizorkin spaces. Our results refine and substantially extend earlier work in this direction.

math.AP