arXiv ScienceSearch

arXiv subjects

Liang Han

Publications and source records attributed to Liang Han.

At least 19 recordsLinked to original sources

Low-energy Muon-Nucleon scattering experiment: LUNE (White Paper)

The HIAF will provide high-intensity, high-quality muon beams with momenta from 0.5 to 7.5 GeV/c. This energy range is uniquely suited for precision muon scattering, bridging the gap between low-energy electron facilities and future high-energy lepton-ion colliders. In particular, HIAF will enable precision measurements with both positive and negative muon beams over a broad kinematic range, complementing existing electron-scattering facilities such as JLab, EicC and EIC. Based on HIAF muon source, the LUNE Collaboration has been established to address several fundamental questions in nuclear and particle physics, including the proton charge radius puzzle, nucleon electromagnetic structure, and the dynamics of quantum electrodynamics and hadronic interactions. The program proceeds in two phases, from elastic scattering to nucleon structure and beyond-Standard-Model searches. The experiment is expected to determine the proton charge radius with a precision of approximately 1.0\% using elastic muon-proton scattering. It will also perform systematic measurements of the proton electromagnetic form factors with both $\mu^+$ and $\mu^-$ beams, enabling precise studies of two-photon exchange effects and stringent tests of quantum electrodynamics. Beyond elastic scattering, LUNE will investigate TMD, gravitational form factors, and nuclear charge radii, providing new insights into the 3D structure of nucleons and nuclei. The experiment will further address important topics including Coulomb-distortion corrections, nuclear medium effects, and possible signatures of physics beyond the Standard Model. This white paper presents the scientific motivation, detector concept, expected performance, and long-term strategy of LUNE.

hep-ex

Surgical Video Generation From Diffusion to World Models: A Survey

Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.

cs.CV

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.

cs.LG

Correct-by-Construction Behavior Tree Synthesis from Signal Temporal Logic Specifications with Application to Robotic Missions

Behavior Trees (BTs) are widely adopted for complex task execution in robotics, providing modular, reactive control but lacking formal guarantees. However, existing correct-by-construction synthesis from Linear Temporal Logic (LTL) cannot express quantitative timing constraints. This letter synthesizes correct-by-construction BTs from Signal Temporal Logic (STL) specifications. The workspace is modeled as a timed transition system and abstracted into a zone graph, and an augmented state space tracking both logical progress and timing constraints is introduced. A hierarchical fixed-point algorithm computes winning sets for an STL fragment encompassing safety, reachability, response, recurrence, and persistence, yielding BT subtrees with a runtime constraint function. Correctness guarantees are proven and complexity bounds are derived. Simulations demonstrate specification satisfaction with strictly positive robustness, and a physical quadrotor experiment with six STL specifications validates practical deployability.

cs.RO

MixDiffusion: Mixing Diffusion-based Uni-condition Text-to-Image Generation Models for Multi-condition Image Synthesis

Recent advances in text-to-image (T2I) generation have enabled controllable image synthesis by incorporating conditions beyond text. However, most existing diffusion-based methods are limited to a single type of control condition (e.g., bounding boxes or keypoints), which restricts their flexibility. To address this limitation, we propose MixDiffusion, a training-free diffusion framework for multi-condition T2I generation. MixDiffusion theoretically supports an arbitrary number of control conditions, including bounding boxes, keypoints, sketches, depth maps, reference images, and text, by collaboratively integrating multiple pre-trained uni-condition diffusion models. The key insight of the proposed approach is to derive the predicted noise distribution in each denoising step of the diffusion-based multi-condition image generation model from the predicted noise distributions of multiple diffusion-based uni-condition models with a derived integration formula, which is supported by rigorous theory proof. Owing to its training-free nature, MixDiffusion is easy to deploy and readily extensible to new control modalities.

cs.CV

Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors

3D reconstruction from sparse views is a challenging task in 3D computer vision. Recent studies on 3D Gaussian Splatting (3DGS) have achieved remarkable results with sparse views in novel view synthesis, yet reconstructing high-quality geometric surfaces from sparse views remains a challenge, due to the limited geometry clues and the discreteness of Gaussians. In this paper, we propose a novel 3DGS-based method for high-fidelity surface reconstruction from sparse views. Our key insight is to introduce a normal-guided depth propagation approach, which can extend depth information from high-confidence regions to constrain the depth in low-confidence areas. Additionally, we propose an abnormal depth edge-aware regularization to address depth discontinuities caused by the discreteness of Gaussians. Extensive experiments on DTU and Tanks-and-Temples datasets demonstrate that our method outperforms the state-of-the-art methods in sparse view surface reconstruction. Project page: https://hanl2010.github.io/DP-GS.

cs.CV

City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion

Multi-view 3D surface reconstruction is a longstanding challenge in computer vision. Although recent large-scale reconstruction methods based on 3D Gaussian Splatting (3DGS) achieve impressive novel-view synthesis, producing high-quality surfaces over large scenes remains difficult, due to complex geometry, long optimization, and limited memory. In this paper, we propose a novel yet simple partitioning method to efficiently and faithfully reconstruct large-scale scene surfaces. Our key insight lies in a scene partitioning method based on viewpoint orientation. This partitioning approach ensures that views with similar orientations are jointly involved for more accurate depth estimations, leading to precise surface reconstructions and balanced computation on multiple GPUs in parallel. In addition, we propose a strategy to detect and repair missing regions in the initial point cloud caused by sparse viewpoints or insufficient textures, thereby further improving the geometric quality. Extensive experiments on the GauU-Scene, MatrixCity, and UrbanScene3D datasets demonstrate that our method outperforms the state-of-the-art approaches in surface reconstruction for large-scale scenes. Project page: https://hanl2010.github.io/VOP-GS.

cs.CV

FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation

Pixel-space diffusion has re-emerged as a promising alternative to latent-space generation because it avoids the representation bottleneck introduced by VAEs. Yet most existing methods still treat image generation as a frequency-homogeneous process, overlooking the distinct roles and learning dynamics of low- and high-frequency components. To address this, we propose FREPix, a FREquency-heterogeneous flow matching framework for Pixel-space image generation. FREPix explicitly decomposes generation into low- and high-frequency components, assigns them separate transport paths, predicts them with a factorized network, and trains them with a frequency-aware objective. In this way, coarse-to-fine generation becomes an explicit design principle rather than an implicit behavior. On ImageNet class-to-image generation, FREPix achieves competitive results among pixel-space generation models, reaching 1.91 FID at $256\times256$ and 2.38 FID at $512\times512$, with particularly strong performance in the early stages of training and in the low-NFE regime.

cs.CV

Doubly charged Higgs production within the Higgs triplet model at future electron-positron colliders

We investigate in detail the discovery potential of the doubly charged Higgs boson at the Compact Linear Collider in $e^-e^-$, $e^-\gamma$, $\gamma\gamma$, and $e^+e^-$ collision modes, within the Higgs triplet model at two extreme benchmark points as representatives of the Yukawa-like and gauge-like regions. In the Yukawa-like region, the most promising production mechanism is the single production via $e^-e^-$ and $e^-\gamma$ collisions. Given the subsequent decay of the doubly charged Higgs into a same-sign lepton pair, CLIC can achieve statistical significance well beyond the discovery threshold, within the parameter space permitted by experimental constraints. In the gauge-like region, with the $\ell^{\pm}\ell^{\pm} + \geq 3j$ final state, CLIC exhibits robust discovery potential for the doubly charged Higgs boson, up to a mass of approximately $1.2~\mathrm{TeV}$. We also investigate the search for doubly charged Higgs at the HL-LHC. Our results demonstrate that CLIC possesses greater advantages and offers superior discovery potential for the doubly charged Higgs boson, compared to the HL-LHC.

hep-ph

4C4D: 4 Camera 4D Gaussian Splatting

This paper tackles the challenge of recovering 4D dynamic scenes from videos captured by as few as four portable cameras. Learning to model scene dynamics for temporally consistent novel-view rendering is a foundational task in computer graphics, where previous works often require dense multi-view captures using camera arrays of dozens or even hundreds of views. We propose \textbf{4C4D}, a novel framework that enables high-fidelity 4D Gaussian Splatting from video captures of extremely sparse cameras. Our key insight lies that the geometric learning under sparse settings is substantially more difficult than modeling appearance. Driven by this observation, we introduce a Neural Decaying Function on Gaussian opacities for enhancing the geometric modeling capability of 4D Gaussians. This design mitigates the inherent imbalance between geometry and appearance modeling in 4DGS by encouraging the 4DGS gradients to focus more on geometric learning. Extensive experiments across sparse-view datasets with varying camera overlaps show that 4C4D achieves superior performance over prior art. Project page at: https://junshengzhou.github.io/4C4D.

cs.CV

InsTraj: Instructing Diffusion Models with Travel Intentions to Generate Real-world Trajectories

The generation of realistic and controllable GPS trajectories is a fundamental task for applications in urban planning, mobility simulation, and privacy-preserving data sharing. However, existing methods face a two-fold challenge: they lack the deep semantic understanding to interpret complex user travel intent, and struggle to handle complex constraints while maintaining the realistic diversity inherent in human behavior. To resolve this, we introduce InsTraj, a novel framework that instructs diffusion models to generate high-fidelity trajectories directly from natural language descriptions. Specifically, InsTraj first utilizes a powerful large language model to decipher unstructured travel intentions formed in natural language, thereby creating rich semantic blueprints and bridging the representation gap between intentions and trajectories. Subsequently, we proposed a multimodal trajectory diffusion transformer that can integrate semantic guidance to generate high-fidelity and instruction-faithful trajectories that adhere to fine-grained user intent. Comprehensive experiments on real-world datasets demonstrate that InsTraj significantly outperforms state-of-the-art methods in generating trajectories that are realistic, diverse, and semantically faithful to the input instructions.

cs.AI

Flavour asymmetry of antiquarks in nucleon and nucleus

Over the years, comprehensive experiments have shown a fact that the nucleons, such as the proton and neutron, are formed by not only the "valence" up and down quarks which were thought to comprise the nucleons in a simple constituent picture, but also "sea" quarks which can be any other flavour. However, it is still unknown how sea quarks are generated inside the nucleons. Since 1990s, measurements on high energy deuterons (formed by a proton and a neutron) indicated that the anti-down quark contribution was higher than the anti-up quark in the proton, based on the assumptions of the proton-neutron isospin symmetry and a small nuclear effect of the deuteron. Henceforth, sea quarks are considered to be generated via some flavour-asymmetrical mechanisms. Here we report an analysis on a series of new measurements from pure proton interactions which are free from those assumptions, unexpectedly showing that the anti-down quark component is rather consistent with the anti-up quark. It appears to be evidence that the previously observed asymmetry was caused by an unknown nuclear effect in the deuteron, rather than by a difference between antiquarks. We anticipate this work to be an essential new discovery and a motivation for studying nuclear structure, both experimentally and theoretically, at high energy scales, as it now appears fundamentally different from our understanding established in the past.

hep-ex

Impact of new measurements of light quarks at hadron colliders

Recently a series of new measurements with both the neutral and charge current Drell--Yan processes have been performed at hadron colliders, showing deviations from the predictions of the current parton distribution functions (PDFs). In this article, the impact of these new measurements is studied by using their results to update the PDFs. Although these new measurements correspond to different boson propagators and colliding energies, they are found to have a similar impact to the light quark parton distributions with the momentum fraction $x$ around 0.1. It manifests that the deviations are consistent with each other and favor a larger valence $d_v/u_v$ ratio than the modern PDF predictions. Further study indicates that such tension arises dominantly from the deep inelastic scattering measurements of NMC and the fixed target experiments of NuSea, both of which play pivotal roles in detecting the relative $u$ and $d$ type quark contributions for modern PDFs. According to the conclusions of the impact study, it would be essential to include these new measurements into the complete PDF global analysis in the future.

hep-ph

SparseRecon: Neural Implicit Surface Reconstruction from Sparse Views with Feature and Depth Consistencies

Surface reconstruction from sparse views aims to reconstruct a 3D shape or scene from few RGB images. The latest methods are either generalization-based or overfitting-based. However, the generalization-based methods do not generalize well on views that were unseen during training, while the reconstruction quality of overfitting-based methods is still limited by the limited geometry clues. To address this issue, we propose SparseRecon, a novel neural implicit reconstruction method for sparse views with volume rendering-based feature consistency and uncertainty-guided depth constraint. Firstly, we introduce a feature consistency loss across views to constrain the neural implicit field. This design alleviates the ambiguity caused by insufficient consistency information of views and ensures completeness and smoothness in the reconstruction results. Secondly, we employ an uncertainty-guided depth constraint to back up the feature consistency loss in areas with occlusion and insignificant features, which recovers geometry details for better reconstruction quality. Experimental results demonstrate that our method outperforms the state-of-the-art methods, which can produce high-quality geometry with sparse-view input, especially in the scenarios with small overlapping views. Project page: https://hanl2010.github.io/SparseRecon/.

cs.CV

Relative difference between up and down quark structure of the proton

We presen a novel determination of the down-to-up composition ratio using the forward-backward asymmetry observed in the proton-proton collisions at the LHC. This method offers unique insights into the flavor-specific difference between down and up quarks, which are difficult to isolate in traditional cross-section measurements due to the inherent mixing of contributions from both flavors. In this study, we systematically measure the down-to-up quark ratio over a broad momentum fraction (x) range of 0.01 to 0.1, utilizing the sensitivity of the forward-backward asymmetry to quark-level couplings. Our findings reveal significant deviations in both the value and x-dependence of this ratio compared to predictions from current parton distribution functions (PDFs). These discrepancies highlight potential limitations in existing PDF parameterization and emphasize the importance of flavor-separated measurements for advancing our understanding of proton structure.

hep-ph

A $p_T$-ratio observable for studies of intrinsic transverse momentum of partons from Drell-Yan $p_T$ spectra

The determination of the intrinsic transverse momentum distribution of partons is central both for applications of parton shower Monte Carlo generators and for QCD studies of transverse momentum dependent (TMD) parton densities. Valuable information on this distribution is provided by experimental measurements of Drell-Yan transverse momentum $p_T$, in the region of low transverse momenta, with fine binning in $p_T$. However, such fine-binning measurements are challenging, as they require an extremely delicate control of systematic uncertainties. We suggest a $p_T$ observable based on measuring ratios between cross sections of suitably defined low-$p_T$ and high-$p_T$ regions. This observable does not rely on any dedicated partition of bins and has lower systematic uncertainties, and is shown to provide a good sensitivity to the intrinsic transverse momentum.

hep-ph

CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion

Stable Diffusion has advanced text-to-image synthesis, but training models to generate images with accurate object quantity is still difficult due to the high computational cost and the challenge of teaching models the abstract concept of quantity. In this paper, we propose CountDiffusion, a training-free framework aiming at generating images with correct object quantity from textual descriptions. CountDiffusion consists of two stages. In the first stage, an intermediate denoising result is generated by the diffusion model to predict the final synthesized image with one-step denoising, and a counting model is used to count the number of objects in this image. In the second stage, a correction module is used to correct the object quantity by changing the attention map of the object with universal guidance. The proposed CountDiffusion can be plugged into any diffusion-based text-to-image (T2I) generation models without further training. Experiment results demonstrate the superiority of our proposed CountDiffusion, which improves the accurate object quantity generation ability of T2I models by a large margin.

cs.CV

MonoInstance: Enhancing Monocular Priors via Multi-view Instance Alignment for Neural Rendering and Reconstruction

Monocular depth priors have been widely adopted by neural rendering in multi-view based tasks such as 3D reconstruction and novel view synthesis. However, due to the inconsistent prediction on each view, how to more effectively leverage monocular cues in a multi-view context remains a challenge. Current methods treat the entire estimated depth map indiscriminately, and use it as ground truth supervision, while ignoring the inherent inaccuracy and cross-view inconsistency in monocular priors. To resolve these issues, we propose MonoInstance, a general approach that explores the uncertainty of monocular depths to provide enhanced geometric priors for neural rendering and reconstruction. Our key insight lies in aligning each segmented instance depths from multiple views within a common 3D space, thereby casting the uncertainty estimation of monocular depths into a density measure within noisy point clouds. For high-uncertainty areas where depth priors are unreliable, we further introduce a constraint term that encourages the projected instances to align with corresponding instance masks on nearby views. MonoInstance is a versatile strategy which can be seamlessly integrated into various multi-view neural rendering frameworks. Our experimental results demonstrate that MonoInstance significantly improves the performance in both reconstruction and novel view synthesis under various benchmarks.

cs.CV