arXiv ScienceSearch

arXiv subjects

Lingxiao Li

Publications and source records attributed to Lingxiao Li.

At least 19 recordsLinked to original sources

GenScale: A Benchmark for Relative Object Scale in Image Generation and Editing

Modern image generation and editing systems can produce photorealistic, prompt-aligned images, but still often render familiar objects at implausible relative sizes. To measure this failure mode, we introduce GenScale, a benchmark and evaluation protocol for real-world relative object scale in image generation and editing. GenScale contains 900 image-level entries and 1,643 pairwise anchor-target scale relations across common-object generation, human-product generation with metric dimensions, and scale correction from failed generations. We further design a human-calibrated ordinal judge for scalable pairwise scale evaluation. Last but not the least, we introduce Rescale, a model-agnostic post-processing agent for localized scale correction without modifying the source generator. Experiments reveal that state-of-the-art image generators and editors cannot reliably observe relative scale yet, while Rescale consistently improves scale plausibility across generated and edited images. Together, GenScale establishes relative object scale as a distinct, measurable, and actionable capability for image generation systems.

cs.CV

Robust Block Preconditioning for 3D nonlinear steady-state radiation transport equations

In this work, based on the discrete ordinate method, we propose a robust block preconditioning strategy for the 3D nonlinear steady-state radiation transport equation with heat diffusion term. The presence of the diffusive term of the temperature equation prevents its elimination into a single equation for the radiation intensity. To overcome this difficulty, all physical variables are assembled into a single monolithic linear system. The heat flux and temperature are treated as independent variables in a mixed $H(\mathrm{div})$-conforming finite element formulation. The equation for radiation intensity is discretised by a discontinuous Galerkin method with upwind flux, where a vectorial finite element space is used to couples the radiation intensity in different directions within each element. We then construct a Newton-Krylov iterative solver to solve the nonlinear equations, for which the core part is efficient preconditioning. To accelerate the convergence of Krylov's method, three block preconditioners are constructed, corresponding to different levels of approximation of the coupling between the temperature and radiation intensity. $P_{\mathrm{Schur}}$ retains the full coupling. $P_{\mathrm{Split}}$ drops the conductive contribution to the radiation block. $P_{\mathrm{BJ}}$ neglects the radiation-to-temperature coupling, retaining only the temperature-to-radiation coupling. Numerical experiments demonstrate the mesh independence and robustness of the proposed preconditioners.

math.NA

High-order Energy-stable and Charge-conservative Lagrangian FEM for 3D Incompressible Inductionless MHD equations with Variable Density

In this paper, we develop a high-order, energy-stable and charge-conservative Lagrangian finite element method for variable-density incompressible inductionless magnetohydrodynamic (MHD) equations. The method utilizes the moving high-order curved tetrahedral mesh to track the material interface. Second-order Backward Differentiation Formula (BDF2) is used for the temporal discretization of material derivative, together with the second-order Adams--Bashforth method (AB2) for the update of control points of the meshes. High-order isoparametric Taylor-Hood elements with grad-div stabilization are used for the velocity-pressure pair. While, to ensure the discrete charge conservation, high-order parametric $\BH(\Div)$-conforming element is adopted for the current density. In the absence of external force, the unconditional energy-stability of the fully discrete scheme is proven. Finally, 3D numerical experiments are conducted to confirm the expected high-order accuracy for smooth solutions, the energy stability property and the capability of the proposed method.

math.NA

Molecular Lead Optimization via Agentic Tool Planning

Drug discovery is a lengthy and resource-intensive process composed of multiple stages. Among these stages, lead optimization plays a critical role in transforming early hit compounds into viable drug candidates. This stage requires improving ADMET-related properties through subtle structural refinement while preserving key molecular substructures responsible for binding affinity to disease targets. Recent advances in artificial intelligence have shown promise in accelerating various aspects of drug discovery; however, most existing approaches to lead optimization rely on one-step molecular optimization, which fail to account for the long-term consequences of sequential design decisions. To address this limitation, we propose TRACE, a trajectory-aware, LLM-reasoning agent for molecular lead optimization that formulates tool selection as a sequential decision-making problem over action trajectories. Given a lead molecule and an optimization objective, TRACE makes trajectory-aware decisions over molecular optimization tools, enabling forward-looking refinement under structural constraints. Experiments on multiple ADMET optimization tasks show that our agent achieves higher optimization success, larger property improvements, and higher validity, while preserving molecular similarity compared to baseline models.

cs.LG

Moir\'e Video Authentication: A Physical Signature Against AI Video Generation

Recent advances in video generation have made AI-synthesized content increasingly difficult to distinguish from real footage. We propose a physics-based authentication signature that real cameras produce naturally, but that generative models cannot faithfully reproduce. Our approach exploits the Moir\'e effect: the interference fringes formed when a camera views a compact two-layer grating structure. We derive the Moir\'e motion invariant, showing that fringe phase and grating image displacement are linearly coupled by optical geometry, independent of viewing distance and grating structure. A verifier extracts both signals from video and tests their correlation. We validate the invariant on both real-captured and AI-generated videos from multiple state-of-the-art generators, and find that real and AI-generated videos produce significantly different correlation signatures, suggesting a robust means of differentiating them. Our work demonstrates that deterministic optical phenomena can serve as physically grounded, verifiable signatures against AI-generated video.

cs.CV

Observation of a Reconstructed Chern Insulator in Twisted Bilayer MoTe2

Twisted bilayer MoTe2 is a prototypical moire material in which long-wavelength superlattices amplify electron correlations, enabling a wealth of emergent quantum phases. To date, experimental efforts have focused primarily on small twist angles (typically smaller than 4deg ), whereas the larger-angle regime-where moire bands become more dispersive and correlations are reduced-has remained largely unexplored. Here we chart the topological phase space of tMoTe2 at a relatively large twist angle of approximately 4.54deg, accessing a moderately correlated regime with enhanced bandwidth. In contrast to small-angle devices that predominantly host fractional quantum anomalous Hall or spin Hall responses, we uncover multiple Chern-insulating states with C = 1 at moire fillings v = -1, -0.53 and -1/2. Strikingly, at v = -2/3 a magnetic field induces a fractional Chern insulator accompanied by an insulator-metal transition. Our results broaden the topological phase diagram of tMoTe2 and establish large-angle moire superlattices as a versatile platform for engineering robust topological states beyond the strong-correlation limit.

cond-mat.mes-hall

Structure-preserving Randomized Neural Networks for Incompressible Magnetohydrodynamics Equations

The incompressible magnetohydrodynamic (MHD) equations are fundamental in many scientific and engineering applications. However, their strong nonlinearity and dual divergence-free constraints make them highly challenging for conventional numerical solvers. To overcome these difficulties, we propose a Structure-Preserving Randomized Neural Network (SP-RaNN) that automatically and exactly satisfies the divergence-free conditions. Unlike deep neural network (DNN) approaches that rely on expensive nonlinear and nonconvex optimization, SP-RaNN reformulates the training process into a linear least-squares system, thereby eliminating nonconvex optimization. The method linearizes the governing equations through Picard or Newton iterations, discretizes them at collocation points within the domain and on the boundaries using finite-difference schemes, and solves the resulting linear system via a linear least-squares procedure. By design, SP-RaNN preserves the intrinsic mathematical structure of the equations within a unified space-time framework, ensuring both stability and accuracy. Numerical experiments on the Navier-Stokes, Maxwell, and MHD equations demonstrate that SP-RaNN achieves higher accuracy, faster convergence, and exact enforcement of divergence-free constraints compared with both traditional numerical methods and DNN-based approaches. This structure-preserving framework provides an efficient and reliable tool for solving complex PDE systems while rigorously maintaining their underlying physical laws.

physics.flu-dyn

Parametric charge-conservative mixed finite element method for 3D incompressible inductionless MHD equations on curved domains

This paper develops a charge-conservative mixed finite element method with optimal convergence rates for the stationary incompressible inductionless MHD equations on three-dimensional curved domains. The discretization employs the isoparametric Taylor-Hood elements with grad-div stabilization for the velocity-pressure pair, and parametric Brezzi-Douglas-Marini elements for the current density. Utilizing the Piola's transformation, the discrete current density is exactly divergence-free. By employing suitable extensions and projections, optimal a priori error estimates are derived in both the energy norm and the $L^2$-norm. Numerical experiments are presented to confirm the theoretical results.

math.NA

Optimal error estimate of an isoparametric upwind discontinuous Galerkin method for radiation transport equation on curved domains

In recent years, high-order finite element methods on high-order meshes have attracted considerable attention. This work investigates the isoparametric upwind discontinuous Galerkin method for the radiation transport equation on a bounded domain with a piecewise $C^{k+1}$ smooth curved boundary. We use the isoparametric mapping to approximate the curved domain and construct a curved upwind discontinuous Galerkin scheme. The first-order hyperbolic nature and the complexity introduced by non-affine transformation, lead to additional difficulties for geometric approximation, numerical stability and the optimal error estimate. To address these issues, with the help of an isoparametric auxiliary operator, we first prove that the bilinear form is continuous with respect to the DG norm when its first argument is the isoparametric projection error. Then the geometric approximation error of inflow boundary of original domain is precisely estimated. The error order between discrete normal vectors and the continuous ones are also proven. Finally, the rigorous analysis yields an optimal convergence rate in the DG norm. Two- and three-dimensional numerical tests are conducted to support the theoretical results.

math.NA

A divergence-free parametric finite element method for 3D Stokes equations on curved domains

The Stokes equations play an important role in the incompressible flow simulation. In this paper, a novel divergence-free parametric mixed finite element method is proposed for solving three-dimensional Stokes equations on domains with piecewise smooth boundaries. The flow velocity and pressure are discretized with high-order parametric Brezzi-Douglas-Marini elements and volume elements, respectively, on curved tetrahedral meshes. Utilizing the interior-penalty discontinuous Galerkin (IPDG) technique, we prove the inf-sup condition for the mixed finite element pair, and high-order optimal error estimates in the energy norm, with the help of the extension and transformation of the true solution to computational domain. Moreover, the discrete velocity is exactly divergence-free, meaning that div uh = 0 holds in the curved computational domain. Numerical experiments are conducted to support the theoretical analyses.

math.NA

Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation

We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video diffusion models conditioned solely on video-based scene context can perform complex spatial tasks. We validate on two tasks: scene navigation - following camera-pose instructions while remaining consistent with 3D geometry of the scene, and object grounding - which requires semantic localization, instruction following, and planning. Both tasks use video-only inputs, without auxiliary modalities such as depth or poses. With simple yet effective design choices in the framework and data curation, Video4Spatial demonstrates strong spatial understanding from video context: it plans navigation and grounds target objects end-to-end, follows camera-pose instructions while maintaining spatial consistency, and generalizes to long contexts and out-of-domain environments. Taken together, these results advance video generative models toward general visuospatial reasoning.

cs.CV

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language models (MLLMs) remains largely untapped, hindered by the absence of large-scale datasets that capture the rich, spatially grounded reasoning intrinsic to visual understanding. Existing visual-CoT resources are typically small, domain-specific, or lack the human-like stepwise structure necessary for compositional visual reasoning. In this paper, we introduce VisReason, a large-scale dataset designed to advance visual Chain-of-Thought reasoning. VisReason comprises 489K annotated examples spanning four diverse domains, each featuring multi-round, human-like rationales that guide MLLMs through interpretable visual reasoning steps. Building upon this, we curate VisReason-Pro, a 165K subset produced with a stronger expert-level GPT annotator, enriched with detailed reasoning traces and 3D spatial grounding via depth-informed annotations. Fine-tuning the state-of-the-art Qwen2.5-VL model on VisReason and VisReason-Pro yields substantial improvements in step-by-step visual reasoning accuracy, interpretability, and cross-benchmark generalization. These results demonstrate that VisReason equips MLLMs with more systematic and generalizable reasoning capabilities. We envision VisReason as a cornerstone for cultivating human-like visual reasoning, paving the way toward the next generation of multimodal intelligence.

cs.CV

Chain-of-Generation: Progressive Latent Diffusion for Text-Guided Molecular Design

Text-conditioned molecular generation aims to translate natural-language descriptions into chemical structures, enabling scientists to specify functional groups, scaffolds, and physicochemical constraints without handcrafted rules. Diffusion-based models, particularly latent diffusion models (LDMs), have recently shown promise by performing stochastic search in a continuous latent space that compactly captures molecular semantics. Yet existing methods rely on one-shot conditioning, where the entire prompt is encoded once and applied throughout diffusion, making it hard to satisfy all the requirements in the prompt. We discuss three outstanding challenges of one-shot conditioning generation, including the poor interpretability of the generated components, the failure to generate all substructures, and the overambition in considering all requirements simultaneously. We then propose three principles to address those challenges, motivated by which we propose Chain-of-Generation (CoG), a training-free multi-stage latent diffusion framework. CoG decomposes each prompt into curriculum-ordered semantic segments and progressively incorporates them as intermediate goals, guiding the denoising trajectory toward molecules that satisfy increasingly rich linguistic constraints. To reinforce semantic guidance, we further introduce a post-alignment learning phase that strengthens the correspondence between textual and molecular latent spaces. Extensive experiments on benchmark and real-world tasks demonstrate that CoG yields higher semantic alignment, diversity, and controllability than one-shot baselines, producing molecules that more faithfully reflect complex, compositional prompts while offering transparent insight into the generation process.

cs.LG

VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how to effectively bridge vision-language (VL) representations to action (A). We introduce VLA-Adapter, a novel paradigm designed to reduce the reliance of VLA models on large-scale VLMs and extensive pre-training. To this end, we first systematically analyze the effectiveness of various VL conditions and present key findings on which conditions are essential for bridging perception and action spaces. Based on these insights, we propose a lightweight Policy module with Bridge Attention, which autonomously injects the optimal condition into the action space. In this way, our method achieves high performance using only a 0.5B-parameter backbone, without any robotic data pre-training. Extensive experiments on both simulated and real-world robotic benchmarks demonstrate that VLA-Adapter not only achieves state-of-the-art level performance, but also offers the fast inference speed reported to date. Furthermore, thanks to the proposed advanced bridging paradigm, VLA-Adapter enables the training of a powerful VLA model in just 8 hours on a single consumer-grade GPU, greatly lowering the barrier to deploying the VLA model. Project page: https://vla-adapter.github.io/.

cs.RO

Correctness-Guaranteed Code Generation via Constrained Decoding

Language Models (LMs) are increasingly being used for code generation, but ensuring the correctness of generated programs remains a significant challenge. Although imperfect code may be acceptable during software development with human oversight, domains such as video games and robotics require one-shot correctness for runtime-critical components. We present a constrained decoding algorithm for generating semantically correct programs that incorporates a context-sensitive parser, which, at each step, outputs a regular expression that satisfies a critical non-extensible property to guide the generation of the next token sequence that can continue to a correct program. To build such a context-sensitive parser, we propose a framework of a dynamic tree of parsers (ToP) during parsing, where each parser corresponds to a modular context-free grammar enriched with contextual information such as variable scopes and type constraints, with tree branches representing ambiguity in the future code segment. We demonstrate our approach through sLua, a strongly typed variant of Lua, showing that our method can generate semantically correct programs conforming to any prescribed scripting API. We further show that, with careful design, our semantic guarantees extend to runtime correctness, as validated in the application of generating game mechanics for a roguelike video game.

cs.PL

QSEA: Quantum Self-supervised Learning with Entanglement Augmentation

As an unsupervised feature representation paradigm, Self-Supervised Learning (SSL) uses the intrinsic structure of data to extract meaningful features without relying on manual annotation. Despite the success of SSL, there are still problems, such as limited model capacity or insufficient representation ability. Quantum SSL has become a promising alternative because it can exploit quantum states to enhance expression ability and learning efficiency. This letter proposes a Quantum SSL with entanglement augmentation method (QSEA). Different from existing Quantum SSLs, QSEA introduces an entanglement-based sample generation scheme and a fidelity-driven quantum loss function. Specifically, QSEA constructs augmented samples by entangling an auxiliary qubit with the raw state and applying parameterized unitary transformations. The loss function is defined using quantum fidelity, quantifying similarity between quantum representations and effectively capturing sample relations. Experimental results show that QSEA outperforms existing quantum self-supervised methods on multiple benchmarks and shows stronger stability in decorrelation noise environments. This framework lays the theoretical and practical foundation for quantum learning systems and advances the development of quantum machine learning in SSL.

quant-ph

Evidence of Mott Insulator with Thermally Induced Melting Behavior in Kagome Compound Nb3Cl8

The kagome lattice provides a playground to explore novel correlated quantum states due to the presence of flat bands in its electronic structure. Recently discovered layered kagome compound Nb3Cl8 has been proposed as a Mott insulator coming from the half-filled flat band. Here we have carried out systematic transport study to uncover the evidence of Mott insulator in Nb3Cl8 thin flakes. Bipolar semiconducting property with Fermi level close to conduction band has been revealed. We have further probed the chemical potential of Nb3Cl8 by tracing the charge neutrality point of the monolayer graphene proximate to Nb3Cl8. The gap of Nb3Cl8 flakes is approximately 1.10 eV at 100 K and shows pronounced temperature dependence, decreasing substantially with increasing temperature to ~0.63 eV at 300 K. The melting behavior of the gapped state is in consistent with theoretically proposed Mott insulator in Nb3Cl8. Our work has demonstrated Nb3Cl8 as a promising platform to study strongly correlated physics at relatively high temperature.

cond-mat.str-el

Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning

In natural language processing (NLP) and computer vision (CV), the successful application of foundation models across diverse tasks has demonstrated their remarkable potential. However, despite the rich structural and textual information embedded in knowledge graphs (KGs), existing research of foundation model for KG has primarily focused on their structural aspects, with most efforts restricted to in-KG tasks (e.g., knowledge graph completion, KGC). This limitation has hindered progress in addressing more challenging out-of-KG tasks. In this paper, we introduce MERRY, a foundation model for general knowledge graph reasoning, and investigate its performance across two task categories: in-KG reasoning tasks (e.g., KGC) and out-of-KG tasks (e.g., KG question answering, KGQA). We not only utilize the structural information, but also the textual information in KGs. Specifically, we propose a multi-perspective Conditional Message Passing (CMP) encoding architecture to bridge the gap between textual and structural modalities, enabling their seamless integration. Additionally, we introduce a dynamic residual fusion module to selectively retain relevant textual information and a flexible edge scoring mechanism to adapt to diverse downstream tasks. Comprehensive evaluations on 28 datasets demonstrate that MERRY outperforms existing baselines in most scenarios, showcasing strong reasoning capabilities within KGs and excellent generalization to out-of-KG tasks such as KGQA.

cs.CL