arXiv ScienceSearch

arXiv subjects

Haoxiang Li

Publications and source records attributed to Haoxiang Li.

At least 19 recordsLinked to original sources

Entwined lattice of atoms and anionic electrons in layered electride LaCl

Controlling the lattice geometry that governs electronic structure is a central theme in condensed-matter physics, yet in crystalline solids this geometry is usually fixed by the atomic framework. Electrides offer an alternative route to electronic structure design in which their excess electrons can organize into anionic electron lattice (AEL) and provide a lattice-like degree of freedom. Recent work has highlighted the standalone limit, where the AEL in YCl yields bands well described by the dice-lattice model. Here, using angle-resolved photoemission spectroscopy (ARPES), we show that LaCl, although isostructural to YCl, realizes a qualitatively different regime where the AEL is entwined with the La cation framework, producing a fully reconstructed electronic structure. Combining the ARPES result with tight-binding model analysis, we demonstrate that this radical divergence stems from the activation of direct hopping channels between the AEL and the La atomic lattice. This coupling reshapes the effective lattice geometry, reconstructs the electronic states, and modifies the associated Chern band topology, transforming the bipartite dice-lattice network in YCl into a tripartite structure in LaCl. Our findings demonstrate that the coupling between the AEL and the atomic lattice can actively shape the effective lattice geometry that governs the electronic structure. This coupling can act as a powerful tuning knob for electronic structure design that is inaccessible in conventional materials.

cond-mat.str-el

Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing

Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additional pretraining. Targeting the dense distribution, multi-scale nature, and large size of remote sensing imagery, we design two core modules. Its Text-aware Laplacian Propagation module de-noises patch-level predictions by combining textual semantic affinities with local visual similarity, improving regional consistency while preserving boundaries. Its Gaussian Splatting Upsampling module reconstructs pixel-level features through RGB-guided anisotropic aggregation and test-time optimization. A global-anchor sliding-window strategy further supports large-scale imagery. Experiments on UDD5, DOTA, and LoveDA demonstrate competitive or superior performance over existing training-free methods, effectively filling the gap of DINO-series models in training-free open-vocabulary segmentation and providing a viable new path for further advances in this direction.

cs.CV

PE-Field 4D: Video Generation Models as Canvas

Diffusion Transformers have recently achieved strong performance in video generation, yet controlling scene geometry under viewpoint changes and camera motion remains challenging. In this work, we revisit the role of positional encoding in video diffusion transformers and show that it provides a useful spatial bias for geometry-aware control. Specifically, if reference tokens are encoded according to their projected locations in the target view, the denoising model is encouraged to retrieve content from position aligned regions of the input video. Building on this observation, we introduce a geometry-aware cross-attention mechanism that enables target video latent tokens to attend to structured context tokens derived from reference images or frames. To establish correspondence between the reference content and the target camera trajectory, we equip the context tokens with a projected positional encoding scheme that combines target-view 2D reprojection with depth-aware disambiguation. At the same time, we preserve the original spatiotemporal positional encoding of the generated video latent, allowing geometric guidance to be injected while maintaining consistency with the video model's native latent structure. The resulting framework provides a simple and effective approach for controllable video generation. It improves spatial controllability in viewpoint-dependent editing tasks, including camera re-trajectory, novel-view video synthesis, and geometry-aware video editing, while preserving the generative prior of the underlying video diffusion model. The code is available at: https://github.com/MTLab/PE-Field.

cs.CV

Natural Language Camera Movement Understanding

Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing vision-language models (VLMs) fail this task in surprising ways, frequently confusing translation with rotation, left with right, and object movement with camera movement. To address these limitations, we establish natural language camera movement understanding as a standalone research task. We introduce a two-level cinematographic taxonomy and an extensive, atomic benchmark featuring both real and synthetic videos. Furthermore, we curate a large-scale, multi-source training set enhanced by targeted camera movement augmentation. Our fine-tuned VLM-8B outperforms Gemini 3.1 Pro by 10% and 11% on our benchmark's real and synthetic videos, respectively. Despite these gains, a significant gap remains relative to human performance, underscoring the need to promote and facilitate future research on natural language camera movement understanding.

cs.CV

The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation

This paper presents an overview of the inaugural PortraitCraft Challenge, held as one of the official competitions at CVPR 2026. The challenge focuses on portrait composition understanding and generation, aiming to advance AI research in portrait aesthetics analysis and controllable image synthesis. Unlike existing datasets and tasks that primarily focus on global aesthetic scoring, PortraitCraft introduces a unified evaluation framework comprising two complementary tracks. Track 1 requires models to perform structured portrait composition understanding, and Track 2 requires models to generate portrait images from structured composition descriptions under explicit compositional constraints. To support the challenge, we constructed and publicly released a large-scale portrait composition dataset consisting of approximately 50,000 curated real portrait images, providing multi-level supervision. This report describes the challenge setup, evaluation protocols, dataset composition, and final results, along with an analysis of the technical characteristics of the submitted solutions. The PortraitCraft Challenge provides a standardized and reproducible platform for research on portrait composition understanding and generation, and is expected to foster further progress in the fields of portrait aesthetics and controllable image generation.

cs.CV

INFANiTE: Implicit Neural representation for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI

Spatio-temporal fetal brain atlases are important for characterizing normative neurodevelopment and identifying congenital anomalies. However, existing atlas construction pipelines necessitate days for slice-to-volume reconstruction (SVR) to generate high-resolution 3D brain volumes and several additional days for iterative volume registration, thereby rendering atlas construction from large-scale cohorts prohibitively impractical. We address these limitations with INFANiTE, an Implicit Neural Representation (INR) framework for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI scans, bypassing both the costly SVR and the iterative non-rigid registration steps entirely, thereby substantially accelerating atlas construction. Extensive experiments demonstrate that INFANiTE outperforms existing baselines in subject consistency, reference fidelity, intrinsic quality and biological plausibility, even under challenging sparse-data settings. Additionally, INFANiTE reduces the end-to-end processing time (i.e., from raw scans to the final atlas) from days to hours compared to the traditional 3D volume-based pipeline (e.g., SyGN), facilitating large-scale population-level fetal brain analysis. Code: https://github.com/hu2274898/INFANiTE

cs.CV

Annotation-free deep learning for detection and segmentation of fetal germinal matrix-intraventricular hemorrhage in brain MRI

Prenatal germinal matrix-intraventricular hemorrhage (GMH-IVH) is a leading cause of infant mortality and neurodevelopmental impairment, yet its manual diagnosis and lesion segmentation on fetal brain MRI are labor-intensive and error-prone. Although supervised deep learning offers potential for automation, it typically requires large amounts of annotated GMH-IVH data, which are challenging to obtain for such a rare condition (0.5-0.9 per 1000 pregnancies). To address these problems, an annotation-free deep learning framework, FreeHemoSeg, was developed for automated detection and segmentation of GMH-IVH without any real patient annotations. Instead of learning from expert labels, FreeHemoSeg was trained on pseudo GMH-IVH images synthesized from normal fetal data guided by medical priors. The framework was evaluated in a retrospective multicentre study of 1,674 stacks of 2D T2-weighted MRI from 558 pregnant women, using data from one hospital for internal training and validation and two hospitals for external validation. FreeHemoSeg achieved the highest diagnostic and segmentation performance in both internal validation (AUROC: 0.959; AUPR: 0.928; sensitivity: 0.914; specificity: 0.966; DSC: 0.559) and external validation (AUROC: 0.930; AUPR: 0.884; sensitivity: 0.824; specificity: 0.943; DSC: 0.512), outperforming a supervised model trained on limited empirical data and unsupervised anomaly detection methods. Moreover, FreeHemoSeg assistance improved radiologists' sensitivity (from 0.882 to 0.941-1.000) and diagnostic confidence, while reducing interpretation time by 16.0-52.7%. We anticipate its immediate utility in supporting earlier diagnosis, prognostic counselling, and perinatal planning for fetal GMH-IVH. Code: https://github.com/Arktis2022/FreeHemoSeg.

eess.IV

APIOT: Autonomous Vulnerability Management Across Bare-Metal Industrial OT Networks

Bare-metal operational technology (OT) devices -- especially the microcontrollers running Modbus/TCP and CoAP at the base of industrial control systems -- have remained outside the reach of autonomous security attacks. Prior autonomous pentesting studies target Linux and web systems, whose shells and filesystems are familiar to LLM agents. Bare-metal OT has neither, so agents must reason directly over protocol fields and parser semantics. This requires new action-space designs and runtime controls, and opens new research questions about protocol-level exploit reasoning and its deployment envelope. We present APIOT (Autonomous Purple-teaming for Industrial OT), the first large language model (LLM) framework demonstrating an autonomous attack and remediation of bare-metal OT devices, achieving the full discovery -> exploitation -> patching -> verification cycle without step-by-step human intervention. We implemented and evaluated this framework on Zephyr RTOS firmware across heterogeneous industrial IoT (IIoT) topologies. Through 290 experiment runs spanning five frontier LLMs, three network topologies, two impairment levels, and guided versus unguided conditions, APIOT achieved a mission success rate of 90.0% on the full attack-remediation cycle. We found that the runtime governance layer (which we call an overseer) is a critical engineering variable: without it, agents exhibit systematic degenerate patterns, including repetition loops, missing crash verification, and reconnaissance deadlocks. Together, these findings carry two implications beyond our testbed. Attacker expertise is no longer the binding constraint on bare-metal OT exploitation, and defender threat models must now assume LLM-augmented adversaries capable of executing autonomous discovery-through-remediation cycles against industrial firmware.

cs.CR

Nonmonotonic Evolution of the Superconducting Transition Temperature and Robust Multigap Extended s-wave + s-wave Pairing in Zn-Substituted FeSe Single Crystals

We report a systematic study of superconductivity on Fe1-xZnxSe single crystals synthesized over a broad Zn doping range (x = 0-0.023). High-quality single crystals across all compositions range exhibit superconducting transitions, while the transition temperature Tc shows a pronounced nonmonotonic dependence on Zn doping concentration, indicating that the underlying mechanism govering Tc its evolution cannot be explained solely by simple impurity pair breaking alone. Magnetization and transport measurements confirm the bulk behavior of superconductivity and reveal enhanced scattering effects with Zn doping. Low-temperature specific heat is consistently described by a two-gap scenario composed of an isotropic s-wave gap and an anisotropic extended s-wave gap, whereas single-gap and alternative pairing symmetries fail to describe the data. The nearly unchanged relative weights of the two gap components suggest the weak interband scattering induced by Zn substitution, thereby preserving multiband superconductivity. These results demonstrate the robustness of multigap superconductivity in FeSe and impose stringent constraints on candidate pairing mechanisms, highlighting the role of multiband electronic structure and anisotropic gap formation.

cond-mat.supr-con

PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation

Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coarse aesthetic scoring, generic image aesthetics, or unconstrained portrait generation. This limits systematic research on structured portrait composition analysis and controllable portrait generation under explicit composition requirements. In this paper, we introduce PortraitCraft, a unified benchmark for portrait composition understanding and generation. PortraitCraft is built on a dataset of approximately 50,000 curated real portrait images with structured multi-level supervision, including global composition scores, annotations over 13 composition attributes, attribute-level explanation texts, visual question answering pairs, and composition-oriented textual descriptions for generation. Based on this dataset, we establish two complementary benchmark tasks for composition understanding and composition-aware generation within a unified framework. The first evaluates portrait composition understanding through score prediction, fine-grained attribute reasoning, and image-grounded visual question answering, while the second evaluates portrait generation from structured composition descriptions under explicit composition constraints. We further define standardized evaluation protocols and provide reference baseline results with representative multimodal models. PortraitCraft provides a comprehensive benchmark for future research on fine-grained portrait understanding, interpretable aesthetic assessment, and controllable portrait generation.

cs.CV

Hi-Light: A Path to high-fidelity, high-resolution video relighting with a Novel Evaluation Paradigm

Video relighting offers immense creative potential and commercial value but is hindered by challenges, including the absence of an adequate evaluation metric, severe light flickering, and the degradation of fine-grained details during editing. To overcome these challenges, we introduce Hi-Light, a novel, training-free framework for high-fidelity, high-resolution, robust video relighting. Our approach introduces three technical innovations: lightness prior anchored guided relighting diffusion that stabilises intermediate relit video, a Hybrid Motion-Adaptive Lighting Smoothing Filter that leverages optical flow to ensure temporal stability without introducing motion blur, and a LAB-based Detail Fusion module that preserves high-frequency detail information from the original video. Furthermore, to address the critical gap in evaluation, we propose the Light Stability Score, the first quantitative metric designed to specifically measure lighting consistency. Extensive experiments demonstrate that Hi-Light significantly outperforms state-of-the-art methods in both qualitative and quantitative comparisons, producing stable, highly detailed relit videos.

cs.CV

UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation

Content-aware layout generation is a critical task in graphic design automation, focused on creating visually appealing arrangements of elements that seamlessly blend with a given background image. The variety of real-world applications makes it highly challenging to develop a single model capable of unifying the diverse range of input-constrained generation sub-tasks, such as those conditioned by element types, sizes, or their relationships. Current methods either address only a subset of these tasks or necessitate separate model parameters for different conditions, failing to offer a truly unified solution. In this paper, we propose UniLayDiff: a Unified Diffusion Transformer, that for the first time, addresses various content-aware layout generation tasks with a single, end-to-end trainable model. Specifically, we treat layout constraints as a distinct modality and employ Multi-Modal Diffusion Transformer framework to capture the complex interplay between the background image, layout elements, and diverse constraints. Moreover, we integrate relation constraints through fine-tuning the model with LoRA after pretraining the model on other tasks. Such a schema not only achieves unified conditional generation but also enhances overall layout quality. Extensive experiments demonstrate that UniLayDiff achieves state-of-the-art performance across from unconditional to various conditional generation tasks and, to the best of our knowledge, is the first model to unify the full range of content-aware layout generation tasks.

cs.CV

Positional Encoding Field

Diffusion Transformers (DiTs) have emerged as the dominant architecture for visual generation, powering state-of-the-art image and video models. By representing images as patch tokens with positional encodings (PEs), DiTs combine Transformer scalability with spatial and temporal inductive biases. In this work, we revisit how DiTs organize visual content and discover that patch tokens exhibit a surprising degree of independence: even when PEs are perturbed, DiTs still produce globally coherent outputs, indicating that spatial coherence is primarily governed by PEs. Motivated by this finding, we introduce the Positional Encoding Field (PE-Field), which extends positional encodings from the 2D plane to a structured 3D field. PE-Field incorporates depth-aware encodings for volumetric reasoning and hierarchical encodings for fine-grained sub-patch control, enabling DiTs to model geometry directly in 3D space. Our PE-Field-augmented DiT achieves state-of-the-art performance on single-image novel view synthesis and generalizes to controllable spatial image editing.

cs.CV

ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering

Hair simulation and rendering are challenging due to complex strand dynamics, diverse material properties, and intricate light-hair interactions. Recent video diffusion models can generate high-quality videos, but they lack fine-grained control over hair dynamics. We present ControlHair, a hybrid framework that integrates a physics simulator with conditional video diffusion to enable precise and controllable dynamic hair rendering. ControlHair adopts a three-stage pipeline: it first encodes physics conditions into per-frame geometry using a simulator, then extracts per-frame control signals, and finally feeds control signals into a video diffusion model to generate videos with desired hair dynamics. This cascaded design decouples physics reasoning from video generation, supports diverse physics, and makes training the video diffusion model easy. Trained on a curated 10K video dataset, ControlHair outperforms text- and pose-conditioned baselines, delivering precisely controlled hair dynamics. We also demonstrate three use cases of ControlHair, including dynamic hairstyle try-on, bullet-time effects, and cinemagraphic. Project page: https://linwk20.github.io/controlhair-web.

cs.GR

GeoRemover: Removing Objects and Their Causal Visual Artifacts

Towards intelligent image editing, object removal should eliminate both the target object and its causal visual artifacts, such as shadows and reflections. However, existing image appearance-based methods either follow strictly mask-aligned training and fail to remove these causal effects which are not explicitly masked, or adopt loosely mask-aligned strategies that lack controllability and may unintentionally over-erase other objects. We identify that these limitations stem from ignoring the causal relationship between an object's geometry presence and its visual effects. To address this limitation, we propose a geometry-aware two-stage framework that decouples object removal into (1) geometry removal and (2) appearance rendering. In the first stage, we remove the object directly from the geometry (e.g., depth) using strictly mask-aligned supervision, enabling structure-aware editing with strong geometric constraints. In the second stage, we render a photorealistic RGB image conditioned on the updated geometry, where causal visual effects are considered implicitly as a result of the modified 3D geometry. To guide learning in the geometry removal stage, we introduce a preference-driven objective based on positive and negative sample pairs, encouraging the model to remove objects as well as their causal visual artifacts while avoiding new structural insertions. Extensive experiments demonstrate that our method achieves state-of-the-art performance in removing both objects and their associated artifacts on two popular benchmarks. The code is available at https://github.com/buxiangzhiren/GeoRemover.

cs.CV

YCl Electride as a Multi-Orbital Correlated Topological Dice Lattice System

The long-sought dice lattice flat band has recently been discovered for the first time in two-dimensional layered electride yttrium monochloride (YCl) [Nature Communications 17, 2213 (2026)]. While essential flat band features of YCl were captured by an idealized simple dice lattice model, we reveal in this Letter that a unique layer-orbital-valley coupling in YCl puts up a fundamental obstruction against a simple three-band dice lattice description of the flat band, and necessitates a multi-orbital description that faithfully represents the symmetry, topology, and correlation physics in the first-ever dice metal. Using an ab initio based multi-orbital Hubbard model with local interactions, we predict that the multi-orbital flat band supports a robust ferromagnetic ground state and electrically tunable correlated quantum anomalous Hall phases that are absent in an interacting single-orbital dice lattice. Our findings open a new avenue for exploring correlation and topology in electride systems.

cond-mat.mtrl-sci

Experimental realization of dice-lattice flat band at the Fermi level in layered electride YCl

Flat electronic bands, where interactions among electrons overwhelm their kinetic energies, hold the promise for exotic correlation physics. The dice lattice has long been theorized as a host of flat bands with intriguing band topology. However, to date, no material has ever been found to host the characteristic flat bands of a dice lattice. Here, using angle-resolved photoemission spectroscopy (ARPES), we discover a dice-lattice flat band at $E_F$ in the van der Waals (vdW) electride [YCl]$^{2+}$: 2e-. In this system, excess valence electrons from Y deconfine from the cation framework to form an interstitial anionic electron lattice that constitutes the dice lattice. Our ARPES measurements unambiguously identify two sets of dice-lattice bands in YCl, including a nearly dispersionless band at the Fermi level. The flat bands and other dispersive bands observed in ARPES find excellent agreement with first-principles calculations, and theoretical analysis reveals that the near-$E_F$ electronic structure is well captured by a simple dice-lattice model. Our findings thus end the long quest of a real dice flat band material and establish vdW electride YCl as a prototype of dice metals. Our results further demonstrate the anionic electron lattice as a novel scheme for realizing lattice geometries and electronic structures rare to find in conventional crystalline systems.

cond-mat.str-el

LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation

Controllable layout generation aims to create plausible visual arrangements of element bounding boxes within a graphic design according to certain optional constraints, such as the type or position of a specific component. While recent diffusion or flow-matching models have achieved considerable advances in multifarious conditional generation tasks, there remains considerable room for generating optimal arrangements under given conditions. In this work, we propose to carry out layout generation through retrieving by conditions and reference-guided generation. Specifically, we retrieve appropriate layout templates according to given conditions as references. The references are then utilized to guide the denoising or flow-based transport process. By retrieving layouts compatible with the given conditions, we can uncover the potential information not explicitly provided in the given condition. Such an approach offers more effective guidance to the model during the generation process, in contrast to previous models that feed the condition to the model and let the model infer the unprovided layout attributes directly. Meanwhile, we design a condition-modulated attention that selectively absorbs retrieval knowledge, adapting to the difference between retrieved templates and given conditions. Extensive experiment results show that our method successfully produces high-quality layouts that meet the given conditions and outperforms existing state-of-the-art models. Code will be released upon acceptance.

cs.CV