arXiv ScienceSearch

arXiv subjects

Si Li

Publications and source records attributed to Si Li.

At least 19 recordsLinked to original sources

On the Formality of Configuration Spaces of $\mathbb{R}^{n'} \times \mathbb{C}^{n}$

This paper presents a complete classification of the formality of configuration spaces of $\mathbb{R}^{n'} \times \mathbb{C}^{n}$. We define a constructible de Rham-Dolbeault cohomology theory which provides a constructible CDGA (commutative differential graded algebra) model of $\Conf_m(\mathbb{R}^{n'} \times \mathbb{C}^{n})$. For $(n'=0,n\ge2)$ or $(n'=1,n\ge1)$, the CDGAs are non-formal. For $n'\ge2,n\ge1$, we establish an explicit quasi-isomorphism between the constructible CDGA and its cohomology by using a diagrammatic CDGA of admissible diagrams and a regularized configuration space integral, which leads to the formality. As an application, we show that the local operator algebra of a topological-holomorphic field theory on $\mathbb{R}^{n'} \times \mathbb{C}^{n}$ ($n'\ge2,n\ge1$) is homotopically equivalent to a higher dimensional analog of vertex algebras.

math.AT

Mirror Chern insulators in two-dimensional altermagnetic Tc$_2$Cl$_2$O and Tc$_2$Br$_2$O

The interplay between altermagnetism and crystalline band topology provides an intriguing avenue for realizing unconventional topological phases with distinctive spin-dependent properties. Here, based on first-principles calculations and theoretical analysis, we identify monolayer $\mathrm{Tc}_2X_2\mathrm{O}$ ($X$ = Cl, Br) as a family of two-dimensional altermagnetic mirror Chern insulators. In the absence of spin--orbit coupling (SOC), both monolayers exhibit robust altermagnetism with mirror-spin coupling and host two symmetry-protected Weyl points in each spin channel near the Fermi level. The Weyl points in opposite spin channels carry distinct mirror-symmetry eigenvalues, $m_z=\pm i$. Upon inclusion of SOC, the Weyl points are gapped, and the two mirror sectors acquire opposite Chern numbers, ${\cal {C}}_{+}=1$ and ${\cal {C}}_{-}=-1$, resulting in a nonzero mirror Chern number ${\cal {C}}_m=1$. A low-energy $k\cdot p$ model captures the symmetry protection of the Weyl points and elucidates their SOC-induced mass gaps and topological character. Furthermore, the resulting mirror Chern insulating phases host helical edge states within the bulk band gap and exhibit a quantized spin Hall conductivity. Our work establishes a direct connection between altermagnetism and mirror Chern topology and provides a promising platform for exploring unconventional topological and spin-dependent phenomena in two-dimensional altermagnetic materials.

cond-mat.mtrl-sci

Valley- and Spin-Dependent Electronic and Transport Properties of Two-Dimensional Altermagnetic Titanium-Based Chalcogenide Halides

Altermagnets (AMs) combine fully compensated magnetization with momentum-dependent spin splitting, yet intrinsic altermagnetic materials exhibiting exceptional valley characteristics remain scarce. Here, we identify monolayer titanium-based chalcogenide halides, Ti$_2X_2Y$ ($X$ = F, Cl, Br, I; $Y$ = O, S, Se, Te), as a new family of two-dimensional altermagnetic valley materials. These monolayers exhibit robust $d$-wave altermagnetic order, semiconducting band gaps, and pronounced spin-polarized valley characteristics. We show that uniaxial strain breaks the valley degeneracy, inducing giant valley polarization together with a tunable piezomagnetic response. An in-plane electric field generates noncollinear spin currents, while spin--orbit coupling gives rise to the anomalous Hall effect, valley-selective linear dichroism, and the magneto-optical Kerr effect. These findings establish Ti$_2X_2Y$ monolayers as a versatile platform for exploring spin- and valley-dependent electronic, optical, and transport phenomena in two-dimensional altermagnets.

cond-mat.mtrl-sci

FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams

Music-driven dance video synthesis aims to animate a reference person according to a given music clip. The task is challenging because it requires a model to jointly learn music-to-motion correspondence, identity-preserving human animation, temporal coherence, and visually realistic video generation. We present FlowDance, a music-driven dance video generation framework that integrates explicit motion modeling with reference-preserving visual synthesis through parallel pose and RGB streams. We further introduce timestep-aware pose injection to adapt structural guidance across denoising steps and persistent identity injection to preserve the reference appearance over long video. To support this task, we further build a popularity-curated, high-resolution in-the-wild dance video dataset with synchronized music, RGB videos, 3D body motion, camera parameters, and projected 2D pose annotations. Extensive experiments show that FlowDance achieves strong performance in both dance motion generation and music-driven dance video synthesis.

cs.CV

Heat Kernel and Resurgence

We study the resurgent structure of short-time heat kernel asymptotics from the viewpoint of Picard-Lefschetz theory. For a real analytic Riemannian manifold, we show the heat kernel admits a 1-Gevrey small-time expansion whose Borel transform detects complex-geometric data beyond the real geodesic sector. We formulate an infinite-dimensional Picard-Lefschetz problem of Morse-Floer type for the holomorphic energy functional on the complexified path space, and propose a heat-kernel analogue of the Picard-Lefschetz/Alien correspondence. In this framework, pointed alien operators acting on the asymptotic expansion associated with the real geodesic are predicted to produce the formal heat-kernel sectors associated with other holomorphic geodesics, with coefficients given by signed counts of connecting trajectories of the Morse flow. We perform a confirming test of this proposal on the hyperbolic plane $H^2$.

math-ph

Deep Learning-Based Automated Quantification of TIMI Myocardial Perfusion Frame Count (DL-TMPFC) from Coronary Angiography: A Novel Framework for Rapid Assessment of Microvascular Dysfunction

Aims: Coronary microvascular dysfunction (CMVD) affects approximately 40%-60% of patients with ischemia and non-obstructive coronary arteries, yet diagnosis remains challenging due to reliance on invasive functional testing or subjective Thrombolysis In Myocardial Infarction (TIMI) flow grade. The TIMI Myocardial Perfusion Frame Count (TMPFC) offers an objective, angiography-based quantitative measure of CMVD, but its clinical translation is hindered by cumbersome manual calculation and insufficient validation. This study aims to develop and validate a deep learning-powered TMPFC calculation (DL-TMPFC), enabling integration into clinical workflows. Methods and results: DL-TMPFC framework comprised two components. A stenosis detection network first excluded obstructive coronary artery disease (CAD). A territory-aware segmentation network then identified perfusion territories and TMPFC calculation module automatically determined the first and last frames from angiographic sequences. The framework was validated in a cohort of 655 patients (445 of obstructive CAD, 100 of confirmed CMVD, 110 of control group) from three independent institutions. DL-TMPFC showed excellent agreement with expert manual measurements (bias: -0.93 frames; 95% LoA: -5.33 to +3.47; r =0.98). DL-TMPFC markedly enhanced clinical feasibility by fully automating TMPFC and removing observer dependence. Clinically, DL-TMPFC accurately identified CMVD across a full spectrum of coronary pathologies and captured the continuous severity of CMVD beyond binary classification, enabling quantitative risk stratification. Conclusion: DL-TMPFC enabled automatic, standardized, and accurate quantification of CMVD directly from routine angiography. By providing an automatic and objective measure, this tool provided immediate diagnostic information for timely recognition and management of CMVD in clinical practice.

cs.CV

Picard-Lefschetz theory and alien calculus: a case study

We compare Picard--Lefschetz theory and resurgence in three basic one-dimensional exponential integrals: the Airy model, the Bessel model, and the Gamma model. On the Picard--Lefschetz side, we describe the Lefschetz thimbles and compute the connecting trajectories between critical points appearing at Stokes phases. On the resurgent side, we analyze the Borel singularities of the saddle expansions and use alien operators to recover the same Stokes coefficients. These examples serve as explicit finite-dimensional test cases for the dictionary between thimble wall-crossing and alien calculus.

math-ph

CDSA-Net:Collaborative Decoupling of Vascular Structure and Background for High-Fidelity Coronary Digital Subtraction Angiography

Digital subtraction angiography (DSA) in coronary imaging is fundamentally challenged by physiological motion, forcing reliance on raw angiograms cluttered with anatomical noise. Existing deep learning methods often produced images with two critical clinically unacceptable flaws: persistent boundary artifacts and a loss of native tissue grayscale fidelity that undermined diagnostic confidence. We propose a novel framework termed as CDSA-Net that for the first time explicitly decouples and jointly optimizes vascular structure preservation and realistic background restoration. CDSA-Net introduces two core innovations: (i) A hierarchical geometric prior guidance (HGPG) mechanism, embedded in our coronary structure extraction network (CSENet). It synergistically combines integrated geometric prior (IGP) with gated spatial modulation (GSM) and centerline-aware topology (CAT) loss supervision, ensuring structural continuity. (ii) An adaptive noise module (ANM) within our coronary background restoration network (CBResNet). Unlike standard restoration, ANM uniquely models the stochastic nature of clinical X-ray noise, bridging the domain gap to enable seamless background intensity estimation and the complete elimination of boundary artifacts. The final subtraction is obtained by removing the restored background from the raw angiogram. Quantitatively, it significantly outperformed state-of-the-art methods in vascular intensity correlation and perceptual quality. A 25.6% improvement in morphology assessment efficiency and a 42.9% gain in hemodynamic evaluation speed set a new benchmark for utility in interventional cardiology, while maintaining diagnostic results consistent with raw angiograms. The project code is available at https://github.com/DrThink-ai/CDSA-Net.

cs.CV

ReContraster: Making Your Posters Stand Out with Regional Contrast

Effective poster design requires rapidly capturing attention and clearly conveying messages. Inspired by the ``contrast effects'' principle, we propose ReContraster, the first training-free model to leverage regional contrast to make posters stand out. By emulating the cognitive behaviors of a poster designer, ReContraster introduces the compositional multi-agent system to identify elements, organize layout, and evaluate generated poster candidates. To further ensure harmonious transitions across region boundaries, ReContraster integrates the hybrid denoising strategy during the diffusion process. We additionally contribute a new benchmark dataset for comprehensive evaluation. Seven quantitative metrics and four user studies confirm its superiority over relevant state-of-the-art methods, producing visually striking and aesthetically appealing posters.

cs.CV

A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation

The surging demand for adapting long-form cinematic content into short videos has motivated the need for versatile automatic video compilation systems. However, existing compilation methods are limited to predefined tasks, and the community lacks a comprehensive benchmark to evaluate the cinematic compilation. To address this, we introduce CineBench, the first benchmark for instruction-driven cinematic video compilation, featuring diverse user instructions and high-quality ground-truth compilations annotated by professional editors. To overcome contextual collapse and temporal fragmentation, we present CineAgents, a multi-agent system that reformulates cinematic video compilation into ``design-and-compose'' paradigm. CineAgents performs script reverse-engineering to construct a hierarchical narrative memory to provide multi-level context and employs an iterative narrative planning process that refines a creative blueprint into a final compiled script. Extensive experiments demonstrate that CineAgents significantly outperforms existing methods, generating compilations with superior narrative coherence and logical coherence.

cs.CV

Higher-order topological insulators in two-dimensional antiferromagnetic and altermagnetic chromium-based group-IV chalcogenides

Based on first-principles calculations combined with theoretical analysis, we identify a family of monolayer chromium-based group-IV chalcogenides as a new class of two-dimensional (2D) magnetic higher-order topological insulators (HOTIs). Specifically, the CrC$X_3$ ($X=$ S, Se, Te) and CrSiS$_3$ monolayers are found to host conventional antiferromagnetic ground states with $\mathcal{PT}$ symmetry, whereas the Janus compounds Cr$_2$C$_2$S$_3$Se$_3$ and Cr$_2$Si$_2$S$_3$Se$_3$ exhibit altermagnetic ground states. We demonstrate that all these monolayer magnetic materials realize 2D HOTI phases, in which the nontrivial topology is protected by lattice $C_3$ rotational symmetry and manifests as zero-dimensional corner states carrying quantized fractional charges. Moreover, upon inclusion of spin-orbit coupling, these systems remain in the HOTI phase and continue to host robust corner-localized states, confirming the stability of their higher-order topological nature. Our results reveal an intrinsic connection between higher-order topology and magnetic order in 2D antiferromagnetic and altermagnetic systems, identifying chromium-based group-IV chalcogenide monolayers as promising platforms for exploring higher-order topological phases and their potential relevance for future topological and spintronic applications.

cond-mat.mtrl-sci

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled, restricting their applicability in domains like filmmaking and virtual production where explicit scene control is essential. We present LiVER, a diffusion-based framework for scene-controllable video generation. To achieve this, we introduce a novel framework that conditions video synthesis on explicit 3D scene properties, supported by a new large-scale dataset with dense annotations of object layout, lighting, and camera parameters. Our method disentangles these properties by rendering control signals from a unified 3D representation. We propose a lightweight conditioning module and a progressive training strategy to integrate these signals into a foundational video diffusion model, ensuring stable convergence and high fidelity. Our framework enables a wide range of applications, including image-to-video and video-to-video synthesis where the underlying 3D scene is fully editable. To further enhance usability, we develop a scene agent that automatically translates high-level user instructions into the required 3D control signals. Experiments show that LiVER achieves state-of-the-art photorealism and temporal consistency while enabling precise, disentangled control over scene factors, setting a new standard for controllable video generation.

cs.CV

Electric-Field-induced Two-Dimensional Fully Compensated Ferrimagnetism and Emergent Transport Phenomena

The recent discovery of altermagnetism has demonstrated that spin-split electronic band structures can emerge in magnetic systems with zero net magnetization. In contrast, fully compensated ferrimagnetic (fFIM) systems remain far less explored, despite exhibiting similar characteristics such as vanishing magnetization and spin-split bands. Here, based on first-principles calculations combined with theoretical analysis, we demonstrate that monolayer CoS and CoSe can be driven into fFIM states by an external electric field. These materials possess collinear antiferromagnetic ground states with out-of-plane N\'eel vectors, and their electronic bands are spin degenerate due to $\mathcal{PT}$ symmetry. When an out-of-plane electric field is applied, $\mathcal{PT}$ symmetry is broken, inducing fFIM states with pronounced spin splitting. Moreover, we show that the resulting fFIM states host fully spin-polarized currents, anomalous Hall effects, and magneto-optical Kerr and Faraday effects. Our results establish monolayer CoS and CoSe as promising platforms for electric-field-controlled fFIM states and spintronic applications.

cond-mat.mtrl-sci

PolGS++: Physically-Guided Polarimetric Gaussian Splatting for Fast Reflective Surface Reconstruction

Accurate reconstruction of reflective surfaces remains a fundamental challenge in computer vision, with broad applications in real-time virtual reality and digital content creation. Although 3D Gaussian Splatting (3DGS) enables efficient novel-view rendering with explicit representations, its performance on reflective surfaces still lags behind implicit neural methods, especially in recovering fine geometry and surface normals. To address this gap, we propose PolGS++, a physically-guided polarimetric Gaussian Splatting framework for fast reflective surface reconstruction. Specifically, we integrate a polarized BRDF (pBRDF) model into 3DGS to explicitly decouple diffuse and specular components, providing physically grounded reflectance modeling and stronger geometric cues for reflective surface recovery. Furthermore, we introduce a depth-guided visibility mask acquisition mechanism that enables angle-of-polarization (AoP)-based tangent-space consistency constraints in Gaussian Splatting without costly ray-tracing intersections. This physically guided design improves reconstruction quality and efficiency, requiring only about 10 minutes of training. Extensive experiments on both synthetic and real-world datasets validate the effectiveness of our method.

cs.CV

Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors

Social media platforms enable users to express emotions by posting text with accompanying images. In this paper, we propose the Affective Image Filter (AIF) task, which aims to reflect visually-abstract emotions from text into visually-concrete images, thereby creating emotionally compelling results. We first introduce the AIF dataset and the formulation of the AIF models. Then, we present AIF-B as an initial attempt based on a multi-modal transformer architecture. After that, we propose AIF-D as an extension of AIF-B towards deeper emotional reflection, effectively leveraging generative priors from pre-trained large-scale diffusion models. Quantitative and qualitative experiments demonstrate that AIF models achieve superior performance for both content consistency and emotional fidelity compared to state-of-the-art methods. Extensive user study experiments demonstrate that AIF models are significantly more effective at evoking specific emotions. Based on the presented results, we comprehensively discuss the value and potential of AIF models.

cs.CV

STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative

While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approaches have emerged as a promising alternative to computationally intensive end-to-end methods, offering the advantages of fine-grained control and greater efficiency. However, these methods often fail to maintain cross-shot consistency and capture cinematic language. In this paper, we introduce STAGE, a SToryboard-Anchored GEneration workflow to reformulate the keyframe-based multi-shot video generation task. Instead of using sparse keyframes, we propose STEP2 to predict a structural storyboard composed of start-end frame pairs for each shot. We introduce the multi-shot memory pack to ensure long-range entity consistency, the dual-encoding strategy for intra-shot coherence, and the two-stage training scheme to learn cinematic inter-shot transition. We also contribute the large-scale ConStoryBoard dataset, including high-quality movie clips with fine-grained annotations for story progression, cinematic attributes, and human preferences. Extensive experiments demonstrate that STAGE achieves superior performance in structured narrative control and cross-shot coherence. Our code will be available at this url.

cs.CV

Valley-dependent electronic properties in two-dimensional altermagnetic iron-based transition metal chalcogenides

Altermagnets represent a newly identified third class of collinear magnets and have recently emerged as a focal point in condensed matter physics. In this work, through first-principles calculations and theoretical analysis, we identify monolayer Fe$_2$MoX$_4$ (X = S, Se, Te) and Fe$_2$WTe$_4$, a class of iron-based transition metal chalcogenides, as promising altermagnetic materials. These systems are found to be semiconductors exhibiting spin splitting in their nonrelativistic band structures, indicative of intrinsic altermagnetic ordering. Remarkably, their valence bands feature a pair of valleys at the time-reversal-invariant momenta X and Y points. Unlike conventional valley systems, these valleys are related by crystal symmetries rather than time-reversal symmetry. We investigate valley-dependent physical phenomena in these materials, including Berry curvature and optical circular dichroism, revealing strong valley-contrasting behavior. Furthermore, we investigate the effect of uniaxial strain and show that it effectively lifts the valley degeneracy, resulting in pronounced valley polarization. Under hole doping, this strain-induced asymmetry gives rise to a piezomagnetic response. We also explore the generation of anisotropic noncollinear spin currents in these systems, expanding the scope of their spin-related functionalities. Our findings unveil rich valley physics in monolayer Fe$_2$MoX$_4$ (X = S, Se, Te) and Fe$_2$WTe$_4$, highlighting their significant potential for applications in valleytronics, spintronics, and multifunctional nanoelectronic devices.

cond-mat.mtrl-sci

GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning

With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges. However, real-world GUI environments, such as PC software and mobile Apps, are often complex and proprietary, making it difficult to obtain the comprehensive environment information needed for agent training and evaluation. This limitation hinders systematic investigation and benchmarking of agent navigation capabilities. To address this limitation, we introduce GUI Exploration Lab, a simulation environment engine for GUI agent navigation research that enables flexible definition and composition of screens, icons, and navigation graphs, while providing full access to environment information for comprehensive agent training and evaluation. Through extensive experiments, we find that supervised fine-tuning enables effective memorization of fundamental knowledge, serving as a crucial foundation for subsequent training. Building on this, single-turn reinforcement learning further enhances generalization to unseen scenarios. Finally, multi-turn reinforcement learning encourages the development of exploration strategies through interactive trial and error, leading to further improvements in screen navigation performance. We validate our methods on both static and interactive benchmarks, demonstrating that our findings generalize effectively to real-world scenarios. These findings demonstrate the advantages of reinforcement learning approaches in GUI navigation and offer practical guidance for building more capable and generalizable GUI agents.

cs.CV