arXiv ScienceSearch

arXiv subjects

Xinran Yu

Publications and source records attributed to Xinran Yu.

8 recordsLinked to original sources

Spectral asymmetry, supersymmetry and the equivariant Riemann-Roch defect

We investigate the relationship between two interpretations of equivariant Riemann-Roch defects of complex spaces with conic singularities; as (i) equivariant $\eta_{T}$ and $\xi_{T}$ invariants, and as (ii) supertraces over local cohomology groups. This leads to a novel threefold partitioning of the $L^{2}$-spinor space on the link and a corresponding splitting of $\xi_{T}$. Two partitions correspond to cohomological contributions coming from $\bar\partial$-Neumann and $\bar\partial$-Dirichlet operators on the cone, while the third partition makes no contribution to the equivariant index defect, which we show is due to supersymmetric cancellations on the cone that we call lifted supersymmetry. We use this to define complex equivariant $\xi_T$ and $\eta_T$ invariants, which are equivalent to the usual invariants but are easier to compute. We highlight connections to related algebraic and analytic descriptions of Riemann-Roch defects in the literature, both at the level of numbers and their categorifications, and explore connections to existing notions of supersymmetric cancellations in physics and mathematics.

math.DG

SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass

Visual token pruning is a promising approach for reducing the computational cost of vision-language models (VLMs), and existing methods often rely on early pruning decisions to improve efficiency. While effective on coarse-grained reasoning tasks, they suffer from significant performance degradation on tasks requiring fine-grained visual details. Through layer-wise analysis, we reveal substantial discrepancies in visual token importance across layers, showing that tokens deemed unimportant at shallow layers can later become highly relevant for text-conditioned reasoning. To avoid irreversible critical information loss caused by premature pruning, we introduce a new pruning paradigm, termed bypass, which preserves unselected visual tokens and forwards them to subsequent pruning stages for re-evaluation. Building on this paradigm, we propose SwiftVLM, a simple and training-free method that performs pruning at model-specific layers with strong visual token selection capability, while enabling independent pruning decisions across layers. Experiments across multiple VLMs and benchmarks demonstrate that SwiftVLM consistently outperforms existing pruning strategies, achieving superior accuracy-efficiency trade-offs and more faithful visual token selection behavior.

cs.CV

Prompting Lipschitz-constrained network for multiple-in-one sparse-view CT reconstruction

Despite significant advancements in deep learning-based sparse-view computed tomography (SVCT) reconstruction algorithms, these methods still encounter two primary limitations: (i) It is challenging to explicitly prove that the prior networks of deep unfolding algorithms satisfy Lipschitz constraints due to their empirically designed nature. (ii) The substantial storage costs of training a separate model for each setting in the case of multiple views hinder practical clinical applications. To address these issues, we elaborate an explicitly provable Lipschitz-constrained network, dubbed LipNet, and integrate an explicit prompt module to provide discriminative knowledge of different sparse sampling settings, enabling the treatment of multiple sparse view configurations within a single model. Furthermore, we develop a storage-saving deep unfolding framework for multiple-in-one SVCT reconstruction, termed PromptCT, which embeds LipNet as its prior network to ensure the convergence of its corresponding iterative algorithm. In simulated and real data experiments, PromptCT outperforms benchmark reconstruction algorithms in multiple-in-one SVCT reconstruction, achieving higher-quality reconstructions with lower storage costs. On the theoretical side, we explicitly demonstrate that LipNet satisfies boundary property, further proving its Lipschitz continuity and subsequently analyzing the convergence of the proposed iterative algorithms. The data and code are publicly available at https://github.com/shibaoshun/PromptCT.

cs.CV

OpenMoCap: Rethinking Optical Motion Capture under Real-world Occlusion

Optical motion capture is a foundational technology driving advancements in cutting-edge fields such as virtual reality and film production. However, system performance suffers severely under large-scale marker occlusions common in real-world applications. An in-depth analysis identifies two primary limitations of current models: (i) the lack of training datasets accurately reflecting realistic marker occlusion patterns, and (ii) the absence of training strategies designed to capture long-range dependencies among markers. To tackle these challenges, we introduce the CMU-Occlu dataset, which incorporates ray tracing techniques to realistically simulate practical marker occlusion patterns. Furthermore, we propose OpenMoCap, a novel motion-solving model designed specifically for robust motion capture in environments with significant occlusions. Leveraging a marker-joint chain inference mechanism, OpenMoCap enables simultaneous optimization and construction of deep constraints between markers and joints. Extensive comparative experiments demonstrate that OpenMoCap consistently outperforms competing methods across diverse scenarios, while the CMU-Occlu dataset opens the door for future studies in robust motion solving. The proposed OpenMoCap is integrated into the MoSen MoCap system for practical deployment. The code is released at: https://github.com/qianchen214/OpenMoCap.

cs.CV

edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer

Vision-Language Models (VLMs) are increasingly deployed in real-time applications such as autonomous driving and human-computer interaction, which demand fast and reliable responses based on accurate perception. To meet these requirements, existing systems commonly employ cloud-edge collaborative architectures, such as partitioned Large Vision-Language Models (LVLMs) or task offloading strategies between Large and Small Vision-Language Models (SVLMs). However, these methods fail to accommodate cloud latency fluctuations and overlook the full potential of delayed but accurate LVLM responses. In this work, we propose a novel cloud-edge collaborative paradigm for VLMs, termed Context Transfer, which treats the delayed outputs of LVLMs as historical context to provide real-time guidance for SVLMs inference. Based on this paradigm, we design edgeVLM, which incorporates both context replacement and visual focus modules to refine historical textual input and enhance visual grounding consistency. Extensive experiments on three real-time vision-lanuage reasoning tasks across four datasets demonstrate the effectiveness of the proposed framework. The new paradigm lays the groundwork for more effective and latency-aware collaboration strategies in future VLM systems.

cs.CV

Conformally compact metrics and the Lovelock tensors

We study conformally compact metrics satisfying the Lovelock equations, which generalize the Einstein equation. We show that these metrics admit polyhomogeneous expansions, thereby naturally realizing the Fefferman-Graham expansion, which is an important tool in conformal geometry and the AdS/CFT correspondence. In even dimensions, we identify a boundary obstruction to smoothness near the boundary that generalizes the ambient obstruction tensor in the Einstein setting. Under appropriate regularity and curvature conditions, we also construct a formal solution to the singular Yamabe-(2q) problem and provide an index obstruction for the conformally compact Lovelock filling problem of spin manifolds.

math.DG

Witten instanton complex and Morse-Bott inequalities on stratified pseudomanifolds

In this paper we construct Witten instanton complexes on stratified pseudomanifolds with wedge metrics, for all choices of mezzo-perversities which classify the self-adjoint extensions of the Hodge Dirac operator. In this singular setting we introduce a generalization of the Morse-Bott condition and in so doing can consider a class of functions with certain non-isolated critical point sets which arise naturally in many examples. This construction of the instanton complex extends the Morse polynomial to this setting from which we prove the corresponding Morse inequalities. This work proceeds by constructing Hilbert complexes and normal cohomology complexes, including those corresponding to the Witten deformed complexes for such critical point sets and all mezzo-perversities; these in turn are used to express local Morse polynomials as polynomial trace formulas over their cohomology groups. Under a technical assumption of `flatness' on our Morse-Bott functions we then construct the instanton complex by extending the local harmonic forms to global quasimodes. We also study the Poincar\'e dual complexes and in the case of self-dual complexes extract refined Morse inequalities generalizing those in the smooth setting. We end with a guide for computing local cohomology groups and Morse polynomials.

math.DG

A Dual-domain Regularization Method for Ring Artifact Removal of X-ray CT

Ring artifacts in computed tomography images, arising from the undesirable responses of detector units, significantly degrade image quality and diagnostic reliability. To address this challenge, we propose a dual-domain regularization model to effectively remove ring artifacts, while maintaining the integrity of the original CT image. The proposed model corrects the vertical stripe artifacts on the sinogram by innovatively updating the response inconsistency compensation coefficients of detector units, which is achieved by employing the group sparse constraint and the projection-view direction sparse constraint on the stripe artifacts. Simultaneously, we apply the sparse constraint on the reconstructed image to further rectified ring artifacts in the image domain. The key advantage of the proposed method lies in considering the relationship between the response inconsistency compensation coefficients of the detector units and the projection views, which enables a more accurate correction of the response of the detector units. An alternating minimization method is designed to solve the model. Comparative experiments on real photon counting detector data demonstrate that the proposed method not only surpasses existing methods in removing ring artifacts but also excels in preserving structural details and image fidelity.

eess.IV