arXiv ScienceSearch

arXiv subjects

Chao Shi

Publications and source records attributed to Chao Shi.

At least 19 recordsLinked to original sources

Artemis: Anatomy-Resolved inTervention for Eliminating Multimodal NeuroImage confounderS

Multimodal neuroimaging, integrating functional connectivity from fMRI and structural connectivity from DTI, enables non-invasive analysis of brain networks using graph neural networks. However, demographic factors such as age and sex systematically confound the relationship between brain connectivity and clinical outcomes, causing GNNs to exploit spurious shortcuts rather than learning causally invariant representations. While recent causal GNN methods introduce causality at the graph-modeling level, their causal mechanisms remain domain-agnostic without accounting for the real-world confounders inherent in clinical neuroimaging data. Moreover, brain networks are constructed from atlas-based parcellations where each region exhibits distinct sensitivity to demographic factors, necessitating region-aware adjustment. We propose Artemis, a region-level causal framework that bridges this gap with causal intervention at each brain region independently by learning region-specific confounder representations with lightweight parameters. Our adjustment comprehensively utilized the multimodal functional and structural features for graph reasoning as a plug-in module compatible with arbitrary GNN backbones. Experiments on three benchmarks, ADNI for disease diagnosis, OASIS for dementia staging, and HCP for sex classification, demonstrate consistent improvements over representative GNN-based baselines. Multiple supporting experiments further demonstrate statistical significance and neuroscientific interpretability.

cs.LG

MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models

In the progress of industrial anomaly detection, general anomaly detection (GAD) is an emerging trend and also the ultimate goal. Unlike the conventional single- and multi-class AD, general AD aims to train a general AD model that can directly detect anomalies in diverse novel classes without any retraining or fine-tuning on the target data. Recently, Multimodal Large Language Models (MLLMs) have shown great promise in achieving general anomaly detection due to their revolutionary visual understanding and language reasoning capabilities. However, MLLM's general AD ability remains underexplored due to: (1) MLLMs are pretrained on amounts of data sourced from the Web, these data still have significant gaps with the data in AD scenarios. Moreover, the image-text pairs during pretraining are also not specifically for AD tasks. (2) The current mainstream AD datasets are image-based and not yet suitable for post-training MLLMs. To facilitate MLLM-based general AD research, we present MMR-AD, which is a comprehensive benchmark for both training and evaluating MLLM-based AD models. With MMR-AD, we reveal that the AD performance of current SOTA generalist MLLMs still falls far behind the industrial requirements. Based on MMR-AD, we also propose a baseline model, Anomaly-R1, which is a reasoning-based AD model that learns from the CoT data in MMR-AD and is further enhanced by reinforcement learning. Extensive experiments show that our Anomaly-R1 achieves remarkable improvements over generalist MLLMs in both anomaly detection and localization.

cs.CV

Pion and $\rho$ meson's unpolarized quark distribution functions from $q\bar{q}$ and all Fock-states within Dyson--Schwinger equations

We compute the twist-2 unpolarized quark parton distribution functions (PDFs) of the pion and the $\rho$ meson within the Dyson-Schwinger equations (DSEs) framework using the rainbow-ladder (RL) truncation. A new DSE for the spin-1 hadron's quark-quark correlation matrix is derived, from which the PDFs can be extracted. For the $\rho$ meson, we obtain for the first time within RL-DSEs the unpolarized PDFs corresponding to different helicity states. A pronounced difference is observed between the $|\Lambda|=1$ and $\Lambda=0$ cases, where $\Lambda$ denotes the meson helicity, leading to a nonvanishing and numerically sizable tensor-polarized PDF $f_{1LL}(x)$. We further compare these results with those obtained under a leading Fock-state ($|q\bar{q}\rangle$) truncation and find substantial deviations. This comparison demonstrates that the present RL-DSEs framework incorporates higher Fock-state contributions associated with gluonic degrees of freedom, which has a significant impact on the quark distributions.

hep-ph

An imprint of intrinsic quark--gluon correlations: a nonmonotonic feature in $e(x)$

We present the first rainbow-ladder Dyson-Schwinger equations study of the pion's chiral-odd, twist-3 parton distribution $e_{\rm q}(x)$. By deriving a novel Dyson-Schwinger equation for the pion's quark-quark correlation matrix, we simultaneously extract from it both the unpolarized twist-2 parton distribution function $f_{\rm q}(x)$ and the twist-3 distribution $e_{\rm q}(x)$. Our results show that chiral symmetry strongly suppresses the twist-2 component of $e_{\rm q}(x)$, leaving the genuine twist-3 quark-gluon component dominant and featuring a node. We therefore argue that the genuine twist-3 term can make a substantial contribution to hadronic $e(x)$, producing a nonmonotonic structure that is a clear imprint of intrinsic quark-gluon correlations. We point out that the existence of a hump-like feature is compatible with, although not uniquely indicated by, recent proton extractions and awaits more precise determination.

hep-ph

HRM^2Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular Phone Scans

We present HRM$^2$Avatar, a framework for creating high-fidelity avatars from monocular phone scans, which can be rendered and animated in real time on mobile devices. Monocular capture with smartphones provides a low-cost alternative to studio-grade multi-camera rigs, making avatar digitization accessible to non-expert users. Reconstructing high-fidelity avatars from single-view video sequences poses challenges due to limited visual and geometric data. To address these limitations, at the data level, our method leverages two types of data captured with smartphones: static pose sequences for texture reconstruction and dynamic motion sequences for learning pose-dependent deformations and lighting changes. At the representation level, we employ a lightweight yet expressive representation to reconstruct high-fidelity digital humans from sparse monocular data. We extract garment meshes from monocular data to model clothing deformations effectively, and attach illumination-aware Gaussians to the mesh surface, enabling high-fidelity rendering and capturing pose-dependent lighting. This representation efficiently learns high-resolution and dynamic information from monocular data, enabling the creation of detailed avatars. At the rendering level, real-time performance is critical for animating high-fidelity avatars in AR/VR, social gaming, and on-device creation. Our GPU-driven rendering pipeline delivers 120 FPS on mobile devices and 90 FPS on standalone VR devices at 2K resolution, over $2.7\times$ faster than representative mobile-engine baselines. Experiments show that HRM$^2$Avatar delivers superior visual realism and real-time interactivity, outperforming state-of-the-art monocular methods.

cs.GR

ResAD++: Towards Class Agnostic Anomaly Detection via Residual Feature Learning

This paper explores the problem of class-agnostic anomaly detection (AD), where the objective is to train one class-agnostic AD model that can generalize to detect anomalies in diverse new classes from different domains without any retraining or fine-tuning on the target data. When applied for new classes, the performance of current single- and multi-class AD methods is still unsatisfactory. One fundamental reason is that representation learning in existing methods is still class-related, namely, feature correlation. To address this issue, we propose residual features and construct a simple but effective framework, termed ResAD. Our core insight is to learn the residual feature distribution rather than the initial feature distribution. Residual features are formed by matching and then subtracting normal reference features. In this way, we can effectively realize feature decorrelation. Even in new classes, the distribution of normal residual features would not remarkably shift from the learned distribution. In addition, we think that residual features still have one issue: scale correlation. To this end, we propose a feature hypersphere constraining approach, which learns to constrain initial normal residual features into a spatial hypersphere for enabling the feature scales of different classes as consistent as possible. Furthermore, we propose a novel logbarrier bidirectional contraction OCC loss and vector quantization based feature distribution matching module to enhance ResAD, leading to the improved version of ResAD (ResAD++). Comprehensive experiments on eight real-world AD datasets demonstrate that our ResAD++ can achieve remarkable AD results when directly used in new classes, outperforming state-of-the-art competing methods and also surpassing ResAD. The code is available at https://github.com/xcyao00/ResAD.

cs.CV

Diffractive electroproduction of light vector particles: leading Fock-state contribution in the presence of significant higher Fock-state effects

We study exclusive diffractive production of vector mesons and photon using the color dipole model with leading Fock state light front wave functions derived from Dyson Schwinger and Bethe Salpeter equations. New results for the $\phi$ meson and real photon are presented. Without data fitting, our calculation well matches HERA data in certain kinematical domains. The key finding of this paper is that in a color dipole model study for $\rho/\gamma$ and $\phi$, where light quarks are involved, the leading $q\bar{q}$ approximation is valid only when $Q^2$ exceeds $20$ and 10 GeV$^2$ respectively, unlike $J/\psi$ which can be well described for $Q^2\approx 0$ GeV$^2$. This underscores the special role of $\phi$ electroproduction in color dipole picture: it strikes a balance between the large dipole size typical of light mesons and the smaller size associated with high $Q^2$ photons, making it potentially well suited for probing gluon saturation effects.

hep-ph

Convergence in charmonium structure: light-front wave functions from basis light-front quantization and Dyson-Schwinger equations

We present a systematic comparison of charmonium light-front wave functions obtained through two complementary non-perturbative approaches: Basis Light-Front Quantization (BLFQ) and Dyson-Schwinger equations (DSE). Key observables include the charge form factor, gravitational form factors, light-cone distribution amplitudes, decay constants, and two-photon transition form factors. Despite their distinct theoretical foundations and model parameters, the predictions from BLFQ and DSE exhibit remarkable agreement across all observables. This convergence validates both frameworks for studying charmonium structure and highlights the complementary strengths of Hamiltonian-based (BLFQ) and Lagrangian-based (DSE) methods in addressing non-perturbative QCD.

hep-ph

FlexiNS: A SmartNIC-Centric, Line-Rate and Flexible Network Stack

As the gap between network and CPU speeds rapidly increases, the CPU-centric network stack proves inadequate due to excessive CPU and memory overhead. While hardware-offloaded network stacks alleviate these issues, they suffer from limited flexibility in both control and data planes. Offloading network stack to off-path SmartNIC seems promising to provide high flexibility; however, throughput remains constrained by inherent SmartNIC architectural limitations. To this end, we design FlexiNS, a SmartNIC-centric network stack with software transport programmability and line-rate packet processing capabilities. To grapple with the limitation of SmartNIC-induced challenges, FlexiNS introduces: (a) a header-only offloading TX path; (b) an unlimited-working-set in-cache processing RX path; (c) a high-performance DMA-only notification pipe; and (d) a programmable offloading engine. We prototype FlexiNS using Nvidia BlueField-3 SmartNIC and provide out-of-the-box RDMA IBV verbs compatibility to users. FlexiNS achieves 2.2$\times$ higher throughput than the microkernel-based baseline in block storage disaggregation and 1.3$\times$ higher throughput than the hardware-offloaded baseline in KVCache transfer.

cs.NI

Revisiting Computational Storage for Data Integrity and Security

The idea of computational storage device (CSD) has come a long way since at least 1990s [1], [2]. By embedding computing resources within storage devices, CSDs could potentially offload computational tasks from CPUs and enable near-data processing (NDP), reducing data movements and/or energy consumption significantly. While the initial hard-disk-based CSDs suffer from severe limitations in terms of on-drive resources, programmability, etc., the storage market has witnessed the commercialization of solid-state-drive (SSD) based CSDs (e.g., Samsung SmartSSD [3], ScaleFlux CSDs [4]) recently, which has enabled CSD-based optimizations for avariety of application scenarios (e.g., [5], [6], [7]).

cs.DC

Noncollinear ferroelectric and screw-type antiferroelectric phases in a metal-free hybrid molecular crystal

Noncollinear dipole textures greatly extend the scientific merits and application perspective of ferroic materials. In fact, noncollinear spin textures have been well recognized as one of the core issues of condensed matter, e.g. cycloidal/conical magnets with multiferroicity and magnetic skyrmions with topological properties. However, the counterparts in electrical polarized materials are less studied and thus urgently needed, since electric dipoles are usually aligned collinearly in most ferroelectrics/antiferroelectrics. Molecular crystals with electric dipoles provide a rich ore to explore the noncollinear polarity. Here we report an organic salt (H2Dabco)BrClO4 (H2Dabco = N,N'-1,4-diazabicyclo[2.2.2]octonium) that shows a transition between the ferroelectric and antiferroelectric phases. Based on experimental characterizations and ab initio calculations, it is found that its electric dipoles present nontrivial noncollinear textures with $60^\circ$-twisting angle between the neighbours. Then the ferroelectric-antiferroelectric transition can be understood as the coding of twisting angle sequence. Our study reveals the unique science of noncollinear electric polarity.

cond-mat.mtrl-sci

Heavy flavor-asymmetric pseudoscalar mesons on the light front

We extract the leading Fock-state light front wave functions (LF-LFWFs) of heavy flavor-asymmetric pseudoscalar mesons $D$, $B$ and $B_c$ from their Bethe-Salpeter wave functions based on Dyson-Schwinger equations approach, and study their leading twist parton distribution amplitudes, generalized parton distribution functions and transverse momentum dependent parton distributions. The spatial distributions of the quark and antiquark on the transverse plane are given, along with their charge and energy distributions on the light front. We find that in the considered mesons, the heavier quarks carry most longitudinal momentum fraction and yield narrow $x$-distributions, while the lighter quarks play an active role in shaping the transverse distributions within both spatial and momentum space, exhibiting a duality embodying characteristics from both light mesons and heavy quarkonium.

hep-ph

HGIC: A Hand Gesture Based Interactive Control System for Efficient and Scalable Multi-UAV Operations

As technological advancements continue to expand the capabilities of multi unmanned-aerial-vehicle systems (mUAV), human operators face challenges in scalability and efficiency due to the complex cognitive load and operations associated with motion adjustments and team coordination. Such cognitive demands limit the feasible size of mUAV teams and necessitate extensive operator training, impeding broader adoption. This paper developed a Hand Gesture Based Interactive Control (HGIC), a novel interface system that utilize computer vision techniques to intuitively translate hand gestures into modular commands for robot teaming. Through learning control models, these commands enable efficient and scalable mUAV motion control and adjustments. HGIC eliminates the need for specialized hardware and offers two key benefits: 1) Minimal training requirements through natural gestures; and 2) Enhanced scalability and efficiency via adaptable commands. By reducing the cognitive burden on operators, HGIC opens the door for more effective large-scale mUAV applications in complex, dynamic, and uncertain scenarios. HGIC will be open-sourced after the paper being published online for the research community, aiming to drive forward innovations in human-mUAV interactions.

cs.RO

Demystifying Datapath Accelerator Enhanced Off-path SmartNIC

Network speeds grow quickly in the modern cloud, so SmartNICs are introduced to offload network processing tasks, even application logic. However, typical multicore SmartNICs such as BlueFiled-2 are only capable of processing control-plane tasks with their embedded processors that have limited memory bandwidth and computing power. On the other hand, cloud applications evolve rapidly, such that a limited number of fixed hardware engines in a SmartNIC cannot satisfy the requirements of cloud applications. Therefore, SmartNIC programmers call for a programmable datapath accelerator (DPA) to process network traffic at line rate. However, no existing work has unveiled the performance characteristics of the existing DPA. To this end, we present the first architectural characterization of the latest DPA-enhanced BlueFiled-3 (BF3) SmartNIC. Our evaluation results indicate that BF3's DPA is significantly wimpier than the off-path Arm processor and the host CPU. However, we still identify that DPA has three unique architectural characteristics that unleash the performance potential of DPA. Specifically, we demonstrate how to take advantage of DPA's three architectural characteristics regarding computing, networking, and memory subsystems. Then we propose three important guidelines for programmers to fully unleash the potential of DPA. To demonstrate the effectiveness of our approach, we conduct detailed case studies regarding each guideline. Our case study on key-value aggregation achieves up to 4.3$\times$ higher throughput by using our guidelines to optimize memory combinations.

cs.NI

Strongly interacting matter in a sphere at nonzero magnetic field

We investigate the chiral phase transition within a sphere under a uniform background magnetic field. The Nambu--Jona-Lasinio (NJL) model is employed and the MIT boundary condition is imposed for the spherical confinement. Using the wave expansion method, the diagonalizable Hamiltonian and energy spectrum are derived for the system. By solving the gap equation in the NJL model, the influence of magnetic field on quark matter in a sphere is studied. It is found that inverse magnetic catalysis occurs at small radii, while magnetic catalysis occurs at large radii. Additionally, both magnetic catalysis and inverse magnetic catalysis are observed at the intermediate radii ($R\approx4$ fm).

nucl-th

Nonperturbative photon $q\bar{q}$ light front wave functions from a contact interaction model

We propose a method to calculate the $q\bar{q}$ light front wave functions (LFWFs) of photon at low-virtuality, i.e., the light front amplitude of $\gamma^*\rightarrow q\bar{q}$ at low $Q^2$, based on a light front projection approach. We exemplify this method using a contact interaction model within Dyson-Schwinger equations formalism and obtain the nonperturbative photon $q\bar{q}$ LFWFs. In this case, we find the nonperturbative effects are encoded in the enhanced quark mass and a dressing function of covariant quark-photon vertex, as compared to the leading order quantum electrodynamics photon $q\bar{q}$ LFWFs. We then use nonperturbative-effect modified photon $q\bar{q}$ LFWFs to study the inclusive deep inelastic scattering HERA data in the framework of the color dipole model. The results demonstrate that the theoretical description of data at low $Q^2$ can be significantly improved once the nonperturbative corrections are included in the photon LFWFs.

hep-ph

Federated Learning in Big Model Era: Domain-Specific Multimodal Large Models

Multimodal data, which can comprehensively perceive and recognize the physical world, has become an essential path towards general artificial intelligence. However, multimodal large models trained on public datasets often underperform in specific industrial domains. This paper proposes a multimodal federated learning framework that enables multiple enterprises to utilize private domain data to collaboratively train large models for vertical domains, achieving intelligent services across scenarios. The authors discuss in-depth the strategic transformation of federated learning in terms of intelligence foundation and objectives in the era of big model, as well as the new challenges faced in heterogeneous data, model aggregation, performance and cost trade-off, data privacy, and incentive mechanism. The paper elaborates a case study of leading enterprises contributing multimodal data and expert knowledge to city safety operation management , including distributed deployment and efficient coordination of the federated learning platform, technical innovations on data quality improvement based on large model capabilities and efficient joint fine-tuning approaches. Preliminary experiments show that enterprises can enhance and accumulate intelligent capabilities through multimodal model federated learning, thereby jointly creating an smart city model that provides high-quality intelligent services covering energy infrastructure safety, residential community security, and urban operation management. The established federated learning cooperation ecosystem is expected to further aggregate industry, academia, and research resources, realize large models in multiple vertical domains, and promote the large-scale industrial application of artificial intelligence and cutting-edge research on multimodal federated learning.

cs.LG

Transvese momentum dependent parton distributions of pion at leading twist

We calculate the leading twist pion unpolarized transverse momentum distribution $f_1(x,k_T^2)$ and the Boer-Mulders function $h_1^\perp(x,k_T^2)$, using leading Fock-state light front wave functions (LF-LFWFs) based on Dyson-Schwinger and Bethe-Salpeter equations. These DS-BSEs based LF-LFWFs provide dynamically generated s- and p-wave components, which are indispensable in producing chirally odd Boer-Mulders function that has one parton spin flipped. Employing a non-perturbative SU(3) gluon rescattering kernel to treat the gauge link of the Boer-Mulders function, we thus obtain both TMDs at hadronic scale and then evolve them to the scale of $\mu^2=4.0$ GeV$^2$. We finally calculate the generalized Boer-Mulders shift and find it to be in agreement with the lattice prediction.

hep-ph