arXiv ScienceSearch

arXiv subjects

Ningsheng Xu

Publications and source records attributed to Ningsheng Xu.

At least 19 recordsLinked to original sources

MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline

Recent advances in deep research systems enable large language models to retrieve, synthesize, and reason over large-scale external knowledge. In medicine, developing clinical guidelines critically depends on such deep evidence integration. However, existing benchmarks fail to evaluate this capability in realistic workflows requiring multi-step evidence integration and expert-level judgment. To address this gap, we introduce MedProbeBench, the first benchmark leveraging high-quality clinical guidelines as expert-level references. Medical guidelines, with their rigorous standards in neutrality and verifiability, represent the pinnacle of medical expertise and pose substantial challenges for deep research agents. For evaluation, we propose MedProbe-Eval, a comprehensive evaluation framework featuring: (1) Holistic Rubrics with 1,200+ task-adaptive rubric criteria for comprehensive quality assessment, and (2) Fine-grained Evidence Verification for rigorous validation of evidence precision, grounded in 5,130+ atomic claims. Evaluation of 17 LLMs and deep research agents reveals critical gaps in evidence integration and guideline generation, underscoring the substantial distance between current capabilities and expert-level clinical guideline development. Project: https://github.com/uni-medical/MedProbeBench

cs.CV

MedQ-Engine: A Closed-Loop Data Engine for Evolving MLLMs in Medical Image Quality Assessment

Medical image quality assessment (Med-IQA) is a prerequisite for clinical AI deployment, yet multimodal large language models (MLLMs) still fall substantially short of human experts, particularly when required to provide descriptive assessments with clinical reasoning beyond simple quality scores. However, improving them is hindered by the high cost of acquiring descriptive annotations and by the inability of one-time data collection to adapt to the model's evolving weaknesses. To address these challenges, we propose MedQ-Engine, a closed-loop data engine that iteratively evaluates the model to discover failure prototypes via data-driven clustering, explores a million-scale image pool using these prototypes as retrieval anchors with progressive human-in-the-loop annotation, and evolves through quality-assured fine-tuning, forming a self-improving cycle. Models are evaluated on complementary perception and description tasks. An entropy-guided routing mechanism triages annotations to minimize labeling cost. Experiments across five medical imaging modalities show that MedQ-Engine elevates an 8B-parameter model to surpass GPT-4o by over 13% and narrow the gap with human experts to only 4.34%, using only 10K annotations with more than 4x sample efficiency over random sampling.

cs.CV

MedQ-UNI: Toward Unified Medical Image Quality Assessment and Restoration via Vision-Language Modeling

Existing medical image restoration (Med-IR) methods are typically modality-specific or degradation-specific, failing to generalize across the heterogeneous degradations encountered in clinical practice. We argue this limitation stems from the isolation of Med-IR from medical image quality assessment (Med-IQA), as restoration models without explicit quality understanding struggle to adapt to diverse degradation types across modalities. To address these challenges, we propose MedQ-UNI, a unified vision-language model that follows an assess-then-restore paradigm, explicitly leveraging Med-IQA to guide Med-IR across arbitrary modalities and degradation types. MedQ-UNI adopts a multimodal autoregressive dual-expert architecture with shared attention: a quality assessment expert first identifies degradation issues through structured natural language descriptions, and a restoration expert then conditions on these descriptions to perform targeted image restoration. To support this paradigm, we construct a large-scale dataset of approximately 50K paired samples spanning three imaging modalities and five restoration tasks, each annotated with structured quality descriptions for joint Med-IQA and Med-IR training, along with a 2K-sample benchmark for evaluation. Extensive experiments demonstrate that a single MedQ-UNI model, without any task-specific adaptation, achieves state-of-the-art restoration performance across all tasks while generating superior descriptions, confirming that explicit quality understanding meaningfully improves restoration fidelity and interpretability.

cs.CV

Open World MRI Reconstruction with Bias-Calibrated Adaptation

Real-world MRI reconstruction systems face the open-world challenge: test data from unseen imaging centers, anatomical structures, or acquisition protocols can differ drastically from training data, causing severe performance degradation. Existing methods struggle with this challenge. To address this, we propose BiasRecon, a bias-calibrated adaptation framework grounded in the minimal intervention principle: preserve what transfers, calibrate what does not. Concretely, BiasRecon formulates open-world adaptation as an alternating optimization framework that jointly optimizes three components: (1) frequency-guided prior calibration that introduces layer-wise calibration variables to selectively modulate frequency-specific features of the pre-trained score network via self-supervised k-space signals, (2) score-based denoising that leverages the calibrated generative prior for high-fidelity image reconstruction, and (3) adaptive regularization that employs Stein's Unbiased Risk Estimator to dynamically balance the prior-measurement trade-off, matching test-time noise characteristics without requiring ground truth. By intervening minimally and precisely through this alternating scheme, BiasRecon achieves robust adaptation with fewer than 100 tunable parameters. Extensive experiments across four datasets demonstrate state-of-the-art performance on open-world reconstruction tasks.

eess.IV

MedQ-Deg: A Multidimensional Benchmark for Evaluating MLLMs Across Medical Image Quality Degradations

Despite impressive performance on standard benchmarks, multimodal large language models (MLLMs) face critical challenges in real-world clinical environments where medical images inevitably suffer various quality degradations. Existing benchmarks exhibit two key limitations: (1) absence of large-scale, multidimensional assessment across medical image quality gradients and (2) no systematic confidence calibration analysis. To address these gaps, we present MedQ-Deg, a comprehensive benchmark for evaluating medical MLLMs under image quality degradations. MedQ-Deg provides multi-dimensional evaluation spanning 18 distinct degradation types, 30 fine-grained capability dimensions, and 7 imaging modalities, with 24,894 question-answer pairs. Each degradation is implemented at 3 severity degrees, calibrated by expert radiologists. We further introduce Calibration Shift metric, which quantifies the gap between a model's perceived confidence and actual performance to assess metacognitive reliability under degradation. Our comprehensive evaluation of 40 mainstream MLLMs reveals several critical findings: (1) overall model performance degrades systematically as degradation severity increases, (2) models universally exhibit the AI Dunning-Kruger Effect, maintaining inappropriately high confidence despite severe accuracy collapse, and (3) models display markedly differentiated behavioral patterns across capability dimensions, imaging modalities, and degradation types. We hope MedQ-Deg drives progress toward medical MLLMs that are robust and trustworthy in real clinical practice.

cs.CV

Methods for characterization of atomic-scale field emission point-electron-source

Field emission (FE) electron sources are made close to atomic-scale to reach the highest spatial resolution as well as stable emission for electron microscopy, electron beam inspection and lithography. At present, no single agreed method exists of using FE current-voltage data to extract the apparent emission area, which is needed for predicting some beam properties. The 1956 theory of Murphy and Good (MG) is better physics than the 1920s theory of Fowler and Nordheim (FN) and colleagues, but many researchers use simplified FN theory to analyse experimental data. The present paper reports an experimental method of finding apparent emission area, based on using field ion and field electron microscopes (FIM-FEM). The discrepancy of emission area between the FIM-FEM method and MG-based analysis is a factor of 7.4, while that with simplified FN-based analysis is about 25, confirming MG theory is better for FE data analysis. The result allows deduction of key indicators, including source energy spread, reduced brightness and emission efficiency. A downloadable program is made available to help analysis. Our work provides a new experimental method of characterizing FE electron sources, especially the atomic-scale cold cathode, for which existing plot-based data-analysis methods are not suitable.

quant-ph

Deep-Subwavelength Plasmon Polariton Atomic Cavity Detector for Frequency- and Polarization-Sensitive Terahertz Detection and Imaging

Room-temperature, miniaturized, polarization-resolved terahertz (THz) detection of high speed is vital for high-resolution imaging in radar, remote sensing, and semiconductor inspection, and is essential for large-scale THz focal plane arrays. However, miniaturization below deep-subwavelength scales (< 1/50 wavelength) remain challenging due to weak light-matter interaction, which degrades responsivity and polarization sensitivity. Here, we present a graphene plasmon polariton atomic cavity (PPAC) monolithic detector that overcomes this limitation by maintaining and even enhancing performance at a deep-subwavelength channel length of just 2 micrometers (1/60 wavelength). The device integrates graphene rectangle PPAC arrays with dissimilar metal contacts, where graphene functions as both absorber and conductor, simplifying the architecture. Exploiting plasmon polariton resonances and the photothermoelectric (PTE) effect, the detector achieves polarization-sensitive, frequency-selective, and fast THz detection spanning 0.53 to 4.24 THz with a polarization ratio of 93, featuring a responsivity (RV) of 1007 V/W, a noise-equivalent power (NEP) of 16 pW/Hz^0.5, a specific detectivity (D*) of 2.9 x 10^7 Jones, and a response time of 230 ps. We further demonstrate monolithic integration for polarization imaging and non-destructive semiconductor chip inspection, advancing room-temperature, compact, and polarization-sensitive THz technologies.

physics.optics

Value of Multi-pursuer Single-evader Pursuit-evasion Game with Terminal Cost of Evader's Position: Relaxation of Convexity Condition

In this study, we consider a multi-pursuer single-evader quantitative pursuit-evasion game with payoff function that includes only the terminal cost. The terminal cost is a function related only to the terminal position of the evader. This problem has been extensively studied in target defense games. Here, we prove that a candidate for the value function generated by geometric method is the viscosity solution of the corresponding Hamilton-Jacobi-Isaacs partial differential equation (HJI PDE) Dirichlet problem. Therefore, the value function of the game at each point can be computed by a mathematical program. In our work, the convexity of the terminal cost or the target is not required. The terminal cost only needs to be locally Lipschitz continuous. The cases in which the terminal costs or the targets are not convex are covered. Therefore, our result is more universal than those of previous studies, and the complexity of the proof is improved. We also discuss the optimal strategies in this game and present an intuitive explanation of this value function.

math.OC

GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution Detection

Recent advancements have explored text-to-image diffusion models for synthesizing out-of-distribution (OOD) samples, substantially enhancing the performance of OOD detection. However, existing approaches typically rely on perturbing text-conditioned embeddings, resulting in semantic instability and insufficient shift diversity, which limit generalization to realistic OOD. To address these challenges, we propose GOOD, a novel and flexible framework that directly guides diffusion sampling trajectories towards OOD regions using off-the-shelf in-distribution (ID) classifiers. GOOD incorporates dual-level guidance: (1) Image-level guidance based on the gradient of log partition to reduce input likelihood, drives samples toward low-density regions in pixel space. (2) Feature-level guidance, derived from k-NN distance in the classifier's latent space, promotes sampling in feature-sparse regions. Hence, this dual-guidance design enables more controllable and diverse OOD sample generation. Additionally, we introduce a unified OOD score that adaptively combines image and feature discrepancies, enhancing detection robustness. We perform thorough quantitative and qualitative analyses to evaluate the effectiveness of GOOD, demonstrating that training with samples generated by GOOD can notably enhance OOD detection performance.

cs.CV

MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs

Medical Image Quality Assessment (IQA) serves as the first-mile safety gate for clinical AI, yet existing approaches remain constrained by scalar, score-based metrics and fail to reflect the descriptive, human-like reasoning process central to expert evaluation. To address this gap, we introduce MedQ-Bench, a comprehensive benchmark that establishes a perception-reasoning paradigm for language-based evaluation of medical image quality with Multi-modal Large Language Models (MLLMs). MedQ-Bench defines two complementary tasks: (1) MedQ-Perception, which probes low-level perceptual capability via human-curated questions on fundamental visual attributes; and (2) MedQ-Reasoning, encompassing both no-reference and comparison reasoning tasks, aligning model evaluation with human-like reasoning on image quality. The benchmark spans five imaging modalities and over forty quality attributes, totaling 2,600 perceptual queries and 708 reasoning assessments, covering diverse image sources including authentic clinical acquisitions, images with simulated degradations via physics-based reconstructions, and AI-generated images. To evaluate reasoning ability, we propose a multi-dimensional judging protocol that assesses model outputs along four complementary axes. We further conduct rigorous human-AI alignment validation by comparing LLM-based judgement with radiologists. Our evaluation of 14 state-of-the-art MLLMs demonstrates that models exhibit preliminary but unstable perceptual and reasoning skills, with insufficient accuracy for reliable clinical use. These findings highlight the need for targeted optimization of MLLMs in medical IQA. We hope that MedQ-Bench will catalyze further exploration and unlock the untapped potential of MLLMs for medical image quality evaluation.

cs.CV

Deep Learning Empowered Sub-Diffraction Terahertz Backpropagation Single-Pixel Imaging

Terahertz single-pixel imaging (THz SPI) has garnered widespread attention for its potential to overcome challenges associated with THz focal plane arrays. However, the inherently long wavelength of THz waves limits imaging resolution, while achieving subwavelength resolution requires harsh experimental conditions and time-consuming processes. Here, we propose a sub-diffraction THz backpropagation SPI technique. We illuminate the object with continuous-wave 0.36-THz radiation ({\lambda}0 = 833.3 {\mu}m). The transmitted THz wave is modulated by prearranged patterns generated on a 500-{\mu}m-thick silicon wafer and subsequently recorded by a far-field single-pixel detector. An untrained neural network constrained with the physical SPI process iteratively reconstructs the THz images with an ultralow sampling ratio of 1.5625%, significantly reducing the long sampling times. To further suppress the THz diffraction-field effects, a backpropagation SPI from near field to far field is implemented by integrating with a THz physical propagation model into the output layer of the network. Notably, using the thick wafer where THz evanescent field cannot be fully recorded, we achieve a spatial resolution of 118 {\mu}m (~{\lambda}0/7) through backpropagation SPI, thus eliminating the need for ultrathin photomodulators. This approach provides an efficient solution for advancing THz microscopic imaging and addressing other inverse imaging challenges.

eess.IV

Dominance Regions of Pursuit-evasion Games in Non-anticipative Information Patterns

The evader's dominance region is an important concept and the foundation of geometric methods for pursuit-evasion games. This article mainly reveals the relevant properties of the evader's dominance region, especially in non-anticipative information patterns. We can use these properties to research pursuit-evasion games in non-anticipative information patterns. The core problem is under what condition the pursuer has a non-anticipative strategy to prevent the evader leaving its initial dominance region before being captured regardless of the evader's strategy. We first define the evader's dominance region by the shortest path distance, and we rigorously prove for the first time that the initial dominance region of the evader is the reachable region of the evader in the open-loop sense. Subsequently, we prove that there exists a non-anticipative strategy by which the pursuer can capture the evader before the evader leaves its initial dominance region's closure in the absence of obstacles. For cases with obstacles, we provide a counter example to illustrate that such a non-anticipative strategy does not always exist, and provide a necessary condition for the existence of such strategy. Finally, we consider a scenario with a single corner obstacle and provide a sufficient condition for the existence of such a non-anticipative strategy. At the end of this article, we discuss the application of the evader's dominance region in target defense games. This article has important reference significance for the design of non-anticipative strategies in pursuit-evasion games with obstacles.

math.OC

Generalized Huang's Equation for Phonon Polariton in Polyatomic Polar Crystal

The original theory of phonon polariton is Huang's equation which is suitable for diatomic polar crystals only. We proposed a generalized Huang's equation without fitting parameters for phonon polariton in polyatomic polar crystals. We obtained the dispersions of phonon polariton in GaP (bulk), hBN (bulk and 2D), {\alpha}-MoO3 (bulk and 2D) and ZnTeMoO6 (2D), which agree with the experimental results in the literature and of ourselves. We also obtained the eigenstates of the phonon polariton. We found that the circular polarization of the ion vibration component of these eigenstates is nonzero in hBN flakes. The result is different from that of the phonon in hBN.

cond-mat.mtrl-sci

Harmonizing Material Quantity and Terahertz Wave Interference Shielding Efficiency with Metallic Borophene Nanosheets

Materials with electromagnetic interference (EMI) shielding in the terahertz (THz) regime, while minimizing the quantity used, are highly demanded for future information communication, healthcare and mineral resource exploration applications. Currently, there is often a trade-off between the amount of material used and the absolute EMI shielding effectiveness (EESt) for the EMI shielding materials. Here, we address this trade-off by harnessing the unique properties of two-dimensional (2D) beta12-borophene (beta12-Br) nanosheets. Leveraging beta12-Br's light weight and exceptional electron mobility characteristics, which represent among the highest reported values to date, we simultaneously achieve a THz EMI shield effectiveness (SE) of 70 dB and an EESt of 4.8E5 dB cm^2/g (@0.87 THz) using a beta12-Br polymer composite. This surpasses the values of previously reported THz shielding materials with an EESt less than 3E5 dB cm^2/g and a SE smaller than 60 dB, while only needs 0.1 wt.% of these materials to realize the same SE value. Furthermore, by capitalizing on the composite's superior mechanical properties, with 158% tensile strain at a Young's modulus of 33 MPa, we demonstrate the high-efficiency shielding performances of conformably coated surfaces based on beta12-Br nanosheets, suggesting their great potential in EMI shielding area.

physics.app-ph

Multi-modal MRI Translation via Evidential Regression and Distribution Calibration

Multi-modal Magnetic Resonance Imaging (MRI) translation leverages information from source MRI sequences to generate target modalities, enabling comprehensive diagnosis while overcoming the limitations of acquiring all sequences. While existing deep-learning-based multi-modal MRI translation methods have shown promising potential, they still face two key challenges: 1) lack of reliable uncertainty quantification for synthesized images, and 2) limited robustness when deployed across different medical centers. To address these challenges, we propose a novel framework that reformulates multi-modal MRI translation as a multi-modal evidential regression problem with distribution calibration. Our approach incorporates two key components: 1) an evidential regression module that estimates uncertainties from different source modalities and an explicit distribution mixture strategy for transparent multi-modal fusion, and 2) a distribution calibration mechanism that adapts to source-target mapping shifts to ensure consistent performance across different medical centers. Extensive experiments on three datasets from the BraTS2023 challenge demonstrate that our framework achieves superior performance and robustness across domains.

eess.IV

Monolithic Multi-parameter Terahertz Nano-micro Detector Based on Plasmon Polariton Atomic Cavity

Terahertz signals hold significant potential for ultra-wideband communication and high-resolution radar, necessitating miniaturized detectors capable of multi-parameter detection of intensity, frequency, polarization, and phase. Conventional detectors cannot meet these requirements. Here, we propose plasmon polariton atomic cavities (PPAC) made from single-atom-thick graphene, demonstrating the monolithic multifunctional miniaturized detector. With a footprint one-tenth the incident wavelength, the detector offers benchmarking intensity-, frequency-, and polarization-sensitive detection, rapid response, and sub-diffraction spatial resolution, all operating at room temperature across 0.22 to 4.24 THz. We present the monolithic detection applications for free-space THz polarization-coded communication and stealth imaging of physical properties. These results showcase the PPAC's unique ability to achieve strong absorption and weak signal detection with a thickness of only 10^-5 of the excitation wavelength, which is inaccessible with other approaches.

physics.app-ph

Digital Twin Brain: a simulation and assimilation platform for whole human brain

In this work, we present a computing platform named digital twin brain (DTB) that can simulate spiking neuronal networks of the whole human brain scale and more importantly, a personalized biological brain structure. In comparison to most brain simulations with a homogeneous global structure, we highlight that the sparseness, couplingness and heterogeneity in the sMRI, DTI and PET data of the brain has an essential impact on the efficiency of brain simulation, which is proved from the scaling experiments that the DTB of human brain simulation is communication-intensive and memory-access intensive computing systems rather than computation-intensive. We utilize a number of optimization techniques to balance and integrate the computation loads and communication traffics from the heterogeneous biological structure to the general GPU-based HPC and achieve leading simulation performance for the whole human brain-scaled spiking neuronal networks. On the other hand, the biological structure, equipped with a mesoscopic data assimilation, enables the DTB to investigate brain cognitive function by a reverse-engineering method, which is demonstrated by a digital experiment of visual evaluation on the DTB. Furthermore, we believe that the developing DTB will be a promising powerful platform for a large of research orients including brain-inspiredintelligence, rain disease medicine and brain-machine interface.

cs.NE

N\'eel-type optical skyrmions inherited from evanescent electromagnetic fields with rotational symmetry

Optical skyrmions, the optical analogue of topological configurations formed by three-dimensional vector fields covering the whole 4{\pi} solid angle but confined in a two-dimensional (2D) domain, have recently attracted growing interest due to their potential applications in high-density data transfer, storage, and processing. While the optical skyrmions have been successfully demonstrated using different field vectors in both of free-space propagating and near-field evanescent electromagnetic fields, the study on generation and control of the optical skyrmions, and their general correlation with the electromagnetic (EM) fields, are still in infancy. Here, we theoretically propose that an evanescent transverse-magnetic-polarized (TM-polarized) EM fields with rotational symmetry are actually N\'eel-type optical skyrmions of the electric field vectors. Such optical skyrmions maintain the rotation symmetry that are independent on the operation frequency and medium. Our proposal was verified by numerical simulations and real-space nano-imaging experiments performed on a graphene monolayer. Such a discovery can therefore not only further our understanding on the formation mechanisms of EM topological textures, but also provide a guideline for facile construction of EM skyrmions that may impact future information technologies.

physics.optics