arXiv ScienceSearch

arXiv subjects

Hui Zhu

Publications and source records attributed to Hui Zhu.

At least 19 recordsLinked to original sources

Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization

Large Vision-Language Models (LVLMs) are prone to hallucinations: they fluently describe objects, attributes, and scenes that are not in the image. We connect part of this failure to a measurable property of their representations, feature instability, where mild semantics-preserving perturbations of the input cause large changes in the learned embeddings; hallucination rates rise together with this variability. Existing stability-motivated remedies are explicit, in the sense that they intervene at inference time through latent steering or constrained decoding, and pay for it on every query. We propose implicit stabilization instead: perturbation-invariance is built into the model weights during fine-tuning, and nothing extra runs at deployment. Our framework, INFUSE, first stabilizes visual and textual representations around perturbation-averaged and ground-truth anchors, then aligns the stabilized representations across modalities with bidirectional contrastive objectives. We prove that the anchor's root-mean-square deviation from the perturbation-mean representation shrinks at rate $1/\sqrt{K}$ in the number of views, and that under a Lipschitz decoder, this bounds how much any perturbation can change the model's hallucination behavior. On LLaVA-1.5, LLaVA-1.6, and Qwen3-VL-8B-Instruct, INFUSE reduces AMBER CHAIR by 46-63% relative to each base model, improves ObjHal, MMHal, HallusionBench, and POPE, and preserves VQA-v2 and TextVQA, all with no inference-time overhead.

cs.CV

MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class-level and region-level concept alignment to organize the shared representation at complementary granularities. Class-level alignment anchors each anatomical target to an aggregated clinical concept profile, while region-level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class-specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late-stage cue. MedPlex achieves state-of-the-art performance across CT and MR benchmarks for multi-organ, cardiac substructure, and tumor segmentation, including settings with real free-text clinical supervision. Code: https://github.com/rafiibnsultan/MedPlex.

cs.CV

LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine

Biomedical knowledge graphs (KGs) are pivotal for knowledge organization, yet traditional binary relations often struggle to represent the conditional nature of biomedical knowledge. Symptoms provide a shared phenotypic layer for linking Traditional Chinese Medicine (TCM), which relies on symptom patterns for syndrome differentiation and treatment selection, with modern biomedicine, which connects clinical manifestations to diseases and molecular mechanisms. We present LingShu, a large-scale symptom-centric contextualized knowledge graph designed to bridge TCM and modern biomedicine. The exported version of LingShu analyzed in this study comprises 17.33 million atom-level entity records and 39.47 million relation records, including 17.19 million semantic triples and 22.29 million contextualized quadruples. LingShu integrates multi-source data, including clinical electronic medical records, authoritative TCM texts, biomedical ontologies, and curated knowledge bases, through a pipeline combining natural language processing, terminology normalization, and human-in-the-loop verification. A key innovation of LingShu is its hybrid data model: it maintains 64 typed triple relation patterns to ensure broad connectivity, while incorporating 35 contextual quadruple relation patterns to capture conditional medical associations. This dual-structure approach explicitly encodes conditional knowledge, providing a granular representation of the contexts associated with medical relations. These contextualized relations cover syndrome-dependent herb efficacy, disease-contextualized drug effects, population-specific clinical associations, and mechanism-related therapeutic responses. Furthermore, we developed a web platform (http://www.tcmkg.com/) that integrates graph visualization, graph-based reasoning, and an evidence-grounded knowledge question-answering agent.

cs.CL

FAST Discovery of $\mu$Jy Radio Pulsations from PSR J2238+5903, Providing a DM Distance Anchor for the Candidate TeV Halo 1LHAASO J2238+5900

We report the first detection of radio pulsations from PSR J2238+5903, a gamma-ray pulsar spatially coincident with the extended TeV source 1LHAASO J2238+5900. Our 3000 s FAST L-band observation reveals a weak periodic signal at the known Fermi-LAT spin period, with $P=162.76568$ ms and $\mathrm{DM}=247.5\pm3.0~\mathrm{pc~cm^{-3}}$. The signal is independently confirmed by both FFT-based and Fast Folding Algorithm searches. The radiometer equation gives a flux density of $S_{1250}\simeq3\,\mu$Jy, placing PSR J2238+5903 among the faintest radio-detected Fermi pulsars. Interpreting the DM with Galactic electron-density models gives $d_{\rm DM}=7.4\pm3.9$ kpc. At this distance, the LHAASO WCDA 39\% containment radius corresponds to a characteristic diameter of $\sim132$ pc, and the $>1$ TeV luminosity is $L_{\rm TeV}\simeq7.1\times10^{34}$ erg s$^{-1}$, about 8\% of the pulsar's spin-down power. The radio DM thus provides the first pulsar-specific distance constraint for assessing whether 1LHAASO J2238+5900 is a young relic-PWN / TeV-halo transition system.

astro-ph.HE

MyGO-Splat: Multi-Objective Closed-Loop Geometric Feedback for RGB-Only Gaussian SLAM

Real-time monocular Simultaneous Localization and Mapping (SLAM) fundamentally suffers from scale ambiguity and a lack of geometric self-correction. While 3D Gaussian Splatting (3DGS) enables high-fidelity rendering, existing RGB-only systems remain open-loop because depth priors are injected into mapping but refined geometry cannot effectively regulate tracking drift. We present MyGO-Splat, a closed-loop Gaussian SLAM framework that analytically rasterizes Gaussian primitives into pixel-wise depth and surface normals, allowing the map to actively supervise camera pose optimization. To bridge monocular priors and scale consistency, our framework introduces scale-aware adaptive alignment that projects foundation-model depth estimates into the globally optimized Gaussian space, forming a self-correcting cycle for scale feedback. Extensive evaluations show that this closed-loop design improves scale stability and appearance-geometry consistency, achieving performance comparable to RGB-D methods while using only monocular input.

cs.RO

MMD-SLAM: Structure-Enhanced Multi-Meta Gaussian Distribution-Guided Visual SLAM

3D Gaussian Splatting (3DGS) has significantly boosted novel view synthesis and high-fidelity scene reconstruction, expanding the potential of 3DGS-based Visual Simultaneous Localization and Mapping (SLAM) methods. However, most existing systems fail to fully exploit the underlying structural information, which limits rendering quality and often leads to inconsistent maps. To address these limitations, we propose MMD-SLAM, a structure-enhanced Visual SLAM framework that leverages the Atlanta World (AW) assumption to guide a Multi-Meta Gaussian representation for photorealistic mapping. First, we introduce a point-line fusion strategy for pose optimization, where 3D line segments are incorporated to improve tracking robustness and provide additional constraints for mapping. Second, we design a Multi-Meta Gaussian representation with dominant directions, explicitly encoding structural priors from the AW hypothesis. Finally, we propose a Gaussian evolution strategy that adapts to scene geometry and incorporates structural cues into global optimization. Extensive experiments demonstrate that these innovations enable MMD-SLAM to achieve state-of-the-art performance in both tracking accuracy and mapping quality. e.g., our method achieves a 48.56% reduction in ATE RMSE on ScanNet and a 5.71% improvement in PSNR on Replica, compared with MonoGS.

cs.RO

WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation

Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision-Language Models (LVLMs) struggle to meet. Although these models can describe visual content, their lack of explicit grounding leads to object hallucinations and unreliable depth reasoning, limiting their usefulness for accessibility guidance. We introduce WalkGPT, a pixel-grounded LVLM for the new task of Grounded Navigation Guide, unifying language reasoning and segmentation within a single architecture for depth-aware accessibility guidance. Given a pedestrian-view image and a navigation query, WalkGPT generates a conversational response with segmentation masks that delineate accessible and harmful features, along with relative depth estimation. The model incorporates a Multi-Scale Query Projector (MSQP) that shapes the final image tokens by aggregating them along text tokens across spatial hierarchies, and a Calibrated Text Projector (CTP), guided by a proposed Region Alignment Loss, that maps language embeddings into segmentation-aware representations. These components enable fine-grained grounding and depth inference without user-provided cues or anchor points, allowing the model to generate complete and realistic navigation guidance. We also introduce PAVE, a large-scale benchmark of 41k pedestrian-view images paired with accessibility-aware questions and depth-grounded answers. Experiments show that WalkGPT achieves strong grounded reasoning and segmentation performance. The source code and dataset are available on the \href{https://sites.google.com/view/walkgpt-26/home}{project website}.

cs.CV

LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning

Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are fragmented in coverage and scale and rely on non-unified or generation-heavy scoring that hinders fair comparison. We present the LingLanMiDian (LingLan) benchmark, a large-scale, expert-curated, multi-task suite that unifies evaluation across knowledge recall, multi-hop reasoning, information extraction, and real-world clinical decision-making. LingLan introduces a consistent metric design, a synonym-tolerant protocol for clinical labels, a per-dataset 400-item Hard subset, and a reframing of diagnosis and treatment recommendation into single-choice decision recognition. We conduct comprehensive, zero-shot evaluations on 14 leading open-source and proprietary LLMs, providing a unified perspective on their strengths and limitations in TCM commonsense knowledge understanding, reasoning, and clinical decision support; critically, the evaluation on Hard subset reveals a substantial gap between current models and human experts in TCM-specialized reasoning. By bridging fundamental knowledge and applied reasoning through standardized evaluation, LingLan establishes a unified, quantitative, and extensible foundation for advancing TCM LLMs and domain-specific medical AI research. All evaluation data and code are available at https://github.com/TCMAI-BJTU/LingLan and http://tcmnlp.com.

cs.AI

A Study Revealing Physical Attributes of Supernova Remnant in G321.3-3.9

We present a radio analysis of the recently identified supernova remnant G321.3-3.9 using archival multi-wavelength data spanning 88-2304 MHz. The source exhibits an elliptical shell-like morphology (1.3 deg x 1.7 deg) and a relatively flat non-thermal spectral index of alpha = -0.40 +/- 0.03. The distance is estimated using both the Sigma-D relation (1.6-2.9 kpc) and tentative associations with HI structures, the latter suggesting a near-side solution of 2.5-3.3 kpc, though the physical connection remains uncertain.

astro-ph.GA

Resolvent bounds imply observability from measurable time sets for Schr\"odinger equations

We prove that on a compact Riemannian manifold, resolvent bounds for the Laplace--Beltrami operator imply observability, and thus controllability, for the Schr\"odinger propagator from time sets of positive Lebesgue measure. Applications include almost all cases where observability and controllability hold from time intervals, particularly when the geometric control condition is satisfied or when the manifold is a compact surface of negative curvature.

math.AP

Observability of Schr\"odinger propagators on tori in rough settings

On tori of arbitrary dimensions, Schr\"odinger propagators with bounded potentials are conjectured to be observable from space-time domains of positive Lebesgue measure. We reduce this conjecture to certain integrability bounds for free Schr\"odinger waves, thereby proving the conjecture on the one-dimensional torus and producing new examples of observation domains. These bounds are far weaker than Bourgain's conjectured periodic Strichartz estimates, yet remain highly nontrivial.

math.AP

Kangaroo: A Private and Amortized Inference Framework over WAN for Large-Scale Decision Tree Evaluation

With the rapid adoption of Models-as-a-Service, concerns about data and model privacy have become increasingly critical. To solve these problems, various privacy-preserving inference schemes have been proposed. In particular, due to the efficiency and interpretability of decision trees, private decision tree evaluation (PDTE) has garnered significant attention. However, existing PDTE schemes suffer from significant limitations: their communication and computation costs scale with the number of trees, the number of nodes, or the tree depth, which makes them inefficient for large-scale models, especially over WAN networks. To address these issues, we propose Kangaroo, a private and amortized decision tree inference framework build upon packed homomorphic encryption. Specifically, we design a novel model hiding and encoding scheme, together with secure feature selection, oblivious comparison, and secure path evaluation protocols, enabling full amortization of the overhead as the number of nodes or trees scales. Furthermore, we enhance the performance and functionality of the framework through optimizations, including same-sharing-for-same-model, latency-aware, and adaptive encoding adjustment strategies. Kangaroo achieves a $14\times$ to $59\times$ performance improvement over state-of-the-art (SOTA) one-round interactive schemes in WAN environments. For large-scale decision tree inference tasks, it delivers a $3\times$ to $44\times$ speedup compared to existing schemes. Notably, Kangaroo enables the evaluation of a random forest with $969$ trees and $411825$ nodes in approximately $60$ ms per tree (amortized) under WAN environments.

cs.CR

FGO-SLAM: Enhancing Gaussian SLAM with Globally Consistent Opacity Radiance Field

Visual SLAM has regained attention due to its ability to provide perceptual capabilities and simulation test data for Embodied AI. However, traditional SLAM methods struggle to meet the demands of high-quality scene reconstruction, and Gaussian SLAM systems, despite their rapid rendering and high-quality mapping capabilities, lack effective pose optimization methods and face challenges in geometric reconstruction. To address these issues, we introduce FGO-SLAM, a Gaussian SLAM system that employs an opacity radiance field as the scene representation to enhance geometric mapping performance. After initial pose estimation, we apply global adjustment to optimize camera poses and sparse point cloud, ensuring robust tracking of our approach. Additionally, we maintain a globally consistent opacity radiance field based on 3D Gaussians and introduce depth distortion and normal consistency terms to refine the scene representation. Furthermore, after constructing tetrahedral grids, we identify level sets to directly extract surfaces from 3D Gaussians. Results across various real-world and large-scale synthetic datasets demonstrate that our method achieves state-of-the-art tracking accuracy and mapping performance.

cs.RO

Trace and Observability Inequalities for Laplace Eigenfunctions on the Torus

We investigate trace and observability inequalities for Laplace eigenfunctions on the d-dimensional torus, with respect to arbitrary Borel measures $\mu$. Specifically, we characterize the measures $\mu$ for which the inequalities $$ \int |u|^2 d \mu \lesssim \int |u|^2 d x \quad \text{(trace)}, \qquad \int |u|^2 d \mu \gtrsim \int |u|^2 d x \quad \text{(observability)}$$ hold uniformly for all eigenfunctions $u$ of the Laplacian. Sufficient conditions are derived based on the integrability and regularity of $\mu$, while necessary conditions are formulated in terms of the dimension of the support of the measure. These results generalize classical theorems of Zygmund and Bourgain--Rudnick to higher dimensions. Applications include results in the spirit of Cantor--Lebesgue theorems, constraints on quantum limits, and control theory for the Schr\"odinger equation. Our approach combines several tools: the cluster structure of lattice points on spheres; decoupling estimates; and the construction of eigenfunctions exhibiting strong concentration or vanishing behavior, tailored respectively to the trace and observability inequalities.

math.AP

High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery

Object detection in Unmanned Aerial Vehicle (UAV) imagery is fundamentally challenged by a prevalence of small, densely packed, and occluded objects within cluttered backgrounds. Conventional detectors struggle with this domain, as they rely on hand-crafted components like pre-defined anchors and heuristic-based Non-Maximum Suppression (NMS), creating a well-known performance bottleneck in dense scenes. Even recent end-to-end frameworks have not been purpose-built to overcome these specific aerial challenges, resulting in a persistent performance gap. To bridge this gap, we introduce HEDS-DETR, a holistically enhanced real-time Detection Transformer tailored for aerial scenes. Our framework features three key innovations. First, we propose a novel High-Frequency Enhanced Semantics Network (HFESNet) backbone, which yields highly discriminative features by preserving critical high-frequency details alongside robust semantic context. Second, our Efficient Small Object Pyramid (ESOP) counteracts information loss by efficiently fusing high-resolution features, significantly boosting small object detection. Finally, we enhance decoder stability and localization precision with two synergistic components: Selective Query Recollection (SQR) and Geometry-Aware Positional Encoding (GAPE), which stabilize optimization and provide explicit spatial priors for dense object arrangements. On the VisDrone dataset, HEDS-DETR achieves a +3.8% AP and +5.1% AP50 gain over its baseline while reducing parameters by 4M and maintaining real-time speeds. This demonstrates a highly competitive accuracy-efficiency balance, especially for detecting dense and small objects in aerial scenes.

cs.CV

Primes of the form $ax+by$

For two coprime positive integers $a,b$, let $T(a,b)=\{ ax+by : x,y\in \mathbb{Z}_{\ge 0} \} $ and let $s(a,b)=ab-a-b$. It is well known that all integers which are greater than $s(a,b)$ are in $T(a,b)$. Let $\pi (a, b)$ be the number of primes in $T(a,b)$ which are less than or equal to $s(a,b)$. It is easy to see that $\pi (2, 3)=0$ and $\pi (2, b)=1$ for all odd integers $b\ge 5$. In this paper, we prove that if $b>a\ge 3$ with $\gcd (a, b)=1$, then $\pi (a, b)>0.005 s(a,b)/\log s(a,b)$. We conjecture that $\frac{13}{66}\pi (s(a,b))\le \pi (a, b)\le \frac 12\pi (s(a,b))$ for all $b>a\ge 3$ with $\gcd (a, b)=1$.

math.NT

The number of primes not in a numerical semigroup

For two coprime positive integers $a$ and $b$,let $\pi^* (a, b)$ be the number of primes that cannot be represented as $au+bv$, where $u$ and $v$ are nonnegative integers. It is clear that $\pi^* (a, b)\le \pi (ab-a-b)$, where $\pi (x)$ denotes the number of primes not exceeding $x$. In this paper, we prove that $\pi^* (a, b)\ge 0.04\pi (ab-a-b)$ and pose following conjecture: $\pi^* (a, b)\ge \frac 12 \pi (ab-a-b)$. This conjecture is confirmed for $1\le a\le 10$.

math.NT

Automatic Calibration for Membership Inference Attack on Large Language Models

Membership Inference Attacks (MIAs) have recently been employed to determine whether a specific text was part of the pre-training data of Large Language Models (LLMs). However, existing methods often misinfer non-members as members, leading to a high false positive rate, or depend on additional reference models for probability calibration, which limits their practicality. To overcome these challenges, we introduce a novel framework called Automatic Calibration Membership Inference Attack (ACMIA), which utilizes a tunable temperature to calibrate output probabilities effectively. This approach is inspired by our theoretical insights into maximum likelihood estimation during the pre-training of LLMs. We introduce ACMIA in three configurations designed to accommodate different levels of model access and increase the probability gap between members and non-members, improving the reliability and robustness of membership inference. Extensive experiments on various open-source LLMs demonstrate that our proposed attack is highly effective, robust, and generalizable, surpassing state-of-the-art baselines across three widely used benchmarks. Our code is available at: \href{https://github.com/Salehzz/ACMIA}{\textcolor{blue}{Github}}.

cs.LG