arXiv ScienceSearch

arXiv subjects

Yong Xia

Publications and source records attributed to Yong Xia.

At least 19 recordsLinked to original sources

The Projected Hessian Quantification Theorem: Exact Duality For Constrained Eigenvalues

The classical Projected Hessian Lemma, originating from Finsler's theorem, characterizes definiteness over a constraint null space through quadratic penalties. However, it does not quantify the corresponding constrained eigenvalues or the associated eigenvalue penalty path. This work develops a quantitative penalty theory for constrained symmetric eigenvalue problems and establishes exact characterizations of the extremal eigenvalues of the reduced Hessian via full-space penalized eigenvalue problems. Three proofs are provided based on orthogonal decomposition, Schur complement analysis, and semidefinite programming duality. We characterize finite exact recovery along the extremal-eigenvalue penalty paths and, in the absence of finite recovery, establish asymptotic convergence with a first-order error expansion and an explicit leading coefficient. The associated Hellmann--Feynman sensitivity relation leads to a strategy for predicting penalty parameters. Based on these results, we develop a matrix-free Penalty--Split--Merge method for successive constrained extremal eigenpairs using penalty continuation, Split--Merge iterations, deflation, and projected certification. Numerical experiments illustrate the predicted penalty regimes, evaluate projected certification, and assess computational performance on moderate-scale benchmarks and large-scale matrix-free test instances.

math.OC

Discrete Potential Optimization for Absolute Value Equations: A Sign-Flip Framework with Polynomial Complexity

Solving the AVE $Ax - |x| = b$ is generally NP-hard. Existing approaches mostly operate in continuous variable spaces, with the notable exception of Rohn's sign-accord algorithm, which incurs an exponential worst-case bound of $2^n$ iterations. To overcome this bottleneck, we develop a discrete potential optimization (DPO) framework over the sign-vector set $\{-1, 1\}^n$. Under the $1$-norm condition $\|A^{-1}\|_1 < 1/2$, we establish that the AVE solution corresponds exactly to the global maximizer of this discrete potential function. When $\|A^{-1}\|_1$ is uniformly upper bounded by $1/2$, for rational inputs with maximum magnitude $L$, we develop a unified polynomial-time framework for sign-flip algorithms, which excludes Rohn's sign-accord algorithm. Within this framework, specific single-flip mechanisms (including our new steepest and Gauss-Southwell rules) terminate in $\mathcal{O}(n^2\log(nL))$ iterations, while full-flip updates (equivalent to the classical GNM) require only $\mathcal{O}(n\log(nL))$ iterations. We further relax the assumption by requiring that the spectral radius $ρ(|A^{-1}|)$ be uniformly upper bounded by $1/2$ via rational diagonal scaling. Moreover, a uniformly randomized $m$-flip approach is proven to achieve an expected iteration bound of $\mathcal{O}\bigl(n^3 \log(nL)/m\bigr)$ without requiring explicit diagonal preconditioning. As a corollary, GNM solves the AVE in $\mathcal{O}(n^2 \log(nL))$ iterations, improving on the prior result of finite termination under the stricter condition that $ρ(|A^{-1}|)$ is less than~$1/3$. Crucially, by equivalently reformulating linear complementarity problems (LCPs) as AVEs, we extend this DPO framework to yield GNM and pivot-type methods with polynomial iteration complexity for LCPs. Numerical experiments validate the practical efficiency of GNM and the structural robustness of the proposed sign-flip approaches.

math.OC

A Line-Search-Free Coordinate Proximal Predictor-Corrector Method for Monotone Absolute Value Equations

We consider the absolute value equation (AVE) $Ax-|x|=b$ under residual monotonicity. We characterize this property through the symmetric part of the coefficient matrix, thereby allowing nonsymmetry, and propose a coordinate proximal predictor-corrector (CPPC) method based on an AVE-specific forward-proximal decomposition. An exact scalar proximal update on the largest proximal-residual coordinate generates a predictor point, while a positive-alignment identity certifies the ensuing full-residual separating-hyperplane correction without backtracking. With the current matrix product cached, each iteration requires one new full matrix-vector product. For a nonempty solution set, we prove Fejér monotonicity, whole-sequence convergence, and an $O(K^{-1/2})$ best-iterate residual bound. A positive monotonicity margin further ensures unique solvability for every right-hand side and global linear convergence. Numerical results identify regimes in which the reduced per-iteration work yields shorter solution times.

math.OC

Volumetric Radiology AI in the Era of Multimodal Large Language Models

Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual representations, or report-derived text. Reliable volumetric radiology AI therefore requires representations that preserve task-relevant three-dimensional (3D) information and systems that can access, verify, and integrate this information across clinical workflows. In this Review, we examine more than 200 publications through July 2026. We organize the literature around volumetric representation and multimodal understanding at the model level, agentic orchestration at the system level, and their links to clinical applications and evaluation. We review volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction. We distinguish settings in which selected 2D views or report-mediated reasoning may suffice from those that warrant native volumetric modeling. We also introduce a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation. Across the literature, native volumetric modeling and agentic capabilities depend on the spatial, quantitative, contextual, and workflow requirements of the intended task. Clinical credibility requires faithful volumetric representation, traceable system behavior, claim-aligned validation, and clearly defined human oversight in realistic workflows.

cs.AI

Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation along a corresponding ray path. Reconstructing a CT volume from only a few X-ray views is therefore severely ill-posed, as the projections collapse depth information and leave 3D locations of anatomical regions and their corresponding intensity distributions highly entangled and ambiguous. We observe that once the spatial organization of anatomical regions is established, estimating their CT intensities becomes substantially more tractable. Motivated by this, we propose LiftXR, an interleaved, geometry-guided framework that explicitly incorporates spatial layout recovery into CT reconstruction. Specifically, a layout lifter first generates a 3D anatomical layout from bi-planar X-rays, providing spatial guidance for an intensity renderer to reconstruct a CT volume. An anatomical parser then performs volumetric perception on the reconstruction, exploiting its spatially resolved boundary and intensity cues to recover a refined anatomical layout. This transition from projection-conditioned layout generation to reconstruction-conditioned anatomical perception allows the parsed layout to provide feedback for region-specific intensity calibration. Extensive experiments on two public datasets demonstrate that LiftXR consistently outperforms recent X-ray-to-CT reconstruction methods, establishing a new state of the art. Moreover, the reconstructed CT achieves superior performance in external downstream segmentation, indicating improved anatomical fidelity. Code will be released.

cs.CV

Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy

Medical images require comprehensive and accurate interpretation to support the diagnosis of diverse clincial conditions. Recent vision-language generalist models offer broad task coverage and promising zero-shot capabilities, yet often lack fine-grained anatomical and lesion awareness for reliable diagnosis and spatial interpretability. In contrast, supervised specialist models achieve strong performance on specific tasks but typically lack generalization across diseases and anatomies. In this work, we present SuG, a Super-Generalist framework that unifies generalist vision-language learning with specialist objectives, enabling both broad generalization and specialist-level diagnostic capability. We perform specialist-enhanced vision-language alignment in SuG by incorporating spatial priors from multiple segmentation experts, including anatomy, class-specific lesion and class-agnostic lesion segmentors that captures lesions beyond anatomies annotated during training. To improve lesion grounding capability, we leverage lesion masks as spatial priors to calibrate text-conditioned visual attention, encouraging disease-related semantics to focus on clinically relevant regions. We evaluate SuG on extensive chest and abdominal CT benchmarks, including CT-RATE, Merlin, MedVL-CT69K, and several in-house tumor datasets. SuG achieves state-of-the-art performance across a wide range of disease diagnosis tasks and surpasses specialist models on several critical tumor diagnosis benchmarks. Furthermore, SuG demonstrates strong lesion grounding capability, including robust generalization to lesion types lacking class-specific supervision.

cs.CV

The Continuous Relaxation of Sparse PCA is NP-hard

Maximizing a symmetric quadratic form under simultaneous L1 norm inequality and L2 norm equality constraints is a standard and widely used continuous relaxation for Sparse Principal Component Analysis (SPCA). This paper settles the computational complexity of this continuous formulation by proving it is NP-hard. Furthermore, the variant with both L1 and L2 norm inequalities is also shown to be NP-hard.

math.OC

FedBiCross: Personalized One-Shot Federated Learning on Medical Images

Data-free knowledge distillation-based one-shot federated learning (OSFL) trains a model in a single communication round without sharing raw data, making OSFL attractive for privacy-sensitive medical applications. However, existing methods aggregate predictions from all clients to form a global teacher. Under non-IID data, conflicting predictions dilute each other during averaging, yielding less informative soft labels that weaken distillation. We propose FedBiCross, a personalized OSFL framework with three stages: (1) clustering clients by model output similarity to form coherent sub-ensembles, (2) bi-level cross-cluster optimization that learns adaptive weights to selectively leverage beneficial cross-cluster knowledge while suppressing negative transfer, and (3) personalized distillation for client-specific adaptation. Experiments on four medical image datasets demonstrate that FedBiCross consistently outperforms state-of-the-art baselines across different non-IID degrees.

cs.LG

Global optimization of quadratic root-difference minimization under elliptic annulus constraints

This paper studies the nonconvex quadratic root-difference minimization under elliptic annulus constraints {\rm (QR)}. We first establish the Annulus Brickman theorem and equivalently reformulate {\rm (QR)} as a 2-dimensional convex problem {\rm (HP)} with hidden variables. We employ the Frank-Wolfe algorithm to globally solve {\rm (HP)}. A key finding is that the solutions of the Frank-Wolfe subproblems, which are traditionally viewed as mere auxiliary updates, are proven to be $O(1/\sqrt{k})$-approximate solutions of the original problem {\rm (QR)}. This transforms an algorithmic by-product into the primary output and completely bypasses the need to solve the computationally expensive quadratic system required for solution recovery. Leveraging this recovery-free property, we develop the efficient Iterative Minimum Generalized Eigenpair (IMGE) algorithm for globally solving {\rm (QR)}. Numerical experiments confirm that IMGE converges rapidly and significantly outperforms conventional methods, especially for large-scale problems.

math.OC

Split-Merge: A Difference-based Approach for Dominant Eigenvalue Problem

The computation of the dominant eigenpair for symmetric positive semidefinite matrices is fundamental in numerical optimization. This work shifts the paradigm from the classical Rayleigh quotient to an unconstrained difference formulation, whose global optimum recovers the dominant eigenpair. Within this framework, we prove that gradient descent with a constant step-size $α\in (0, 1)$ converges almost surely to the global optimum at a local linear rate. This analysis thereby reinterprets the classical power method as the conservative special case $α=1/2$ and rigorously establishes its asymptotic sub-optimality. To advance this first-order scheme, we propose the Split-Merge algorithm based on the majorization-minimization principle. After splitting the matrix, we introduce auxiliary vectors to effectively merge the decomposition factors, resulting in a matrix-free and parameter-free iteration that captures tighter curvature information. We establish that Split-Merge converges almost surely to a global minimizer, and show that the iteration exhibits a spectral peeling mechanism that suppresses the targeted eigenspace, potentially surpassing the static linear rate of power iterations. Numerical evaluations across synthetic and real-world datasets confirm that our method has scalable efficiency, achieving speed-ups exceeding $10\times$ over the power method, with performance comparable to subspace iterations.

math.OC

Stealthy Patch-Wise Backdoor Attack in 3D Point Cloud via Curvature Awareness

Backdoor attacks pose a severe threat to deep neural networks (DNNs) by implanting hidden backdoors that can be activated with predefined triggers to manipulate model behaviors maliciously. Recent studies have extended backdoor attacks to 3D point clouds, but most existing triggers are sample-wise and often cause visible geometric artifacts or high optimization cost. To address these limitations, we propose the Stealthy Patch-Wise Backdoor Attack (SPBA), a patch-wise backdoor attack framework for 3D point clouds. Specifically, SPBA decomposes a point cloud into local patches, where each patch is formed by a Farthest Point Sampling (FPS) center and its K-nearest neighbors (KNN). Candidate patches are ranked using a patch imperceptibility score derived from local curvature variation, and a unified spectral trigger is injected into the selected patches by perturbing only the coordinates of existing points while preserving the original point cardinality. Extensive experiments on ModelNet40 and ShapeNetPart further demonstrate that SPBA achieves state-of-the-art stealthiness among prior methods and reduces spectral-trigger computation by 98.43% relative to a sample-wise spectral baseline, while maintaining competitive attack performance. These results support localized spectral design as an effective and efficient approach to stealthy backdoor attacks in 3D point cloud models. Code is available at https://github.com/HazardFY/SPBA.

cs.CV

Upper Generalization Bounds for Neural Oscillators

Neural oscillators that originate from second-order ordinary differential equations (ODEs) have shown competitive performance in learning mappings between dynamic loads and responses of complex nonlinear structural systems. Despite this empirical success, theoretically quantifying the generalization capacities of their neural network architectures remains undeveloped. In this study, the neural oscillator consisting of a second-order ODE followed by a multilayer perceptron (MLP) is considered. Its upper probably approximately correct (PAC) generalization bound for approximating causal and uniformly continuous operators between continuous temporal function spaces and that for approximating the uniformly asymptotically incrementally stable second-order dynamical systems are derived by leveraging the Rademacher complexity framework. These bounds are further extended to the squared Wasserstein-1 distances between the probability measures of quantities of interest calculated from target causal operators and the corresponding learned neural oscillators. The theoretical results show that the estimation errors grow polynomially with respect to both MLP sizes and the time length, thereby avoiding the curse of parametric complexity. Furthermore, the derived error bounds demonstrate that constraining the Lipschitz constants of the MLPs via loss function regularization can improve the generalization ability of the neural oscillator. Numerical studies considering a Bouc-Wen nonlinear system under stochastic seismic excitation validates the theoretically predicted power laws of the estimation errors with respect to the sample size and time length, and confirms the effectiveness of constraining MLPs' matrix and vector norms in enhancing the performance of the neural oscillator under limited training data.

cs.LG

Upper Approximation Bounds for Neural Oscillators

Neural oscillators, originating from second-order ordinary differential equations (ODEs), have demonstrated strong performance in stably learning causal mappings between long-term sequences or continuous temporal functions, as well as in accurately approximating physical systems. However, theoretically quantifying the capacities of their neural network architectures remains a significant challenge. In this study, the neural oscillator consisting of a second-order ODE followed by a multilayer perceptron (MLP) is considered. Its upper approximation bound for approximating causal and uniformly continuous operators between continuous temporal function spaces and that for approximating uniformly asymptotically incrementally stable second-order dynamical systems are derived. The established proof method of the approximation bound for approximating the causal continuous operators can also be directly applied to state-space models consisting of a linear time-continuous complex recurrent neural network followed by an MLP. Theoretical results reveal that the approximation error of the neural oscillator for approximating the second-order dynamical systems scales polynomially with the reciprocals of the widths of two utilized MLPs, thus overcoming the curse of parametric complexity. The convergence rates of two established approximation error bounds are validated through four numerical cases. These results provide a robust theoretical foundation for the effective application of the neural oscillator in science and engineering.

cs.LG

SPP-SBL: Space-Power Prior Sparse Bayesian Learning for Block Sparse Recovery

The recovery of block-sparse signals with unknown structural patterns remains a fundamental challenge in structured sparse signal reconstruction. By proposing a variance transformation framework, this paper unifies existing pattern-based block sparse Bayesian learning methods, and introduces a novel space power prior based on undirected graph models to adaptively capture the unknown patterns of block-sparse signals. By combining the EM algorithm with high-order equation root-solving, we develop a new structured sparse Bayesian learning method, SPP-SBL, which effectively addresses the open problem of space coupling parameter estimation in pattern-based methods. We further demonstrate that learning the relative values of space coupling parameters is key to capturing unknown block-sparse patterns and improving recovery accuracy. Experiments validate that SPP-SBL successfully recovers various challenging structured sparse signals (e.g., chain-structured signals and multi-pattern sparse signals) and real-world multi-modal structured sparse signals (images, audio), showing significant advantages in recovery accuracy across multiple metrics.

math.OC

InViC: Intent-aware Visual Cues for Medical Visual Question Answering

Medical visual question answering (Med-VQA) aims to answer clinically relevant questions grounded in medical images. However, existing multimodal large language models (MLLMs) often exhibit shortcut answering, producing plausible responses by exploiting language priors or dataset biases while insufficiently attending to visual evidence. This behavior undermines clinical reliability, especially when subtle imaging findings are decisive. We propose a lightweight plug-in framework, termed Intent-aware Visual Cues (InViC), to explicitly enhance image-based answer generation in medical VQA. InViC introduces a Cue Tokens Extraction (CTE) module that distills dense visual tokens into a compact set of K question-conditioned cue tokens, which serve as structured visual intermediaries injected into the LLM decoder to promote intent-aligned visual evidence. To discourage bypassing of visual information, we further design a two-stage fine-tuning strategy with a cue-bottleneck attention mask. In Stage I, we employ an attention mask to block the LLM's direct view of raw visual features, thereby funneling all visual evidence through the cue pathway. In Stage II, standard causal attention is restored to train the LLM to jointly exploit the visual and cue tokens. We evaluate InViC on three public Med-VQA benchmarks (VQA-RAD, SLAKE, and ImageCLEF VQA-Med 2019) across multiple representative MLLMs. InViC consistently improves over zero-shot inference and standard LoRA fine-tuning, demonstrating that intent-aware visual cues with bottlenecked training is a practical and effective strategy for improving trustworthy Med-VQA.

cs.CV

Rethinking the Efficiency and Effectiveness of Reinforcement Learning for Radiology Report Generation

Radiologists highly desire fully automated AI for radiology report generation (R2G), yet existing approaches fall short in clinical utility. Reinforcement learning (RL) holds potential to address these shortcomings, but its adoption in this task remains underexplored. In this paper, we revisit RL in terms of data efficiency and optimization effectiveness for R2G tasks. First, we explore the impact of data quantity and quality on the performance of RL in medical contexts, revealing that data quality plays a more critical role than quantity. To this end, we propose a diagnostic diversity-based data sampling strategy that enables comparable performance with fewer samples. Second, we observe that the majority of tokens in radiology reports are template-like and diagnostically uninformative, whereas the low frequency of clinically critical tokens heightens the risk of being overlooked during optimization. To tackle this, we introduce Diagnostic Token-weighted Policy Optimization (DiTPO), which directly optimizes for clinical accuracy by using a diagnostic F1 score as the reward signal. Unlike standard RL approaches that treat all tokens equally, DiTPO explicitly models the varying importance of different tokens through rule- or gradient-based mechanisms to prioritize clinically relevant content. Extensive experiments on the MIMIC-CXR, IU-Xray, and CheXpert Plus datasets demonstrate that our framework achieves state-of-the-art (SOTA) performance while requiring substantially fewer training samples in RL. Notably, on MIMIC-CXR, our framework attains an F1 score of 0.516 using only 20% of the RL training samples.

cs.CV

Subgradient Gliding Method for Nonsmooth Convex Optimization

We identify and analyze a fundamental limitation of the classical projected subgradient method in nonsmooth convex optimization: the inevitable failure caused by the absence of valid subgradients at boundary points. We show that, under standard step sizes for both convex and strongly convex objectives, the method can fail after a single iteration with probability arbitrarily close to one, even on simple problem instances. To overcome this limitation, we propose a novel alternative termed the \textit{subgradient gliding method}, which remains well defined without boundary subgradients and avoids premature termination. Beyond resolving this foundational issue, the proposed framework encompasses the classical projected subgradient method as a special case and substantially enlarges its admissible step-size design space, providing greater flexibility for algorithmic design. We establish optimal ergodic convergence rates, $\mathcal{O}(1/\sqrt{t})$ for convex problems and $\mathcal{O}(1/t)$ for strongly convex problems, and further extend the framework to stochastic settings. Notably, our analysis does not rely on global Lipschitz continuity of the objective function, requiring only mild control on subgradient growth. Numerical experiments demonstrate that, in scenarios where the classical projected subgradient method fails completely, the proposed method converges reliably with a $100\%$ success rate and achieves orders-of-magnitude improvements in accuracy and convergence speed. These results substantially expand the scope of subgradient-based optimization methods to non-Lipschitz nonsmooth convex problems.

math.OC

UniVRSE: Unified Vision-conditioned Response Semantic Entropy for Hallucination Detection in Medical Vision-Language Models

Vision-language models (VLMs) have great potential for medical image understanding, particularly in Visual Report Generation (VRG) and Visual Question Answering (VQA), but they may generate hallucinated responses that contradict visual evidence, limiting clinical deployment. Although uncertainty-based hallucination detection methods are intuitive and effective, they are limited in medical VLMs. Specifically, Semantic Entropy (SE), effective in text-only LLMs, becomes less reliable in medical VLMs due to their overconfidence from strong language priors. To address this challenge, we propose UniVRSE, a Unified Vision-conditioned Response Semantic Entropy framework for hallucination detection in medical VLMs. UniVRSE strengthens visual guidance during uncertainty estimation by contrasting the semantic predictive distributions derived from an original image-text pair and a visually distorted counterpart, with higher entropy indicating hallucination risk. For VQA, UniVRSE works on the image-question pair, while for VRG, it decomposes the report into claims, generates verification questions, and applies vision-conditioned entropy estimation at the claim level. To evaluate hallucination detection, we propose a unified pipeline that generates responses on medical datasets and derives hallucination labels via factual consistency assessment. However, current evaluation methods rely on subjective criteria or modality-specific rules. To improve reliability, we introduce Alignment Ratio of Atomic Facts (ALFA), a novel method that quantifies fine-grained factual consistency. ALFA-derived labels provide ground truth for robust benchmarking. Experiments on six medical VQA/VRG datasets and three VLMs show UniVRSE significantly outperforms existing methods with strong cross-modal generalization.

cs.CV