arXiv ScienceSearch

arXiv subjects

Yuanyuan Li

Publications and source records attributed to Yuanyuan Li.

At least 19 recordsLinked to original sources

Two-Stage Mixture-of-LoRA for Multi-Task Medical Vision-Language Learning

Medical vision-language models (VLMs) allow a single model to perform clinical image analysis tasks ranging from diagnosis classification to report generation. However, joint adaptation is challenged by heterogeneous output formats, conflicting task gradients, and imbalanced training data. Hence, we present \textbf{Two-Stage Mixture-of-LoRA}, a framework built on MedGemma-1.5-4B. The framework uses a shared-specific Mixture-of-LoRA architecture comprising one shared LoRA and six task-specific expert LoRAs, together with a two-stage training procedure. In Stage 1, we jointly train the shared LoRA and all task-specific expert LoRAs on all tasks. In Stage 2, we first freeze the backbone, the shared LoRA, and all non-target experts, and refine one task expert at a time. Classification and regression then receive an additional modality-balanced continuation, in which smaller modality groups are repeated to match the largest group. In the FLARE 2026 Task 3 test sets, the proposed method achieves 0.85 balanced accuracy for classification, 0.48 micro-F1 for multi-label classification, 0.79 detection F1, and 17.39 regression MAE. Code is available at https://github.com/YuanYL03/MICCAI-FLARE-2026-Challenge-Task3-2D.

cs.CV

Global $W^{2,p}$ Regularity in Optimal Transport

In this paper we establish global $W^{2,p}$ estimates for the convex potentials of quadratic optimal transport between bounded convex domains with continuous positive densities. All the assumptions are optimal. The main new ideas include a blow-up analysis that allows for lower-dimensional collapse of the limiting source measure and reduces the limiting problem to a transport problem on its affine hull, and a good-bad scale decomposition in which rigidity controls the good scales while a counting argument shows that the proportion of bad scales tends to zero.

math.AP

BifrostUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation

High-quality demonstration data are essential for humanoid robot skill learning, especially for whole-body behaviors that require coordinated perception, locomotion, and manipulation. Existing data-collection methods largely rely on robot teleoperation, which is constrained by hardware accessibility, operator expertise, and limited efficiency. Inspired by the Universal Manipulation Interface (UMI), we propose BifrostUMI, a portable and robot-free framework for humanoid whole-body data collection. BifrostUMI uses lightweight VR devices and UMI-inspired grippers to collect sparse human keypoint trajectories, wrist-view observations, and gripper actions. These demonstrations train a high-level policy to predict future keypoints, which are retargeted to robot-native whole-body references and executed by a whole-body controller. Experiments in five real-world scenarios demonstrate the effectiveness of the proposed framework and validate the collected demonstrations for transferable humanoid whole-body skill learning.

cs.RO

HumanoidUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation

High-quality demonstration data are essential for humanoid robot skill learning, especially for whole-body behaviors that require coordinated perception, locomotion, and manipulation. Existing data-collection methods largely rely on robot teleoperation, which is constrained by hardware accessibility, operator expertise, and limited efficiency. Inspired by the Universal Manipulation Interface (UMI), we propose HumanoidUMI, a portable and robot-free framework for humanoid whole-body data collection. HumanoidUMI uses lightweight VR devices and UMI-inspired grippers to collect sparse human keypoint trajectories, wrist-view observations, and gripper actions. These demonstrations train a high-level policy to predict future keypoints, which are retargeted to robot-native whole-body references and executed by a whole-body controller. Experiments in five real-world scenarios demonstrate the effectiveness of the proposed framework and validate the collected demonstrations for transferable humanoid whole-body skill learning.

cs.RO

The many-body Blaschke-Santaló type inequality via optimal transport

Let $K_1,\ldots,K_k\subset\mathbb R^n$ be origin-symmetric measurable sets of finite volume such that \[ \sum_{1\le i<j\le k}\langle x_i,x_j\rangle\le \binom{k}{2}, \qquad \forall\,x_i\in K_i, x_j\in K_j. \] We prove the sharp many-body Blaschke--Santaló type inequality \[ \prod_{i=1}^k |K_i|\le |B^n|^k \] proposed by Kalantzopoulos and Saroglou, and characterize all equality cases. The proof combines multi-marginal optimal transport with a pseudo-Euclidean volume estimate. Using the geometric--functional equivalence of Kalantzopoulos and Saroglou, we also establish the functional version inequality proposed by Kolesnikov and Werner.

math.AP

The Mahler Conjecture in Three Dimensions

The Mahler conjecture dates back to 1938. This paper solves the conjecture for general convex bodies in three dimensions by developing a method called the shadow flow. The equality case is characterized as well. This method is also applied to give a new proof of the three-dimensional symmetric case, which was first proved by Iriyeh--Shibata.

math.MG

Domain-Shift-Aware Conformal Prediction for Large Language Models

Large language models have achieved impressive performance across diverse tasks. However, their tendency to produce overconfident and factually incorrect outputs, known as hallucinations, poses risks in real-world applications. Conformal prediction provides finite-sample, distribution-free coverage guarantees, but standard conformal prediction breaks down under domain shift, often leading to under-coverage and unreliable prediction sets. We propose a new framework called Domain-Shift-Aware Conformal Prediction (DS-CP). Our framework adapts conformal prediction to large language models under domain shift, by systematically reweighting calibration samples based on their proximity to the test prompt, thereby preserving validity while enhancing adaptivity. Our theoretical analysis and experiments on the MMLU benchmark demonstrate that the proposed method delivers more reliable coverage than standard conformal prediction, especially under substantial distribution shifts, while maintaining efficiency. This provides a practical step toward trustworthy uncertainty quantification for large language models in real-world deployment.

stat.ML

Optimal (partial) transport to non-convex polygonal domains

In this paper, we investigate optimal (partial) transport problems for which the target is a non-convex polygonal domain in \(\mathbb{R}^2\). For the complete optimal transport problem, we prove that the singular set is locally a smooth one-dimensional curve away from finitely many points. For the optimal partial transport problem, we prove that the free boundary is smooth away from finitely many singular points. In higher dimensions, we formulate two conjectures concerning the structure of singularities when the target is a non-convex polytope.

math.AP

The Symmetric Mahler Inequality in Dimension Three via Admissible Shadow Systems

The three-dimensional symmetric Mahler inequality states that, for every origin-symmetric convex body \(K=-K\subset \mathbb{R}^3\), \[ \VP(K)= |K|\,|K^\circ|\geq \frac{32}{3}. \] It was recently proved by Iriyeh--Shibata \cite{IS2020}, and a shorter proof was later given by Fradelizi--Hubard--Meyer--Roldán-Pensado--Zvavitch \cite{FHMRZ}. Both proofs combine ingenious equipartition arguments of algebraic-topological origin with delicate geometric estimates inspired by Meyer's argument for unconditional bodies. In this paper, we give a new proof of this inequality using a purely geometric approach, based on what we call symmetric admissible shadow systems. This is a natural extension of the new techniques developed in our proof of the three-dimensional non-symmetric Mahler conjecture \cite{CLXX-Mahler}.

math.MG

OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction

UMI-style interfaces enable scalable robot learning, but existing systems remain largely visuomotor, relying primarily on RGB observations and trajectory while providing only limited access to physical interaction signals. This becomes a fundamental limitation in contact-rich manipulation, where success depends on contact dynamics such as tactile interaction, internal grasping force, and external interaction wrench that are difficult to infer from vision alone. We present OmniUMI, a unified framework for physically grounded robot learning via human-aligned multimodal interaction. OmniUMI synchronously captures RGB, depth, trajectory, tactile sensing, internal grasping force, and external interaction wrench within a compact handheld system, while maintaining collection--deployment consistency through a shared embodiment design. To support human-aligned demonstration, OmniUMI enables natural perception and modulation of internal grasping force, external interaction wrench, and tactile interaction through bilateral gripper feedback and the handheld embodiment. Built on this interface, we extend diffusion policy with visual, tactile, and force-related observations, and deploy the learned policy through impedance-based execution for unified regulation of motion and contact behavior. Experiments demonstrate reliable sensing and strong downstream performance on force-sensitive pick-and-place, interactive surface erasing, and tactile-informed selective release. Overall, OmniUMI combines physically grounded multimodal data acquisition with human-aligned interaction, providing a scalable foundation for learning contact-rich manipulation.

cs.RO

On the monotonicity of affine quermassintegrals

Lutwak's affine quermassintegral theory is a foundational component of modern affine Brunn--Minkowski theory. Developed in the 1980s, it provides affine analogues of the classical quermassintegrals and has led to a rich family of sharp affine isoperimetric inequalities. A central question in this program, going back to Lutwak's 1988 work, is an Alexandrov--Fenchel-type monotonicity principle for the normalized $L^{-n}$-moment quermassintegrals $I_{k,-n}$. In one form, this principle predicts that \[ I_{m,-n}(K)^{1/m}\ge I_{k,-n}(K)^{1/k}, \qquad 1\le m (m+2)(k+2)-2$, there exists an origin-symmetric $C^2_+$ convex body $K\subset\mathbb R^n$ such that \[ I_{m,-n}(K)^{1/m} < I_{k,-n}(K)^{1/k}. \] The example is obtained from the Euclidean ball by an arbitrarily small degree-four spherical harmonic perturbation. On the positive side, we prove that the endpoint chain is true in dimension three: for every convex body $K\subset\mathbb R^3$, \[ I_{1,-3}(K)\ge I_{2,-3}(K)^{1/2}\ge I_{3,-3}(K)^{1/3}=1. \] The equality cases in both non-trivial inequalities are exactly ellipsoids, up to translation and nonsingular affine transformations.

math.AP

Uniqueness of Blow-ups for the Superconductivity Free Boundary Problem

We study the free-boundary equation \[ Δu=χ_{\{|\nabla u|>0\}} \] near the origin. We prove that, at a singular point of \(\partial\{|\nabla u|>0\}\), the quadratic blow-up is unique. As noted in \cite[Notes to Chapter 7]{PSU2012}, little is known about the singular set for this problem. The usual Weiss--Monneau monotonicity argument does not seem to apply directly, because the inactive set is determined by the vanishing of \(\nabla u\), rather than by a sign condition on \(u\). The proof follows the quadratic part of the rescalings. Projecting onto the trace-free quadratic harmonics yields a finite-dimensional differential equation for the quadratic coefficient. Together with a Lyapunov identity and estimates on dyadic annuli, this implies convergence of the quadratic coefficient, and hence uniqueness of the blow-up.

math.AP

MOGeo: Beyond One-to-One Cross-View Object Geo-localization

Cross-View Object Geo-Localization (CVOGL) aims to locate an object of interest in a query image within a corresponding satellite image. Existing methods typically assume that the query image contains only a single object, which does not align with the complex, multi-object geo-localization requirements in real-world applications, making them unsuitable for practical scenarios. To bridge the gap between the realistic setting and existing task, we propose a new task, called Cross-View Multi-Object Geo-Localization (CVMOGL). To advance the CVMOGL task, we first construct a benchmark, CMLocation, which includes two datasets: CMLocation-V1 and CMLocation-V2. Furthermore, we propose a novel cross-view multi-object geo-localization method, MOGeo, and benchmark it against existing state-of-the-art methods. Extensive experiments are conducted under various application scenarios to validate the effectiveness of our method. The results demonstrate that cross-view object geo-localization in the more realistic setting remains a challenging problem, encouraging further research in this area.

cs.CV

Counterfactually Fair Conformal Prediction

While counterfactual fairness of point predictors is well studied, its extension to prediction sets--central to fair decision-making under uncertainty--remains underexplored. On the other hand, conformal prediction (CP) provides efficient, distribution-free, finite-sample valid prediction sets, yet does not ensure counterfactual fairness. We close this gap by developing Counterfactually Fair Conformal Prediction (CF-CP) that produces counterfactually fair prediction sets. Through symmetrization of conformity scores across protected-attribute interventions, we prove that CF-CP results in counterfactually fair prediction sets while maintaining the marginal coverage property. Furthermore, we empirically demonstrate that on both synthetic and real datasets, across regression and classification tasks, CF-CP achieves the desired counterfactual fairness and meets the target coverage rate with minimal increase in prediction set size. CF-CP offers a simple, training-free route to counterfactually fair uncertainty quantification.

cs.LG

Efficient Graph Knowledge Distillation from GNNs to Kolmogorov--Arnold Networks via Self-Attention Dynamic Sampling

Recent success of graph neural networks (GNNs) in modeling complex graph-structured data has fueled interest in deploying them on resource-constrained edge devices. However, their substantial computational and memory demands present ongoing challenges. Knowledge distillation (KD) from GNNs to MLPs offers a lightweight alternative, but MLPs remain limited by fixed activations and the absence of neighborhood aggregation, constraining distilled performance. To tackle these intertwined limitations, we propose SA-DSD, a novel self-attention-guided dynamic sampling distillation framework. To the best of our knowledge, this is the first work to employ an enhanced Kolmogorov-Arnold Network (KAN) as the student model. We improve Fourier KAN (FR-KAN+) with learnable frequency bases, phase shifts, and optimized algorithms, substantially improving nonlinear fitting capability over MLPs while preserving low computational complexity. To explicitly compensate for the absence of neighborhood aggregation that is inherent to both MLPs and KAN-based students, SA-DSD leverages a self-attention mechanism to dynamically identify influential nodes, construct adaptive sampling probability matrices, and enforce teacher-student prediction consistency. Extensive experiments on six real world datasets demonstrate that, under inductive and most of transductive settings, SA-DSD surpasses three GNN teachers by 3.05%-3.62% and improves FR-KAN+ by 15.61%. Moreover, it achieves a 16.69x parameter reduction and a 55.75% decrease in average runtime per epoch compared to key benchmarks.

cs.LG

Safer Prompts: Reducing Risks from Memorization in Visual Generative AI

Visual Generative AI models have demonstrated remarkable capability in generating high-quality images from user inputs like text prompts. However, because these models have billions of parameters, they risk memorizing certain parts of the training data and reproducing the memorized content. Memorization often raises concerns about safety of such models -- usually involving intellectual property (IP) infringement risk -- and deters their large scale adoption. In this paper, we evaluate the effectiveness of prompt engineering techniques in reducing memorization risk in image generation. Our findings demonstrate the effectiveness of prompt engineering in reducing the similarity between generated images and the training data of diffusion models, while maintaining relevance and aestheticity of the generated output.

cs.CV

Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting

Large scale text-to-image generation models can memorize and reproduce their training dataset. Since the training dataset often contains copyrighted material, reproduction of training dataset poses a copyright infringement risk, which could result in legal liabilities and financial losses for both the AI user and the developer. The current works explores the potential of chain-of-thought and task instruction prompting in reducing copyrighted content generation. To this end, we present a formulation that combines these two techniques with two other copyright mitigation strategies: a) negative prompting, and b) prompt re-writing. We study the generated images in terms their similarity to a copyrighted image and their relevance of the user input. We present numerical experiments on a variety of models and provide insights on the effectiveness of the aforementioned techniques for varying model complexity.

cs.LG