arXiv ScienceSearch

arXiv subjects

Xin Fu

Publications and source records attributed to Xin Fu.

At least 19 recordsLinked to original sources

A Novel Semantic Manifold Alignment Attack against Embedding-to-Embedding Obfuscation in Privacy-Preserving LLMs

With the widespread applications of large language models (LLMs), privacy-preserving inference has become increasingly essential for sensitive queries. To balance privacy and utility, a series of lightweight obfuscation approaches has recently been proposed, where users locally transform plaintext embeddings into the fixed ciphertext ones. While such Embedding-to-Embedding Obfuscation (E2EO) schemes demonstrate considerable resilience against traditional token frequency and embedding inversion attacks, the core mechanism behind remains to be the large-scale one-to-one substitution, which provides no cryptographic guarantees. In this paper, we propose Proxy Manifold Alignment (PMA), a novel attack against E2EO in privacy-preserving LLMs. Our key observation is that E2EO schemes keep the original semantic structure, so that the obfuscated vector stream can be regarded as an unknown tokenizer-language whose symbols are the vectors themselves. Therefore, the proposed ciphertext to plaintext reconstruction attack can be formulated as a translation task from the unknown tokenizer-language to plaintext. Specifically, by only accessing the obfuscated vector stream, the target tokenizer and a public corpus, the PMA attack first employs Word2Vec to model the co-occurrence patterns within the obfuscated stream and the public corpus independently, and constructs two proxy vector embeddings. Then, the attack aligns the underlying manifolds of these two embeddings based on structural similarity. Finally, it maps the obfuscated vectors back to plaintext. Experimental results demonstrate that PMA consistently achieves higher plaintext recovery than other state-of-the-art attack methods.

cs.CR

On a conjecture of Demailly-Peternell-Schneider: the Kahler case

Let $f:(X,\Delta)\rightarrow Y$ be a surjective holomorphic map between two normal Kahler varieties, where (X,\Delta) is a log canonical pair, -(KX+\Delta) is nef and Y is Q-Gorenstein. In Fu-Guo-Song-Wang [18], it is proved that -KY is pseudo-effective if f is a projective morphism. In this paper, prove that -KY is pseudo-effective without assuming that f is projective

math.AG

Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation

RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challenging because retrieved context increases computation and memory. A direct way to reduce this cost is to retain only one retrieved chunk before generation, but the top-ranked retrieved chunk is not always the most evidence-supporting one, since retrieval similarity does not necessarily imply evidential sufficiency. Existing context-reduction methods can improve context quality, but often require additional LLMs or compressors that are costly under a strict mobile budget. In this paper, we study lightweight RAG chunk selection as an evidence-alignment problem. Our selector combines three complementary feature sources: question hidden states that represent LLM-side query intent, MoE routing-derived expert signals that capture the generator's internal routing structure, and retrieved chunk embeddings that preserve candidate-side evidence geometry. A compact multilayer perceptron maps these features to an evidence prototype in the chunk embedding space, and the candidate most aligned with this prototype is selected by cosine similarity. For stricter deployment budgets, we further introduce an optional task-aware feature selection strategy to reduce the selector input dimension. To support supervised evaluation, we construct semantic chunk-correctness labels based on evidence sufficiency rather than answer-string containment. Experiments show that the proposed selector consistently improves rank-1 evidence selection over mobile-applicable baselines by an average of 2.5%. These results suggest that using LLM-side query representations and MoE routing information and aligning them with retrieval-side candidate embedding is an effective and parameter-efficient strategy for mobile-applicable RAG chunk selection.

cs.LG

On the Datar-Mete-Song minimal slope conjecture

We prove a conjecture of Datar-Mete-Song \cite{DMS} characterizing $J$-slope semi-stability by the minimal $J$-slope. More precisely, for a semi-stable pair of K\"ahler classes $(\alpha,\beta)$ on a compact K\"ahler manifold $X$, every big and nef birational test class has slope at least the topological $J$-slope, whereas an unstable pair admits a test class with strictly smaller slope. We also introduce the $J$-null locus of a semi-stable pair and prove that it is an analytic subset of $X$ if $X$ is a compact K\"ahler surface or a compact toric K\"ahler manifold. In the toric invariant case, we show that Murakami's \cite{Murakami} weak solution to the $J$-equation is smooth and K\"ahler on the dense big torus $(\mathbb{C}^*)^n$ of $X$.

math.AG

MVMD: A Multi-View Approach for Enhanced Mirror Detection

In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models. As 3D reconstruction typically relies on multi-view images to capture different perspectives of a scene, detecting and labeling mirrors in multi-view images before reconstruction can effectively address this issue. However, existing methods focus solely on single-image detection, overlooking the rich information provided by multi-view setups. To overcome this limitation, we propose MVMD, a novel Multi-View Mirror Detection method, along with the first database specifically designed for mirror detection in multi-view scenes. The design of MVMD is grounded in the inherent associations between objects seen from different views and those reflected inside and outside of mirrors. These relationships are learned through cross- and self-attention mechanisms. MVMD consists of three key blocks: the Inter-Views Block tracks the shifts of objects within mirrors caused by changes in viewpoint; the Intra-View Block detects object reflections inside mirrors; and the Refinement Block sharpens mirror boundaries and enhances detected details. Experimental results show that our method improves accuracy by up to 2.6% and IoU by up to 11.1%, compared to single-image mirror detection techniques. This substantial improvement makes MVMD particularly effective for computer vision tasks, especially in enhancing the accuracy of 3D reconstruction in mirror-dense environments.

cs.CV

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot videos contain rich physical interactions but often lack executable robot action labels. We present WALA, a framework for learning executable latent actions from both action-labeled demonstrations and action-free videos. WALA first pretrains a semantic-geometric latent action model from videos by modeling the evolution between current observations and sparsely sampled future observations. Instead of reconstructing raw pixels, WALA predicts future deltas in the DINOv3 feature space and dense depth space, preserving task-relevant semantic and geometric structure while reducing sensitivity to appearance details. During policy training, the pretrained encoder provides stable latent action targets, and the decoder serves as a trainable latent world model. The latent actions generated by the vision-language backbone are jointly supervised by robot action prediction, latent action target matching, and future dynamics prediction. This enables action-labeled demonstrations to provide executable control supervision, while action-free videos contribute dynamics supervision without requiring robot action annotations. Experiments show that WALA achieves strong performance on RoboTwin, sets a new state-of-the-art result on RoboCasa with 75.2% average success, and improves both policy performance and generalization in real-world manipulation tasks.

cs.RO

Minnaert resonances and higher-order acoustic modes for bubbles in a viscous fluid with surface tension

The aim of this paper is to account for viscosity, surface tension, and interactions between micro-bubbles in approximating their resonant behavior in the ultrasonic regime. Original asymptotic formulas for the resonance frequencies are derived in terms of the difference in the acoustic impedance at the interface between the gas and the fluid. Both low-frequency resonances (Minnaert resonances) and higher-frequency resonances, i.e., beyond the subwavelength regime, are considered. We also provide a resonant characterization for a system of several micro-bubbles.

math.AP

DMT: Demographic Conditioning, Morphology-Enhanced Transformer for Cuffless Blood Pressure Estimation from PPG Signals

Blood pressure (BP) is a key marker for cardiovascular risk assessment and therapeutic decision-making, and Photoplethysmography (PPG) enables low-cost, wearable-friendly cuffless BP estimation. However, even with recent progress, many PPG-based models are trained with BP regression alone and may rely on amplitude-dominated shortcuts. In addition, demographic covariates that systematically modulate vascular compliance are often incorporated only via late fusion, limiting subject-specific representation learning. We propose a Transformer-based network for cuffless BP estimation from PPG signal, leveraging self-attention to capture long-range dependencies across multiple cardiac cycles. To account for subject-specific vascular differences, the model is conditioned on demographics via FiLM-style feature modulation applied through the attention and feed-forward sublayers of Transformer blocks. In addition, we add an auxiliary morphology head to guide the model to attend to BP-relevant waveform morphology associated with arterial stiffness and wave reflection. Under calibration-based evaluation protocols on the large-scale PulseDB dataset, the proposed method achieves MAE of 4.56 mmHg for systolic BP and 2.62 mmHg for diastolic BP, reducing errors by 47% and 50% compared with prior demographic-enhanced PPG baselines. The resulting lightweight, single-sensor model supports scalable and clinically grounded cuffless BP estimation in calibration-enabled deployment settings.

eess.SP

Nationwide EHR-Based Chronic Rhinosinusitis Prediction Using Demographic-Stratified Models

Chronic rhinosinusitis (CRS) is a common heterogeneous inflammatory disorder that causes substantial morbidity and healthcare costs. CRS is difficult to identify early from routine encounters, as symptom presentations overlap with common conditions such as allergic rhinitis, and heterogeneous phenotypes further obscure risk patterns. Prior predictive studies often rely on single-institutional cohorts , which reduce population-level generalizability. To overcome this, we leveraged nationwide longitudinal EHR data from the \textit{All of Us} Research Program to predict CRS diagnosis using two years of pre-diagnostic history. To address extreme feature sparsity and dimensionality in coded EHR data, we implemented a hybrid feature-selection pipeline that combines prevalence-based statistical screening with model-based importance ranking, compressing approximately 110,000 candidate codes into 100 interpretable features. To capture demographic heterogeneity, we trained demographic stratified models across six adult sex and life-stage subgroups with subgroup-specific hyperparameter tuning. Our framework achieved an overall AUC of 0.8461, improving discrimination by 0.0168 over the best baseline. These results demonstrate that routinely collected EHR data may support population-representative CRS risk stratification and inform earlier triage and referral prioritization in primary care.

cs.LG

Szczarba's twisted shuffle and equivariant path homology of directed graphs

To a marked simplicial set one can associate its path chain complex, and define its homology to be the homology of this complex, inspired by path homology theories for directed graphs, quivers, and marked categories. Given a marked simplicial set with a simplicial group action preserving the markings and degenerate 1-simplices, together with a twisting function, we define a marked twisted Cartesian product using the box product. Classically, Szczarba's twisted shuffle provides a quasi-isomorphism between the chain complex of a twisted Cartesian product and the corresponding twisted tensor product. In this paper, we prove that in the marked setting, this map restricts to a chain isomorphism on path chain complexes. As an application, for directed graphs with group actions, we obtain a natural Borel construction as a special case of marked twisted Cartesian products. Equivariant path homology is defined as the homology of this construction and is computed by an explicit twisted tensor product.

math.AT

Observability and Semiclassical Control for Schr\"odinger Equations on Non-compact Hyperbolic Surfaces

We study the observability of the Schr\"odinger equation on $X$, a non-compact covering space of a compact hyperbolic surface $M$. Using a generalized Bloch theory, functions on $X$ are identified as sections of flat Hilbert bundles over $M$. We develop a semiclassical analysis framework for such bundles and generalize the result of semiclassical control estimates in [Dyatlov and Jin, Acta Math., 220 (2018), pp. 297-339] to all flat Hilbert bundles over $M$, with uniform constants with respect to the choice of bundle. Furthermore, when the Riemannian cover $X \to M$ is a normal cover with a virtually Abelian deck transformation group $\Gamma$, we combine the uniform semiclassical control estimates on flat Hilbert bundles with the generalized Bloch theory to derive observability from any $\Gamma$-periodic open subsets of $X$. We also discuss applications of the uniform semiclassical control estimates in spectral geometry.

math.AP

Fundamental groups of compact Kahler varieties with nef anti canonical bundle

It is proved by M. Paun (1997, 2017) that the fundamental group of a compact Kahler manifold X is almost Abelian if the anti-canonical bundle -KX is nef. In this paper, we apply the recent geometric analytic theory of Kahler spaces developed by Guo-Phong-Song-Sturm to study fundamental groups of mildly singular compact Kahler varieties. We first extend Paun's result to log canonical pairs (X,Delta) with smooth X and nef -(KX+Delta) as well as to compact Kahler manifolds X with pseudo-effective -KX under a suitable assumption on the singularities of c1(-KX). We further prove that, for a 3-dimensional log canonical pair (X,\Delta) with X being klt, pi 1(X) is almost Abelian if -(KX+\Delta) is nef. Moreover, as one of the main ingredients for the proof of these results, we establish the surjectivity of the Albanese maps of compact normal complex varieties X in Fujiki class C that admits an effective R-divisor \Delta such that the pair (X,\Delta) is log canonical with nef anti-log canonical divisor -(KX+\Delta).This generalizes the corresponding theorems for projective varieties (Zhang, 2005), for klt pairs (Matsumura-Wang-Wu-Zhang, 2025) and for log smooth case (Fu-Han-Zou, 2025)

math.AG

Homogenization of the scattered wave and scattering resonances for periodic high-contrast subwavelength resonators

We study time-harmonic scattering by a periodic array of penetrable, high-contrast obstacles with small period, confined to a bounded Lipschitz domain. The strong contrast between the obstacles and the background induces subwavelength resonances. We derive a frequency-dependent effective model in the vanishing-period limit and prove quantitative convergence of the heterogeneous scattered wave to the effective scattered wave. We also identify the limiting set of scattering resonances and establish convergence rates. Finally, we establish convergence rates for the far-field pattern of the heterogeneous problem to that of the effective model.

math.AP

Fed MobiLLM: Efficient Federated LLM Fine-Tuning over Heterogeneous Mobile Devices via Server Assisted Side-Tuning

Collaboratively fine-tuning (FT) large language models (LLMs) over heterogeneous mobile devices fosters immense potential applications of personalized intelligence. However, such a vision faces critical system challenges. Conventional federated LLM FT approaches place prohibitive computational and memory burdens on mobile hardware, and their synchronous model aggregation protocols stall for slower devices. In this paper, we propose Fed MobiLLM, a novel design to facilitate efficient federated LLM FT across mobile devices with diverse computing/communication speeds and local model architectures. In particular, Fed MobiLLM implements a pioneering server-assisted federated side-tuning paradigm. Briefly, mobile devices perform lightweight forward propagation computations on local data using their frozen pre-scaled backbone LLMs, and then upload selected intermediate activations. The server trains a shared side-network independently, eliminating client-side backpropagation and enabling asynchronous updates. To bridge model heterogeneity across different devices, we introduce an adaptive layer-wise feature alignment method, which ensures consistent representations for collaboratively tuning a shared side network. Extensive experimental results demonstrate that Fed MobiLLM can maintain robust fine-tuning performance while achieving extremely low on-device memory, with at least 95.2% reduction in computation overhead, 93.2% reduction in communication costs and 5.1x faster convergence compared to existing methods, validating its efficacy for practical LLM adaptation over heterogeneous mobile devices.

cs.LG

Variation of Kahler-Einstein metrics with mixed singularities

In this short note, we consider a fiberation f: (X, Delta) to Y between two compact Kahler manifolds with generic fiber of f being a smooth log canonical pair with ample canonical divisor, we prove that the current induced by variation of Kahler Einsteins with mixed cone and Poincare singularities is positive, hence generalize the result of Schumacher in the smooth case [22] and the result of Guenancia in the conic case [14]. As application, we prove the surjectivity of Albanese map for a smooth log canonical pair with -(KX + Delta) being nef.

math.DG

PAE MobiLLM: Privacy-Aware and Efficient LLM Fine-Tuning on the Mobile Device via Additive Side-Tuning

There is a huge gap between numerous intriguing applications fostered by on-device large language model (LLM) fine-tuning (FT) from fresh mobile data and the limited resources of a mobile device. While existing server-assisted methods (e.g., split learning or side-tuning) may enable LLM FT on the local mobile device, they suffer from heavy communication burdens of activation transmissions, and may disclose data and labels to the server. To address those issues, we develop PAE MobiLLM, a a privacy-aware and efficient LLM FT method which can be deployed on the mobile device via server-assisted additive side-tuning. To further accelerate FT convergence and improve computing efficiency, PAE MobiLLM integrates activation caching on the server side, which allows the server to reuse historical activations and saves the mobile device from repeatedly computing forward passes for the recurring data samples. Besides, to reduce communication cost, PAE MobiLLM develops an activation shortcut that transmits only the token involved in the loss calculation instead of full activation matrices to guide the side network tuning. Last but not least, PAE MobiLLM introduces the additive adapter side-network design which makes the server train the adapter modules based on device-defined prediction differences rather than raw ground-truth labels. In this way, the server can only assist device-defined side-network computing, and learn nothing about data and labels. Extensive experimental results demonstrate PAE MobiLLM's superiority.

cs.LG

NeuroMoE: A Transformer-Based Mixture-of-Experts Framework for Multi-Modal Neurological Disorder Classification

The integration of multi-modal Magnetic Resonance Imaging (MRI) and clinical data holds great promise for enhancing the diagnosis of neurological disorders (NDs) in real-world clinical settings. Deep Learning (DL) has recently emerged as a powerful tool for extracting meaningful patterns from medical data to aid in diagnosis. However, existing DL approaches struggle to effectively leverage multi-modal MRI and clinical data, leading to suboptimal performance. To address this challenge, we utilize a unique, proprietary multi-modal clinical dataset curated for ND research. Based on this dataset, we propose a novel transformer-based Mixture-of-Experts (MoE) framework for ND classification, leveraging multiple MRI modalities-anatomical (aMRI), Diffusion Tensor Imaging (DTI), and functional (fMRI)-alongside clinical assessments. Our framework employs transformer encoders to capture spatial relationships within volumetric MRI data while utilizing modality-specific experts for targeted feature extraction. A gating mechanism with adaptive fusion dynamically integrates expert outputs, ensuring optimal predictive performance. Comprehensive experiments and comparisons with multiple baselines demonstrate that our multi-modal approach significantly enhances diagnostic accuracy, particularly in distinguishing overlapping disease states. Our framework achieves a validation accuracy of 82.47\%, outperforming baseline methods by over 10\%, highlighting its potential to improve ND diagnosis by applying multi-modal learning to real-world clinical data.

eess.IV