arXiv Science⌕ Search

arXiv subjects

Jun Luo

Publications and source records attributed to Jun Luo.

At least 37 records · Page 2Linked to original sources

Conflict-Aware Federated Fine-Tuning of Large Language Models with Mixture-of-Experts

The continuous scaling of large language models (LLMs) incurs prohibitive computational costs, making Mixture-of-Experts (MoE) a scalable alternative for efficient fine-tuning via sparse activation. While federated learning (FL) emerges as the paradigm for privacy-preserving collaborative optimization, integrating MoE into FL under data heterogeneity may trigger conflicting expert optimizations. Client-specific data distributions force same-indexed experts to optimize under inconsistent or even conflicting feature-label correlations. This mismatch induces destructive interference during aggregation, thus destabilizing the optimization trajectory and degrading model performance. To address this issue, we propose FC-MoE, a federated conflict-aware framework for MoE fine-tuning. It employs an importance aware weighting scheme to prioritize reliable local updates and utilizes gradient consensus projection to suppress conflicting updates, ensuring a stable global optimization path. Moreover, a local knowledge retention mechanism further preserves specialized client expertise by re-anchoring domain-specific residuals. Extensive experiments demonstrate that FC-MoE accelerates convergence and enhances both global and local model performance in non-IID federated environments.

cs.LG↗

Ultralow shot noise limited giant passive resonant gyroscope for Earth rotation measurement

Optical gyroscopes directly measure the Earth's rotation and are promising instruments for real-time geophysical observations and Earth orientation parameter (EOP) determination requiring both high precision and high temporal resolution. Large-scale ring laser gyroscopes (RLGs) currently reach rotational resolutions around $10^{-11}\,\mathrm{(rad/s)/\sqrt{Hz}}$, but their quantum noise limits make it challenging to meet the requirements of future high-temporal-resolution EOP measurements. Passive resonant gyroscopes (PRGs), on the other hand, offer a potentially lower photon shot noise limit and more flexible power scaling, even if their demonstrated rotational resolutions are still about two orders of magnitude below those of leading RLGs. Here we demonstrate a $64\,\mathrm{m^{2}}$ giant passive resonant gyroscope HUST-2, and develop with an extremely low shot noise level. We experimentally obtain a shot noise limited of $5.7(1)\times10^{-13}\,\mathrm{(rad/s)/\sqrt{Hz}}$ at $1\,\mathrm{mW}$ incident optical power, following the characteristic $1/\sqrt{P}$ scaling. Through systematic suppression of dominant technical noise sources, HUST-2 further achieves a measured rotational resolution of $3\times10^{-11}\,\mathrm{(rad/s)/\sqrt{Hz}}$, bringing PRGs into the performance regime of leading large-scale RLGs for the first time. The gap between the present demonstrated rotational resolution and the shot noise limit indicates nearly two orders of magnitude further improvement potential. Reaching this limit would enable high-precision length-of-day (LOD) measurements with $10$-$100\,\mathrm{s}$ temporal resolution and lays the foundation for future large-scale gyroscope networks dedicated to real-time EOP determination.

physics.optics↗

Cross-Domain Multi-Person Human Activity Recognition via Near-Field Wi-Fi Sensing

Wi-Fi-based human activity recognition (HAR) provides substantial convenience and has emerged as a thriving research field, yet the coarse spatial resolution inherent to Wi-Fi significantly hinders its ability to distinguish multiple subjects. By exploiting the near-field domination effect, establishing a dedicated sensing link for each subject through their personal Wi-Fi device offers a promising solution for multi-person HAR under native traffic. However, due to the subject-specific characteristics and irregular patterns of near-field signals, HAR neural network models require fine-tuning (FT) for cross-domain adaptation, which becomes particularly challenging with certain categories unavailable. In this paper, we propose WiAnchor, a novel training framework for efficient cross-domain adaptation in the presence of incomplete activity categories. This framework processes Wi-Fi signals embedded with irregular time information in three steps: during pre-training, we enlarge inter-class feature margins to enhance the separability of activities; in the FT stage, we innovate an anchor matching mechanism for cross-domain adaptation, filtering subject-specific interference informed by incomplete activity categories, rather than attempting to extract complete features from them; finally, the recognition of input samples is further improved based on their feature-level similarity with anchors. We construct a comprehensive dataset to thoroughly evaluate WiAnchor, achieving over 90% cross-domain accuracy with absent activity categories.

eess.SP↗

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data

Recent advances in language models have established reinforcement learning as the primary paradigm for eliciting self-correction and long-chain reasoning. While group relative policy optimization (GRPO) offers superior scalability by eliminating the critic network, deploying it on a central infrastructure entails collecting a large volume of data from distributed owners, which poses significant privacy risks. To address these concerns, we introduce federated GRPO (FGRPO), a framework designed to decentralize the fine-tuning of reasoning models across heterogeneous data owners. To effectively mitigate the instability caused by divergent reward scales across heterogeneous tasks, FGRPO incorporates an adaptive aggregation mechanism based on relative performance gain. By characterizing each client's improvement relative to its personalized historical baseline, the framework dynamically prioritizes effective learning trajectories regardless of local task difficulty. FGRPO ensures robust convergence on non-IID data while preserving data privacy.

cs.LG↗

DECA: Decentralizing Block-Wise Adam for Efficient LLM Full-Parameter Fine-Tuning on Non-IID Data

Fine-tuning large language models (LLMs) in privacy-sensitive and resource-constrained environments remains challenging. Since training data are often distributed across multiple clients, decentralized fine-tuning offers a natural paradigm for collaborative adaptation without a central server. However, enabling full-parameter fine-tuning (FPFT) in this decentralized setting is difficult: FPFT provides strong adaptation capacity but incurs prohibitive resource consumption for billion-scale models. Existing decentralized LLM fine-tuning methods therefore mainly rely on parameter-efficient updates, which improve efficiency but may restrict downstream performance. Moreover, client data are typically non-IID, making decentralized optimization more vulnerable to client drift and unstable convergence. To address these challenges, we propose DECA, a resource-efficient decentralized FPFT framework for LLMs on non-IID data. DECA partitions model parameters into disjoint blocks and performs sequential block-wise Adam optimization, reducing resource consumption while preserving decentralized full-parameter adaptation. To stabilize training, DECA further introduces first- and second-order block-wise moment estimates with fresh local gradient statistics and consensus-derived discrepancy signals. We provide rigorous theoretical analysis and extensive experiments, showing that DECA achieves fast convergence, strong downstream performance, and significant resource efficiency.

cs.LG↗

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness

Hallucination remains one of the key challenges undermining the reliability of Large Vision-Language Models (LVLMs). But what makes an LVLM hallucinate less? Many existing efforts focus on improving internal components of the model. We argue that hallucination fundamentally stems from how the model architecture is designed. To investigate this, we factor the architecture design into three dimensions: Linguistic Foundation (LF), Visual Representation (VR), and Semantic Alignment (SA), and categorize hallucinations into Co-occurrence, Similarity, and previously overlooked Uncertainty types. Building on this formulation, we propose CoSimUE, a benchmark that creates fine-grained hallucination scenarios through controlled textual perturbations and random perturbations, enabling mapping between design choices and hallucination behaviors. Experiments across 7 design aspects show that: 1) the widely emphasized scaling of model parameters has only limited impact on reducing all three types of hallucinations; 2) larger and better-trained language foundations can reduce co-occurrence hallucinations; 3) stronger visual encoders and higher resolutions mitigate similarity errors; 4) effective alignment strategies alleviate uncertainty hallucinations. 5) Furthermore, cross-dimensional analysis reveals that jointly enhancing visual fidelity and alignment quality yields the most comprehensive improvements. This study provides the first systematic exploration linking architecture-level design to hallucination robustness, offering practical guidance for developing reliable and efficient LVLMs.

cs.CV↗

Fundamental Physics and Cosmology with TianQin

The exploration of the surrounding world and the universe is an important theme in the legacy of humankind. The detection of gravitational waves is adding a new dimension to this grand effort. What are the fundamental physical laws governing the dynamics of the universe? What is the fundamental composition of the universe? How has the universe evolved in the past and how will it evolve in the future? These are the basic questions that press for answers. The space-based gravitational wave detector TianQin will tune in to gravitational waves in the millihertz frequency range ($10^{-4} \sim 1$ Hz, to be specific), opening a new gravitational wave spectrum window to explore many of the previously hidden sectors of the universe. TianQin will discover many astrophysical systems, populating the universe at different redshifts: some will be of new types that have never been detected before, some will have very high signal-to-noise ratios, and some will have very high parameter estimation precision. The plethora of information collected will bring us to new fronts on which to search for the breaking points of general relativity, the possible violation of established physical laws, the signature of possible new gravitational physics and new fundamental fields, and to improve our knowledge on the expansion history of the universe. In this white paper, we highlight the advances that TianQin can bring to fundamental physics and cosmology.

gr-qc↗

Design and Characterization of Racetrack 3D-Trench Silicon Sensor Based on 8-Inch Process with Excellent Time Resolution

In the extreme environments of high-luminosity colliders, traditional planar silicon sensors suffer severe radiation-induced performance degradation and fail to satisfy the stringent demands of high-precision tracking and high-speed timing in particle physics. 3D silicon sensors enhance radiation hardness by shortening charge collection distance, yet conventional designs with columnar or square-cell trench electrodes exhibit non-uniform electric fields, including saddle points and low-field regions, which degrade charge collection efficiency and timing resolution. This work presents a novel racetrack 3D-trench silicon sensor with continuous racetrack electrodes surrounding a long central collection electrode, aiming to eliminate electric field inhomogeneities. For the first time, a 23 $μ$m shallow-etched device was fabricated on an 8-inch platform, which provides a promising basis for its subsequent mass production and engineering applications. The device performance was systematically evaluated through theoretical analysis, 3D TCAD simulations, and characterization using semiconductor parameter analyzers and transient current technique (TCT) measurements. The sensor achieves leakage current below 0.2 nA, breakdown voltage above 110 V, full depletion voltage as low as a few volts, capacitance as low as 650 fF, collected charge of 4 fC, time response of about 640 ps, and time resolution of 50 ps. This large-scale manufacturable, shallow-etched racetrack 3D-trench silicon sensor provides a competitive device solution for portable radiation detection and next-generation 4D tracking under high-radiation and high-event-rate conditions.

physics.ins-det↗

Si/SiGe multi-channel superlattice structure epitaxial growth with segmented temperature control for Next-Generation Logic Devices

Stacking multiple SiSiGe channels in advanced logic devices faces severe thermal budget accumulation, which degrades interfaces via Ge-Si interdiffusion and strain relaxation.This strategy lowers the Ge diffusion coefficient to 5.6-7% of its value at 650C (Arrhenius estimate), suppressing interdiffusion and preserving pseudomorphic strain. The 4 + 4 channel stack exhibits clear XRD satellite peaks, fully coherent strain state (reciprocal space mapping), sharp interfaces (1.5-2.6 nm transition width) and low RMS roughness (0.08 nm). Quantitative analysis from bottom to top reveals that prolonged high-temperature exposure broadens bottom interfaces and dilutes Ge concentration (from 20% to 18.5%), while the top stack maintains design targets. This work provides a process-physics understanding of thermal budget effects in multi-channel superlattices and establishes a high-quality material foundation for advanced logic devices beyond 2 nm node.

cond-mat.mtrl-sci↗

FluxShard: Motion-Aware Feature Cache Reuse for Collaborative Video Analytics in Mobile Edge Computing

Caching and reusing intermediate features across consecutive frames is a common technique to reduce redundant computation and transmission for edge-cloud video analytics in mobile edge computation. Existing methods manage the cache in a fixed or globally shifted coordinate system, treating it as an indivisible whole. Under the non-uniform motion patterns of mobile scenes, this whole-scene granularity invalidates large portions of the cache even when most content has merely shifted spatially, wasting computation and bandwidth. The root cause is a granularity mismatch: the cache is managed per scene, yet motion varies per region. In this paper, we present FluxShard, a motion-aware edge-cloud video analytics system that uses codec-level block motion vectors (MVs) to manage feature cache reuse and recomputation at the granularity of individual motion regions. By re-indexing cached features along per-block MVs, FluxShard separates spatial displacement from content changes, recovering reusable content that whole-scene methods would otherwise discard. To ensure correct reuse under heterogeneous motion, the Receptive Field Alignment Principle (RFAP) identifies, from the input-level MV field alone, the positions that must be recomputed due to inconsistent spatial composition within receptive fields. To maintain cache coherence across frames, MV-guided cache remapping warps the entire feature cache to the current coordinate system each frame, sustaining a high reuse ratio over time. A profiling-driven dispatcher routes the remaining sparse workload between edge and cloud for lower latency. Evaluation across multiple vision tasks, dynamic video benchmarks, and network conditions shows that FluxShard reduces latency by 32.6-83.8% and energy by 14.9-64.0% over all baselines under the prescribed accuracy budget.

cs.NI↗

To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions. Among different methods, inference-time alignment is often cheaper as it intervenes (i.e., offers guidances) only during output generation. Existing proposals apply guidances extracted from certain aligned models without properly assessing their reliability. Nonetheless, our systematic evaluation reveals that guidance effectiveness varies drastically across models; since ineffective guidances lead to further confusion and thus further interventions, the resulting excessive interventions typically indicate poor performance. To make interventions more effective and thus more efficient, we introduce BlendIn, an inference-time alignment framework that shifts from binary decisions to creating hybrid distributions integrating both models' knowledge. BlendIn stabilizes inference-time alignment by performing quality-aware alignment and proportionally weighting each model's contribution based on reliability. Compared with existing works, it preserves beneficial guidance while downweighting unreliable suggestions. BlendIn provides both diagnostic signals and mitigation strategies for misaligned guidance, achieving consistent and up to 50% performance improvement on challenging model pairs. Our code is available at: https://github.com/DecayingSeart/BlendIn.

cs.LG↗

Bulk superconductivity up to 96 K in pressurized nickelate single crystals

Recently, the Ruddlesden-Popper bilayer nickelate $La_3Ni_2O_7$ has emerged as a superconductor with a transition temperature ($T_c$) of approximately 80 K above 14 GPa (Refs. 1-3). Achieving higher $T_c$ in nickelate superconductors, along with the synthesis of reproducible high-quality single crystals without relying on high-oxygen-pressure growth conditions, remains a significant challenge$^{[4-7]}$. Here we report superconductivity up to 96 K under high pressure in bilayer nickelate single crystals synthesized at ambient pressure. Energy-dispersive spectroscopy, single-crystal X-ray diffraction, nuclear quadrupole resonance and scanning transmission electron microscopy evidenced high crystal quality of the flux-grown $La_2SmNi_2O_{7-δ}$ single crystals. $La_2SmNi_2O_7$ exhibits clear bulk superconductivity, including zero resistivity ($T_{c,max}^{onset}$ = 92 K and $T_{c,max}^{zero}$ = 73 K at 21.6 GPa) and the Meissner effect ($T_c$= 60 K at 20.6 GPa). A low-temperature high-pressure structural study indicates that both monoclinic and tetragonal structures can support superconductivity in this bilayer nickelate. Furthermore, we established a correlation between higher $T_c$ under high pressures and larger in-plane lattice distortion under ambient conditions, corroborated by observing even higher $T_c^{onset}$ of 96 K in $La_{1.57}Sm_{1.43}Ni_2O_{7-δ}$. This study overcomes key limitations in growing nickelate superconductor crystals, resolves the crystal structure in the superconducting state and demonstrates an effective pathway towards achieving higher $T_c$.

cond-mat.supr-con↗

Physically-Induced Atmospheric Adversarial Perturbations: Enhancing Transferability and Robustness in Remote Sensing Image Classification

Adversarial attacks pose a severe threat to the reliability of deep learning models in remote sensing (RS) image classification. Most existing methods rely on direct pixel-wise perturbations, failing to exploit the inherent atmospheric characteristics of RS imagery or survive real-world image degradations. In this paper, we propose FogFool, a physically plausible adversarial framework that generates fog-based perturbations by iteratively optimizing atmospheric patterns based on Perlin noise. By modeling fog formations with natural, irregular structures, FogFool generates adversarial examples that are not only visually consistent with authentic RS scenes but also deceptive. By leveraging the spatial coherence and mid-to-low-frequency nature of atmospheric phenomena, FogFool embeds adversarial information into structural features shared across diverse architectures. Extensive experiments on two benchmark RS datasets demonstrate that FogFool achieves superior performance: not only does it exceed in white-box settings, but also exhibits exceptional black-box transferability (reaching 83.74% TASR) and robustness against common preprocessing-based defenses such as JPEG compression and filtering. Detailed analyses, including confusion matrices and Class Activation Map (CAM) visualizations, reveal that our atmospheric-driven perturbations induce a universal shift in model attention. These results indicate that FogFool represents a practical, stealthy, and highly persistent threat to RS classification systems, providing a robust benchmark for evaluating model reliability in complex environments.

cs.CV↗

Atoms of Compacta on Closed Surfaces

For any compact set $K$ lying on a closed surface $\mathcal{S}$ we introduce a closed equivalence relation $\sim$, called the {\em Schönflies equivalence} on $K$. We show that every class $[x]_\sim$ of $\sim$ is a continuum and that the resulting quotient space $K\!/\!\sim$ is a {\em Peano compactum}. By definition, all components of a Peano compactum are locally connected and for any $\varepsilon>0$ only finitely many of them have diameter greater than $\varepsilon$. The decomposition $\mathcal{D}_K=\{[x]_\sim: x\in K\}$ refines every other upper semicontinuous decomposition of $K$ into subcontinua that has a Peano compactum as its quotient space. In other words, $\mathcal{D}_K$ is the {\em core decomposition of $K$} with Peano quotient. The elements of $\mathcal{D}_K$ are called {\em atoms} of $K$. We also show that for any branched covering $f: \mathcal{S}^*\rightarrow \mathcal{S}$ from a closed surface $\mathcal{S}^*$ to $\mathcal{S}$, every atom of $f^{-1}(K)$ is sent into an atom of $K$. If $f$ is even a covering, it sends every atom of $f^{-1}(K)$ onto an atom of $K$. We illustrate our theory with examples and show that it cannot be generalized to $n$-manifolds with $n\ge 3$ by providing a detailed counterexample in~$\mathbb{R}^3$.

math.GN↗

Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents

Large Vision-Language Models (LVLMs) empower autonomous mobile agents, yet their security under realistic mobile deployment constraints remains underexplored. While agents are vulnerable to visual prompt injections, stealthily executing such attacks without requiring system-level privileges remains challenging, as existing methods rely on persistent visual manipulations that are noticeable to users. We uncover a consistent discrepancy between human and agent interactions: automated agents generate near-zero contact touch signals. Building on this insight, we propose a new attack paradigm, agent-only perceptual injection, where malicious content is exposed only during agent interactions, while remaining not readily perceived by human users. To accommodate mobile UI constraints and one-shot interaction settings, we introduce HG-IDA*, an efficient one-shot optimization method for constructing jailbreak prompts that evade LVLM safety filters. Experiments demonstrate that our approach induces unauthorized cross-app actions, achieving 82.5% planning and 75.0% execution hijack rates on GPT-4o. Our findings highlight a previously underexplored attack surface in mobile agent systems and underscore the need for defenses that incorporate interaction-level signals.

cs.CR↗

Dual-Envelope Constrained Nonlinear MPC for Distributed Drive Electric Vehicles Drifting Under Bounded Steering and Direct Yaw-Moment Control

Distributed drive electric vehicles offer superior yaw moment control for autonomous drifting in extreme maneuvers. Conventional drift analysis constructs stability boundaries from open loop equilibria points and assumes a fixed envelope structure. However, coupling among control inputs reshapes the phase plane and shifts saddle point location, which can invalidate open loop envelopes when used for closed loop drifting. To address this issue, a saddle point coordinate model is established in this paper by combining a nonlinear tire model with the handling diagram and explicitly accounting for road adhesion coefficient, longitudinal velocity, front wheel steering angle, and additional yaw moment. Based on saddle point properties, an extended dual envelope framework is constructed in the phase plane of slip angle and yaw rate. Using the convergence tendency of state points toward saddle points under bounded control inputs, the outer envelope defines a recoverable set under constraints on front wheel steering angle and additional yaw moment. The inner envelope characterizes the non-drifting stability region associated with unsaturated tire forces. Finally, a nonlinear model predictive control (NMPC) controller is developed using the extended dual envelope constraint. Hardware-in-the-loop experiments show that, compared with NMPC without envelope constraints, the proposed method enables smoother convergence toward the drift saddle point, reduces the steady-state tracking errors of vehicle speed, sideslip angle, and yaw rate by 33.07%, 71.18%, and 31.27%, respectively, and decreases the peak tracking error by 63.66% under road-friction mismatch.

eess.SY↗

Cascade of Spin Liquids in a Bilayer Triangular-lattice Antiferromagnet Rb_2Co_2(SeO_3)_3

In frustrated Ising magnets, classical spin liquids (CSLs) with macroscopic ground-state degeneracy can survive against conventional magnetic order, as exemplified by systems on triangular, kagome and pyrochlore lattices at zero field. Here we report the discovery of a high-field route toward spin liquids in a bilayer triangular lattice antiferromagnet, Rb$_2$Co$_2$(SeO$_3$)$_3$. We demonstrate that a cascade of CSLs -- characterized by doubly degenerate one-up-one-down local spin configurations and a residual entropy of 1/2(1-M/M_s)Rln2 per mole -- emerges through field-controlled dilution of Ising dimers. Owing to the interplay of intra- and inter-layer interactions, these CSLs are further stabilized by lattice symmetry breaking at fractional magnetization plateaus. Such field-induced spin liquids can be understood as a consequence of generalized ice rules, analogous to those governing in pyrochlore antiferromagnets. In particular, the 5/6-plateau state is a candidate quantum spin liquid. Our results thereby establish a new pathway for exploring diverse spin liquid states across both classical and quantum regimes.

cond-mat.str-el↗

LOPT: Learning Optimal Pigovian Tax in Sequential Social Dilemmas

In multi-agent reinforcement learning, each agent acts to maximize its individual accumulated rewards. Nevertheless, individual accumulated rewards could not fully reflect how others perceive them, resulting in selfish behaviors that undermine global performance. The externality theory, defined as ``the activities of one economic actor affect the activities of another in ways that are not reflected in market transactions,'' is applicable to analyze the social dilemmas in MARL. One of its most profound non-market solutions, ``Pigovian Tax'', which internalizes externalities by taxing those who create negative externalities and subsidizing those who create positive externalities, could aid in developing a mechanism to resolve MARL's social dilemmas. The purpose of this paper is to apply externality theory to analyze social dilemmas in MARL. To internalize the externalities in MARL, the \textbf{L}earning \textbf{O}ptimal \textbf{P}igovian \textbf{T}ax method (LOPT), is proposed, where an additional agent is introduced to learn the tax/allowance allocation policy so as to approximate the optimal ``Pigovian Tax'' which accurately reflects the externalities for all agents. Furthermore, a reward shaping mechanism based on the approximated optimal ``Pigovian Tax'' is applied to reduce the social cost of each agent and tries to alleviate the social dilemmas. Compared with existing state-of-the-art methods, the proposed LOPT leads to higher collective social welfare in both the Escape Room and the Cleanup environments, which shows the superiority of our method in solving social dilemmas.

cs.MA↗