arXiv Science⌕ Search

arXiv subjects

Bo Yang

Publications and source records attributed to Bo Yang.

At least 37 records · Page 2Linked to original sources

SkyNative: A Native Multimodal Architecture for Remote Sensing Vision-Language Understanding

Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the token space of a large language model. Although effective, this modular RS-VLMs separates visual representation from language reasoning, potentially limiting the direct involvement of fine-grained visual evidence in complex spatial inference. This challenge is particularly relevant to remote sensing imagery, which often covers broad geographic areas and contains multi-scale objects, dense target distributions, and intricate spatial layouts. In this paper, we propose SkyNative, the first study to explore a native multimodal architecture for remote sensing vision-language tasks. SkyNative converts remote sensing images into visual tokens through a lightweight patch embedding module and places them together with text tokens in a shared autoregressive sequence, allowing textual tokens to directly access the preceding visual context. To accommodate the heterogeneous characteristics of the two modalities, we further adopt a modality-aware decoupling mechanism that applies modality-specific projections, normalization, and feed-forward transformations while processing both modalities through shared causal self-attention. Extensive experiments demonstrate SkyNative's strong capabilities in dense small-object perception, large-format contextual understanding, complex reasoning, and robustness, with scores of 68.93%, 47.40%, and 63.43% on HRRSD, RSHR reasoning, and OmniEarth, respectively. These results suggest that the native VLM architecture explored in SkyNative represents a promising approach to RS vision-language modeling.

cs.CV↗

Scalar curvature of blow-ups of compact Kähler manifolds along complex submanifolds

Let $(M,ω)$ be a compact Kähler manifold with its scalar curvature $S(ω)$, Assume that $M$ has complex dimension at least $3$ and contains a complex submanifold $X$ of complex codimension at least $2$. Let $σ: Bl_{X} M \rightarrow M$ denote the blow-up of $M$ along $X$. We show that $Bl_{X} M$ admits a sequence of Kähler metrics $\{\widetildeω_{i}\}_{i \geq 1}$ whose scalar curvatures $S(\widetildeω_i)$ converge to $σ^{\ast} (S(ω))$ in the $C^0(Bl_{X} M)$ norm. Our work is motivated by a recent result of Brown, who established the corresponding result for blow-ups at a point. The proof is based on the gluing method for constructing extremal Kähler metrics on blow-ups, together with new analytic tools and several modifications needed in our setting.

math.DG↗

Non-Abelian braiding in Abelian Fractional Quantum Hall Phases from realistic interactions

We propose a method of realizing non-Abelian braiding of fractionalized quasiholes in the Laughlin fractional quantum Hall phase at $ν=1/3$ with realistic two-body interactions within the lowest Landau level. It is numerically shown that low-lying gapped excitations near $ν=1/3$ are contained almost entirely within the null space of the three-body Moore-Read model Hamiltonian. They are thus quantum fluids of non-Abelian quasiholes that are in principle physically accessible. In particular, Laughlin ground state can be described as a fluid of ``$ψ$-type" quasiholes formed by binding a magnetic flux with a Majorana fermion (MF), and the Laughlin quasiholes are described by the ``$1$-type'' quasiholes, which are magnetic fluxes without a MF attached. Within the Laughlin phase, Laughlin quasiholes can be locally fractionalized into non-Abelian quasiholes, when the strong attraction between them is overcome by properly designed one-body electronstatic trapping potentials. Extensive numerics with proper finite-size scaling corroborate this physical picture, and our study points to the possibility of realizing non-Abelian braiding within an Abelian topological phase in experiment without the need for fine-tuning realistic electron-electron interaction.

cond-mat.str-el↗

Holomorphic functions on complete Hermitian manifolds with flat Chern connection

A classical result of Boothby states that any compact Hermitian manifold with flat Chern connection is covered by a complex Lie group. In this work, we prove a sharp generalization: the universal cover of any complete Hermitian manifold with vanishing Chern curvature and torsion of sublinear growth is holomorphically isometric to a complex Lie group equipped with a left-invariant metric. The proof relies on a new gradient estimate for holomorphic functions on complete Hermitian manifolds with nonnegative second Ricci curvature. Combining this estimate with methods from sub-Riemannian geometry, we establish quantitative characterizations of function theory on these manifolds.

math.DG↗

ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism

In financial trading, large language model (LLM)-based agents demonstrate significant potential, but their decisions can be sensitive to noisy and non-stationary market information. We propose ContestTrade, a multi-agent trading system with an internal competitive mechanism inspired by institutional investment workflows. The system consists of two specialized teams: (1) a Data Team that processes and condenses massive market data into diversified textual factors optimized for constrained LLM context windows, and (2) a Research Team that produces parallelized multipath trading decisions via tool-augmented deep research. The core design is a "Quantify-Predict-Allocate" contest mechanism within each team: agent outputs are scored only after market outcomes become observable, future utility is predicted from historical scores, and resources are allocated to agents with positive predicted utility. In a post-2024 A-share backtest, ContestTrade achieves higher backtested return and risk-adjusted performance than the evaluated baselines. We further describe the temporal protocol, implementation choices, and limitations to clarify the scope of these results.

q-fin.TR↗

Many-Anyon Braiding in Non-Abelian Fractional Quantum Hall Effect with Hybrid Monte Carlo Simulation

We employ the hybrid Monte Carlo method to efficiently compute the many-anyon non-Abelian braiding matrices associated with different braiding schemes of the Moore-Read quasiholes. A novel proposal in this work is that anyon braiding schemes based on a global rotation are robust against finite-size effects, as demonstrated by benchmarking their errors in the braiding matrix against those of a simple two-anyon exchange. Moreover, we investigate how electron-electron interactions and local electrostatic trapping potentials influence the energetic preference of different fusion channels. Their effect on the non-Abelian braiding matrices has been verified, a surprising phenomenon that demonstrates long-range entanglement of non-Abelian states. Our results are relevant to the experimental realization of non-Abelian physics in fractional quantum Hall and other analogous systems, including the fast-growing field of fractional quantum anomalous Hall states in moiré materials.

cond-mat.str-el↗

Parameter- and Bandwidth-Efficient Edge--cloud Many-to-Many Speech-to-Text Translation

Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from resource constraints, while centralized cloud systems incur bandwidth bottlenecks and privacy risks by transmitting raw voice data. In this paper, we propose Edge--cloud Speech Recognition and Translation (ESRT), a parameter-efficient, bandwidth-efficient, and privacy-aware collaborative Edge--cloud MLLM framework. First, we introduce a multi-task weighted curriculum learning strategy to mitigate catastrophic forgetting, improve multilingual balance, and train parameter-efficient ESRT-1B, ESRT-4B, and ESRT-12B models. Second, we enable bandwidth-efficient Edge--cloud inference by retaining a lightweight speech encoder and adapter on the device and transmitting only a compressed tensor to the cloud. Extensive experiments on FLEURS demonstrate that ESRT models achieve state-of-the-art S2TT performance across 45 languages ($45 \times 44$ directions). Relative to raw audio, ESRT and ESRT-Lite reduce the transmitted tensor size by $5.1\times$ and $10.2\times$, respectively, while keeping raw speech on-device and avoiding its direct exposure to the cloud. The code and models are released to facilitate reproducible, privacy-aware S2TT research.

cs.AI↗

Rogue-wave and lump patterns associated with the third Painlevé equation

We report rogue-wave and lump patterns associated with Umemura polynomials, which arise in rational solutions of the third Painlevé equation. We first show that in many integrable equations such as the nonlinear Schrödinger equation and the Boussinesq equation, when internal parameters of their rogue wave solutions are large and of certain form, then their rogue patterns in the spatial-temporal plane can be asymptotically predicted by root distributions of Umemura polynomials (or equivalently, pole distributions of rational solutions to the third Painlevé equation). Specifically, every simple root of the Umemura polynomial would induce a fundamental rogue wave whose spatial-temporal location is linearly related to that simple root, while a multiple root of the Umemura polynomial would induce a non-fundamental rogue wave in the $O(1)$ neighborhood of the spatial-temporal origin. Next, we show that in a certain class of higher-order lump solutions of the Kadomtsev-Petviashvili-I (KPI) equation, when their internal parameters are large and of certain form, then their lump patterns at $O(1)$ time can also be predicted asymptotically by root distributions of Umemura polynomials, where simple and multiple roots of the polynomial would give rise to fundamental and non-fundamental lumps in the spatial plane, respectively. These results reveal the importance of the third Painlevé equation in studies of nonlinear wave patterns. We also report a new transformation which turns bilinear rogue-wave solutions of the nonlinear Schrödinger equation to higher-order lump solutions of the KPI equation.

nlin.SI↗

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multilinguality: languages at different resource levels compete for limited representation capacity, leading to strong high-resource performance but substantial degradation on low-resource speech. To address this problem and improve multilingual consistency, we propose MSRT, a novel framework built around a resource-aware Mixture of Speech Encoders (MoSE). MoSE uses an explicit language router to assign each utterance to an appropriate expert encoder. A frozen expert preserves high-resource language capabilities, while a trainable expert adapts to and specializes in medium- and low-resource languages. We further introduce a five-stage curriculum learning strategy that substantially reduces data dependence, requiring only 10 hours of paired S2TT data per language for effective alignment. We conduct extensive experiments on 45 languages, systematically evaluating all $45 \times 44$ translation directions. Our 4B-parameter model achieves state-of-the-art performance, outperforming substantially larger baselines. Empirical analyses show that MoSE improves high-, medium-, and low-resource languages simultaneously, with the largest gains on low-resource speech, thereby breaking the curse of multilinguality without compromising high-resource performance. To support future multilingual S2TT research, we release our code and models.

cs.CL↗

Fairness in Augmented Graph Learning: A Survey

Graph learning has evolved into Augmented Graph Learning (AGL) by integrating specialized machine learning (ML) techniques. Examples include federated learning, graph transformers, and graph condensation. While enhancing model utility, AGL introduces unique intersectional fairness challenges that traditional GNN debiasing frameworks, which primarily focus on message-passing regulations, fail to address. This paper provides a systematic investigation into this emerging field, termed FairGX. We first delineate the shift from conventional fairness-aware graph learning to the FairGX paradigm, identifying novel bias sources inherent in ML augmentations, such as dual-side disparities in federated aggregation and attention-head skewness. A structured taxonomy is established to categorize existing literature based on their technical integration and fairness objectives. Furthermore, we analyze the impact of diverse ML paradigms on algorithmic equity, emphasizing the unique challenges in human-centered applications and the absence of a unified framework. We conclude by identifying five critical future directions, including novel metrics for AGL, fairness-privacy synergy, and Fairness-aware LLM4Graph/Graph4LLM. This survey serves as a foundational roadmap for developing robust and equitable graph systems in complex ML environments.

cs.LG↗

Verifiable blind probabilistic error cancellation

Quantum error mitigation (QEM) is an essential tool for mitigating hardware noise without incurring space overhead. Yet, its reliability depends on modeling, calibration, and implementation, leaving end-to-end security on untrusted quantum hardware unresolved. We address this problem by introducing verifiable blind probabilistic error cancellation (VBPEC), the first secure verification protocol against a fully malicious adversary that integrates QEM. VBPEC brings probabilistic error cancellation (PEC), a widely studied QEM technique, within the scope of composable security by formalizing delegated mitigation as a cryptographic resource in the abstract cryptography framework. The protocol performs PEC with perfect blindness and an exponentially small security error. VBPEC retains the absence of quantum-space overhead from recent statistically-secure verified quantum computation protocols and from PEC. The only overhead takes the form of additional repetitions due to the QEM procedure. To achieve this, we extend trap-based verification from deterministic pass/fail checks to statistical tests that benefit from QEM and develop a new proof technique that integrates the corresponding additional deviation sources. Rather than merely tolerating honest noise below a fixed threshold, VBPEC actively cancels it, enabling correctly mitigated estimates to be accepted with high probability without compromising security. Our framework thus establishes an essential route towards secure, reliable, and practical delegated quantum computation on near-future quantum hardware: VBPEC fundamentally improves the practicality of verification.

quant-ph↗

Three-term Recurrence Relation with Arbitrary Degree Steps for Orthogonal Polynomials

An approach to generate three-term recurrence relations with arbitrary degree steps is proposed for orthogonal polynomials. Specifically, given any class of orthogonal polynomials $\{Q_{p}(x)\}_{p=0}^{\infty}$ defined by Favard's theorem, we employ the adjacent members $Q_{p}(x)$ and $Q_{p-1}(x)$ to compute $Q_{p+s}(x)$ of high degree and the one of low degree $Q_{p-t}(x)$, where $(s,t)$ are parameters for degree step adjustment. The coefficients of both relations are analyzed, revealing novel properties that enable the derivation of three-term recurrence relations with respect to $Q_{p+s}(x)$, $Q_{p}(x)$ and $Q_{p-t}(x)$ by eliminating $Q_{p-1}(x)$. Furthermore, in addition to the standard recursive formula, which is characterized by degree increase, the formulas for degree decrease and end-to-middle directions are also formulated. Moreover, explicit recurrence relations with 2-degree steps are presented for Hermite, Gegenbauer and Legendre polynomials. The computation precision of the proposed recurrence relations is also compared with that of the standard ones.

math.NA↗

On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline

Vision-Language Pre-training Models (VLPMs) are known to be vulnerable to adversarial attacks. Recent transferable attacks on VLPMs have followed a common pipeline with complicated loss functions or multi-stage text/image attacks. However, in this paper, we demonstrate that such a sophisticated attack pipeline can be simpler yet more successful. Specifically, we identify three previously overlooked issues caused by inappropriate cross-modal interactions and excessive operations. To address them, we propose the Simple Vision-Language Attack (SimVLA) pipeline, which observably improves transferability and efficiency. Experiments on four datasets and three downstream tasks validate the superiority of our pipeline. For instance, on Flickr30k text-image retrieval dataset, our SimVLA outperforms the SOTA baseline in R@1 transferability by 8.01\%-14.71\%, while consuming only about 35.73\% of the time and 46.26\% of the max VRAM. Overall, the superiority of our SimVLA highlights the importance of leveraging domain knowledge (e.g., our proposed cross-modal word identification), while blindly pursuing intricate operations (e.g, complex loss functions and redundant multi-stage designs) may even be harmful. We hope our SimVLA can serve as a simple yet effective backbone for future extensions. Code is available at https://github.com/RYC-98/SimVLA.

cs.CV↗

Task-Oriented Sensing and Covert Transmissions for Collaborative Multi-AUV Systems

In underwater covert cooperative missions, autonomous underwater vehicles (AUVs) often cannot rely on active sonar to continuously obtain complete information, since active sensing and frequent communications increase the risk of exposure. As a result, AUVs primarily rely on passive observation, an approach that yields incomplete local perception and limited task efficiency. Although underwater acoustic communications can mitigate this limitation through information sharing, they are simultaneously constrained by long delays, severe interference, low reliability, and the risk of covert exposure. Existing communications-oriented multi-agent reinforcement learning (MARL) studies often model communication as an ideal information flow, whereas traditional communication optimization primarily focuses on link-level performance. However, both are insufficient to characterize the actual contribution of perceptual information to cooperative tasks under realistic conditions of covert physical communications. This paper proposes a Sensed Information Value Realization Multi-Agent Reinforcement Learning (SVR-MARL) framework that leverages practical information to characterize the utility of information for cooperative tasks and learns distributed cooperative policies under realistic communication and covert constraints. Through a case study of covert multi-AUV cooperative localization and tracking, the potential of the proposed framework to improve collaborative task efficiency while reducing unnecessary communication and exposure risks is demonstrated.

cs.LG↗

Observational Evidence for Counter-helicity Magnetic Reconnection in a Solar Eruption

Magnetic reconnection between coronal magnetic systems carrying opposite self-helicity may play a role in solar eruptions, but observational evidence remains limited. We investigate an M7.0 flare in NOAA Active Region 13615 on 2024 March 28 using multiwavelength observations and nonlinear force-free field extrapolations. The reconstructed coronal field reveals a low-lying positive-helicity core field beneath an overlying magnetic system of opposite sign. During the eruption, the footpoint connectivity of these two magnetic systems changes markedly: field lines rooted in the western footpoint region change from positive to negative helicity, and the positive-helicity domain is substantially reduced. These changes are accompanied by a remote chromospheric brightening, intermittent EUV stripe-like brightenings extending from the source region toward the remote chromospheric brightening, the subsequent formation of large-scale coronal loops, and a weak outer hard X-ray source located at a footpoint of the core field. Together, these results suggest that the eruption was closely associated with reconnection between the core field and the overlying counter-helicity system, providing observational evidence that counter-helicity reconnection can contribute to the destabilization of eruptive solar magnetic fields.

astro-ph.SR↗

Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit

We present a full-pipeline inference optimization for the MiMo-V2.5 model family, which combines Hybrid Sliding Window Attention (Hybrid SWA), sparse Mixture-of-Experts (MoE), and multimodal encoders. While Hybrid SWA can ideally reduce both attention compute and KVCache storage significantly compared to Full Attention, realizing these gains in production requires substantial engineering effort. We systematically optimize the KVCache system with layerwise prefetch, SWA-aware prefix cache trees, and specialized placement strategies, achieving strict $O(W)$ SWA storage and high cache hit rates. We further build GCache, a high-performance distributed cache infrastructure with RDMA-optimized networking, and develop a KVCache-affinity router to reduce computation while preserving load balancing. We also optimize for multimodal inputs, including GPU image preprocessing, parallel video decoding, and multimodal cache sharing. Together, these optimizations constitute the first large-scale LLM serving system in production that efficiently covers the Hybrid SWA + MoE + multimodal composite architecture.

cs.AR↗

Modeling Story Expectations: A Generative Framework using LLMs

Consumers' engagement with stories is shaped by their expectations about what will happen next, yet modeling these forward-looking beliefs over unstructured narrative content has remained challenging. We develop a framework that uses large language models to approximate consumers' story expectations. Our method generates multiple imagined story continuations from a pre-trained LLM and extracts interpretable, theory-motivated features from these continuations, such as emotion and narrative path features. We propose two complementary validation procedures suited to different data availability: a survey-based approach that compares LLM-derived expectations to human-reported beliefs, and a rational-expectations approach that compares them to actual story outcomes. Applying the framework to both survey data collected in a controlled lab setting and observational data from an online reading platform, we find that LLM-derived expectations correlate with human-reported beliefs as well as actual story continuations along all features studied. In both settings, forward-looking expectations are associated with reader engagement above and beyond features of the content already consumed. Our framework provides a scalable method for modeling consumer beliefs about narrative content, with implications for content creation, platform strategy, and the study of narrative media.

cs.CL↗

LLM-PDESR: Robust PDE Discovery via Subdomain Weighted Residuals and LLM-Guided Symbolic Hypothesis Generation

Discovering governing partial differential equations (PDEs) from noisy observational data is a fundamental challenge in scientific machine learning. Traditional symbolic regression (SR) methods often struggle to identify accurate equations within vast combinatorial search spaces, largely due to their inability to incorporate essential domain-specific prior knowledge. Furthermore, reliance on pointwise evaluations and discrete finite differences inherently amplifies high-frequency noise, creating deceptive fitness landscapes that derail the optimization process. To resolve these bottlenecks, we propose LLM-PDESR, a framework that integrates the structural hypothesis generation of Large Language Models (LLMs) with a mathematically rigorous evaluation environment. By employing C^4-continuous quintic splines for robust differentiation and subdomain weighted residuals as natural low-pass filters, our approach effectively mitigates the fitness landscape distortion that plagues existing methods. A Pareto-driven feedback loop then enables the LLM to iteratively refine candidate equations, balancing predictive accuracy with structural parsimony. We evaluate LLM-PDESR on 23 canonical PDEs and five structurally novel equations (including a multivariate system) specifically designed to preclude dataset memorization and test true discovery capabilities. Demonstrating real-world applicability, the framework successfully extracts a consistent structural skeleton for an interpretable 1D dynamical surrogate (1D-CACE) directly from noisy ERA5 reanalysis data. Extensive experiments and out-of-distribution testing confirm that LLM-PDESR significantly outperforms state-of-the-art methodologies in structural recovery, noise resilience, and the avoidance of spurious complexity and equation bloat.

cs.LG↗