arXiv ScienceSearch

arXiv subjects

Geng Chen

Publications and source records attributed to Geng Chen.

At least 19 recordsLinked to original sources

Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities under reinforcement learning (RL) paradigm. However, most existing multimodal medical reasoning models focus on basic reasoning, which refers to shallow inference based on visual feature matching. In contrast, real-world clinical diagnosis extends beyond basic reasoning, demanding complex reasoning that integrates heterogeneous clinical information (such as chief complaints and medical history) with multimodal medical imaging data. To bridge this gap, we introduce MM-Retinal-Reason, an ophthalmic multimodal dataset covering the full spectrum of perception and reasoning. Specifically, it is the first dataset in ophthalmology to encompass both basic and complex reasoning tasks with Chain-of-Thought (CoT) trajectories, aiming to enhance visual-centric reasoning and emulate realistic clinical decision-making. Building upon MM-Retinal-Reason, we propose OphthaReason, the first RL-enhanced ophthalmic multimodal reasoning model with step-by-step reasoning traces. To enable flexible adaptation to both basic and complex reasoning tasks, we further introduce Uncertainty-Aware Dynamic Thinking (UADT), which estimates sample-level uncertainty via entropy and dynamically modulates exploration depth through a shaped advantage mechanism. Comprehensive experiments demonstrate the effectiveness of our model on both basic and complex reasoning tasks, outperforming general-purpose MLLMs, medical MLLMs, RL-based medical MLLMs, and ophthalmic MLLMs by at least 15.47\%. Project Page: \href{https://github.com/lxirich/OphthaReason}{link}.

cs.AI

Structurally stable singularities and Lipschitz stable optimal transport metrics for the compressible Euler equations

It is well known that solutions to the compressible Euler equations can develop singularities in finite time. In this paper, we carry out a detailed analysis on behaviors of solutions up to the time of the first singularity for the one-dimensional compressible Euler equations with general smooth initial data. Our main results consist of three parts. First, for an open dense set of $C^3$ initial data, we show that the solution of Euler equations is twice continuously differentiable except at most finitely many points when the first singularity happens, using Thom's Transversality Theorem. Second, for any initial data in the open dense set of $C^3$ functions given in the first result, we provide the precise asymptotic description of the solution in a semi-neighborhood in the $(x,t)$-plane of each singular point at the time of the first singularity, and verify that the solution has a cusp-type singularity with Hölder exponent $1/3$ at each singular point. The proofs of the first two results are based on the representation of the solution in terms of a semilinear system. Third, for smooth initial data with small BV norm, we construct two Finsler type optimal transport metrics, then under these metrics show that the solution depends Lipschitz continuously on the initial data up to the time of the first singularity, with uniformly bounded Lipschitz constants. In particular, the $C^{1/3}$ generic singularity is stable in this sense. On the other hand, since our first two results hold for an open and dense set of initial data, any Hölder continuous cusp singularity with exponent other than $1/3$ is unstable under initial perturbations.

math.AP

Regularity of Structurally Stable Cusp Singularities for Two Families of Quasilinear Wave-type Equations

In this paper, we study two families of quasilinear equations: Hunter-Saxton type and Camassa-Hom type equations, with a paramerter $λ\in(0,1)$ whose solutions form cusp singularities. When $λ=1$, the first system becomes the scalar conservation law. The main result of this paper is to give regularity of two types of structurally stable singularities: Type I on the singular curve, Type II at the point where cusp singularity forms, for some $λ\in(0,1)$. When $λ\rightarrow 1$, our result indicts the $C^{1/3}$ regularity at the point where singularity forms, which agrees with the regularity of the generic pre-shock solution.

math.AP

Correlation Geometry of Quantum Sensor Networks: Local-Global Information Flow and Local Privacy

Quantum sensor networks (QSN) typically encode N unknown parameters while targeting a single linear combination, rendering the N-1 remaining parameters as nuisance directions. To rigorously quantify estimation precision under such nuisances, we use the effective quantum Fisher information (EQFI) and establish a ``barrel-effect'' bottleneck: the global EQFI cannot exceed the weakest weighted local sensing capacity. To elucidate the information allocation mechanism underlying this bottleneck, we derive an exact local--global phase map that delineates how the trade-off between local and global EQFI depends dynamically on quantum correlations, and accordingly we identify concrete conditions for saturating the bottleneck bound. Notably, this geometric map uncovers a counterintuitive ``overcorrelated'' regime where excessive correlations actively degrade both local and global performance. Finally, we apply the phase map to intrinsic local privacy and identify the condition under which every local parameter is inaccessible while the desired global combination remains estimable. Overall, our work provides a principled methodology for engineering optimal network states in quantum sensing architectures.

quant-ph

The $L^2$ contraction of solutions with large perturbation in multiple space dimensions from the oscillatory dispersive planar shock

In this paper, we show the $L^2$ contraction property of the planar oscillatory or monotone dispersive shock profiles of the dissipative Kadomtsev-Petviashvili (KP) equation modelling water waves and the multi-dimensional Korteweg-de Vries (KdV) Burgers equation under arbitrarily large perturbations in two space dimensions, up to Lipschitz time-dependent shifts. This stability result extends the results of two recent papers by Chen, Eun, Kang, and Shen on $L^2$ contraction for the KdV-Burgers equation.

math.AP

ARGON: A GNN-Empowered Compilation Framework for Scalable Neutral Atom Computing

Neutral atom quantum systems offer a promising pathway to large-scale quantum computing due to high qubit uniformity and flexible connectivity. To exploit this architecture, compilers must coordinate dynamic atom transport alongside highly parallel entangling gates. As circuits scale, the interplay between these operations becomes a system bottleneck, introducing denser logical interactions and longer temporal dependencies. Compilers must simultaneously satisfy rigid spatial constraints and complex movement schedules. Existing joint spatiotemporal compilation methods face an exponentially expanding search space, incurring substantial overheads or compromising fidelity as circuit size grows. In this work, we propose ARGON, a scalable compilation framework that introduces a spatiotemporal decoupling paradigm for neutral atom processors. Our key novelty is offloading static geometric conflict resolution to an offline phase, precomputing a library of hardware-certified, high-parallelism spatial layouts. To guide temporal routing, we deploy a Graph Neural Network (GNN) predictor to evaluate candidate layouts against deep temporal horizons, proactively evading downstream kinematic bottlenecks. Finally, a heuristic router translates the selected sequence into collision-free physical transport. Evaluations show ARGON completes compilation in under 10 seconds, delivering up to a >10^4x and 600x average speedup over state-of-the-art baselines. ARGON also minimizes routing decoherence and reduces Rydberg stages, improving execution fidelity by up to 10^2x on dense circuits.

cs.ET

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limited in scope, difficult to extend, and fragmented across institutions. We introduce EgoVerse, a collaborative platform for human data-driven robot learning that unifies data collection, processing, and access under a shared framework, enabling contributions from individual researchers, academic labs, and industry partners. The current release includes 1,362 hours (80k episodes) of human demonstrations spanning 1,965 tasks, 240 scenes, and 2,087 unique demonstrators, with standardized formats, manipulation-relevant annotations, and tooling for downstream learning. Beyond the dataset, we conduct a large-scale study of human-to-robot transfer with experiments replicated across multiple labs, tasks, and robot embodiments under shared protocols. We find that policy performance generally improves with increased human data, but that effective scaling depends on alignment between human data and robot learning objectives. Together, the dataset, platform, and study establish a foundation for reproducible progress in human data-driven robot learning. Videos and additional information can be found at https://egoverse.ai/

cs.RO

Effect of an aligned current on the stability of oscillatory incompressible flow past a circular cylinder

The stability of incompressible flow past a circular cylinder under collinear steady and oscillatory forcing is investigated within a two-dimensional Floquet framework. The flow is parameterised by the Keulegan-Carpenter number $KC \in [4,12]$, the steady-to-oscillatory velocity ratio $m \in [0,1]$, and the oscillatory Reynolds number $Re_m \in [20,100]$. The loci of the leading Floquet multipliers, and hence case-specific bifurcation modes, are examined by progressively reducing $Re_m$ to subcritical values for prescribed $m$. A steady current with $m > 0.5$ gives rise to a period-doubling subharmonic bifurcation that does not occur in purely oscillatory flow, where only synchronous and quasi-periodic modes arise. For $Re_m = 100$, three key features are discernible. First, the neutral stability curve in $(KC,m)$ space is strongly non-monotonic in $m$, separating intrinsically stable regions from those with single unstable modes; a sub-region of striking mode re-stabilisation appears beyond $m \approx 0.9$, where the flow recovers a $Z_2$-symmetric state at peak Reynolds number $\approx 190$, despite the steady and oscillatory components each being individually unstable. Second, a distinct regime supports the coexistence of two unstable modes of different types. Third, complementary direct numerical simulations show that, for a single unstable mode, the linear analysis successfully predicts the saturated nonlinear state even when $Re_m = 100$ substantially exceeds the critical Reynolds number, whereas under mode coexistence the quasi-periodic attractor tends to dominate the developed dynamics.

physics.flu-dyn

Long-Horizon Manipulation via Trace-Conditioned VLA Planning

Long-horizon manipulation remains challenging for vision-language-action (VLA) policies: real tasks are multi-step, progress-dependent, and brittle to compounding execution errors. We present LoHo-Manip, a modular framework that scales short-horizon VLA execution to long-horizon instruction following via a dedicated task-management VLM. The manager is decoupled from the executor and is invoked in a receding-horizon manner: given the current observation, it predicts a progress-aware remaining plan that combines (i) a subtask sequence with an explicit done + remaining split as lightweight language memory, and (ii) a visual trace -- a compact 2D keypoint trajectory prompt specifying where to go and what to approach next. The executor VLA is adapted to condition on the rendered trace, thereby turning long-horizon decision-making into repeated local control by following the trace. Crucially, predicting the remaining plan at each step yields an implicit closed loop: failed steps persist in subsequent outputs, and traces update accordingly, enabling automatic continuation and replanning without hand-crafted recovery logic or brittle visual-history buffers. Extensive experiments spanning embodied planning, long-horizon reasoning, trajectory prediction, and end-to-end manipulation in simulation and on a real Franka robot demonstrate strong gains in long-horizon success, robustness, and out-of-distribution generalization. Project page: https://www.liuisabella.com/LoHoManip

cs.RO

Scalable UAV Multi-Hop Networking via Multi-Agent Reinforcement Learning with Large Language Models

In disaster scenarios, establishing robust emergency communication networks is critical, and unmanned aerial vehicles (UAVs) offer a promising solution to rapidly restore connectivity. However, organizing UAVs to form multi-hop networks in large-scale dynamic environments presents significant challenges, including limitations in algorithmic scalability and the vast exploration space required for coordinated decision-making. To address these issues, we propose MRLMN, a novel framework that integrates multi-agent reinforcement learning (MARL) and large language models (LLMs) to jointly optimize UAV agents toward achieving optimal networking performance. The framework incorporates a grouping strategy with reward decomposition to enhance algorithmic scalability and balance decision-making across UAVs. In addition, behavioral constraints are applied to selected key UAVs to improve the robustness of the network. Furthermore, the framework integrates LLM agents, leveraging knowledge distillation to transfer their high-level decision-making capabilities to MARL agents. This enhances both the efficiency of exploration and the overall training process. In the distillation module, a Hungarian algorithm-based matching scheme is applied to align the decision outputs of the LLM and MARL agents and define the distillation loss. Extensive simulation results validate the effectiveness of our approach, demonstrating significant improvements in network performance over the MAPPO baseline and other comparison methods, including enhanced coverage and communication quality.

cs.MA

Existence and singularity formation for the supersonic expanding wave of radially symmetric non-isentropic compressible Euler equations

This paper studies the existence and singularity formation of supersonic expanding waves for the radially symmetric non-isentropic compressible Euler equations of polytropic gases. We introduce a suitable pair of gradient variables to characterize the rarefaction and compression properties of the solutions. Based on their Riccati equations, we construct several useful invariant domains to establish a series of priori estimates of solutions under some assumptions on the initial data. We show that the solution is smooth in the characteristic triangle or quadrangle domain if both of these two gradient variables are non-negative at the initial time. On the other hand, when one of these two variables is very negative at some initial point, the solution forms a singularity in finite time.

math.AP

$L^2$-contraction of Shock Waves for KdV-Burgers Equation

The KdV-Burgers equation is a canonical model describing the interplay between nonlinearity, viscosity and dispersion, and it admits viscous-dispersive shocks as traveling wave solutions. In this paper, we establish an $L^2$-contraction property for viscous-dispersive shocks under arbitrarily large perturbations, up to a time-dependent shift. This yields time-asymptotic stability and uniform estimates with respect to the strengths of viscosity and dispersion. We present the proof for the monotone shocks, and introduce the companion work in [6] on the stability and structural properties of oscillatory shocks.

math.AP

Complex Dynamics of Wave-Character Transitions in Radially Symmetric Isentropic Euler Flows: Theory and Numerics

We investigate the qualitative dynamics of smooth solutions to the radially symmetric isentropic compressible Euler equations, focusing specifically on the evolution of rarefactive and compressive wave characters across three distinct configurations: the outward supersonic, subsonic, and inward supersonic regimes. For each case, we establish structural restrictions on wave-character transitions and identify invariant sign domains for gradient variables under specific initial data conditions. While our findings refine existing invariance properties in the outward supersonic regime, they reveal novel asymmetric transition mechanisms in the subsonic and inward regimes that are absent in purely supersonic expanding cases. Consequently, we derive sufficient conditions for finite-time singularity formation. To complement the analytical results where closed-form solutions are unavailable, we provide numerical experiments using a Semi-Discrete Lagrangian-Eulerian (SDLE) formulation. These simulations reproduce the predicted wave-character dynamics and offer qualitative evidence that supports our theoretical findings, providing a unified description of wave transitions in radially symmetric isentropic gas dynamics.

math.AP

Uniform Stability of Oscillatory Shocks for KdV-Burgers Equation

We study viscous-dispersive shock waves with infinite oscillations of the Korteweg-de Vries-Burgers (KdVB) equation. First, we establish detail structures of the shock waves, including the rates at which the local extrema converge to the left end state towards the left far field. Then, by exploiting the structural properties of the shocks, we show the $L^2$-contraction property of the shock profiles under arbitrarily large perturbations, up to time-dependent shifts. This property implies both time-asymptotic stability and uniform stability with respect to the viscosity and dispersion coefficients. This uniformity yields the existence of zero viscosity-dispersion limits, on which Riemann shocks are orbitally stable.

math.AP

Informationally Complete Distributed Metrology Without a Shared Reference Frame

In quantum information processing, implementing arbitrary preparations and measurements on qubits necessitates precise information to identify a specific reference frame (RF). In space quantum communication and sensing, where a shared RF is absent, the interplay between locality and symmetry imposes fundamental restrictions on physical systems. A restriction on realizable unitary operations results in a no-go theorem prohibiting the extraction of locally encoded information in RF-independent distributed metrology. Here, we propose a reversed-encoding method applied to two copies of local-unitary-invariant network states. This approach circumvents the no-go theorem while simultaneously mitigating decoherence-like noise caused by RF misalignment, thereby enabling the complete recovery of the quantum Fisher information (QFI). Furthermore, we confirm local Bell-state measurements as an optimal strategy to saturate the QFI. Our findings pave the way for the field application of distributed quantum sensing, which is inherently subject to unknown RF misalignment and was previously precluded by the no-go theorem.

quant-ph

Non-commutativity as a Universal Characterization for Enhanced Quantum Metrology

A central challenge in quantum metrology is to effectively harness quantum resources to surpass classical precision bounds. Although recent studies suggest that the indefinite causal order may enable sensitivities to attain the super-Heisenberg scaling, the physical origins of such enhancements remain elusive. Here, we introduce the nilpotency index $\mathcal{K}$, which quantifies the depth of non-commutativity between operators during the encoding process, can act as a fundamental parameter governing quantum-enhanced sensing. We show that a finite $\mathcal{K}$ yields an enhanced scaling of root-mean-square error as $N^{-(1+\mathcal{K})}$. Meanwhile, the requirement for indefinite causal order arises only when the nested commutators become constant. Remarkably, in the limit $\mathcal{K} \to \infty$, exponential precision scaling $N^{-1}e^{-N}$ is achievable. We propose experimentally feasible protocols implementing these mechanisms, providing a systematic pathway towards practical quantum-enhanced metrology.

quant-ph

AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning

Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor setups. Existing multi-image benchmarks mainly target basic perception tasks using high-quality single-agent images, thus failing to evaluate MLLMs in more complex, egocentric collaborative scenarios, especially under real-world degraded perception conditions.To address these challenges, we introduce AirCopBench, the first comprehensive benchmark designed to evaluate MLLMs in embodied aerial collaborative perception under challenging perceptual conditions. AirCopBench includes 14.6k+ questions derived from both simulator and real-world data, spanning four key task dimensions: Scene Understanding, Object Understanding, Perception Assessment, and Collaborative Decision, across 14 task types. We construct the benchmark using data from challenging degraded-perception scenarios with annotated collaborative events, generating large-scale questions through model-, rule-, and human-based methods under rigorous quality control. Evaluations on 40 MLLMs show significant performance gaps in collaborative perception tasks, with the best model trailing humans by 24.38% on average and exhibiting inconsistent results across tasks. Fine-tuning experiments further confirm the feasibility of sim-to-real transfer in aerial collaborative perception and reasoning.

cs.CV