arXiv ScienceSearch

arXiv subjects

Mugen Peng

Publications and source records attributed to Mugen Peng.

At least 19 recordsLinked to original sources

Analytical Statistics of Vortex Beams in a Turbulent Channel for OAM-Multiplexed FSO Communications

Orbital angular momentum (OAM) multiplexing can increase the capacity of free-space optical (FSO) communications, whereas atmospheric turbulence causes modal crosstalk and irradiance fluctuations that degrade demultiplexing performance. Analytical modeling is therefore important for characterizing turbulence-induced propagation effects and the demultiplexed port-power statistics of OAM channels. In this paper, we first study the receiver-plane irradiance statistics of vortex beams after propagating through the turbulent channel. The average irradiance is derived using frequency-domain convolution, and a closed-form frequency-domain diffraction kernel is obtained based on extended Rytov theory to evaluate the scintillation index for moderate and strong turbulence. However, receiver-plane irradiance statistics alone are insufficient to describe the performance of OAM-multiplexed FSO communications. We therefore derive the demultiplexed port-power statistics. Specifically, we derive the average port power, modal crosstalk, port-power variance, and cross-port covariance based on the complex Gaussian expansion of general LG vortex fields and the extended Huygens-Fresnel framework. The demultiplexed port-power statistics are then used to evaluate the symbol-error rate (SER) of OAM-multiplexed FSO communications. Numerical results demonstrate that all derived statistics are consistent with those obtained by phase-screen simulations under different turbulence strengths and beam parameters. The resulting SER performance further shows that OAM-multiplexing performance is more sensitive to mode spacing for small receiver apertures than for large apertures.

eess.SP

HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving

Recent advances in large language models and multimodal models have pushed remote sensing (RS) processing from simple perception models to agentic systems designed to tackle complex, long-horizon RS tasks. However, existing systems often rely on monolithic decision-making frameworks, which fail to accommodate the multi-stage, interdependent nature of RS tasks. This centralized approach leads to challenges such as unstable task execution, incorrect tool usage, and error propagation across stages. To address these issues, we propose HiRS-Agent, a hierarchical multi-agent system for long-horizon RS task solving. HiRS-Agent adopts a two-level collaborative architecture: the Manager Layer handles dynamic routing, step-level verification, replanning, and termination control, while the Specialist Layer organizes domain-specific tools according to the RS workflow and is responsible for subtask reasoning and tool execution. To further enhance the system's capability, we introduce a two-stage supervised tuning strategy and a verification-guided hierarchical reinforcement learning stage to jointly optimize coordination and tool-use policies. Experiments on Earth-Agent Benchmark and ThinkGeo show that HiRS-Agent substantially improves long-horizon tool-use capability and final-task correctness, demonstrating the effectiveness of structured multi-agent collaboration for reliable RS agents. The code is publicly available at https://github.com/IntelliSensing/HiRS-Agent.

cs.AI

SCI-D$^2$NN: An Optimization Framework for OAM-Multiplexed FSO Communications

Orbital angular momentum (OAM) multiplexing can increase the capacity of free-space optical (FSO) communications, but its detection performance is strongly affected by impairments such as atmospheric turbulence, transmitter pointing errors, and photodetection noise. The diffractive deep neural network (D$^2$NN) can be used as an all-optical front end to mitigate turbulence-induced distortions before detection. However, existing D$^2$NN compensation schemes are not specifically optimized for communication detection. In this paper, we propose a supervised contrastive inspired D$^2$NN (SCI-D$^2$NN) framework for improving the detection performance of OAM-multiplexed FSO communications under these impairments. The proposed framework introduces two training branches: a projection branch that maps the optical field to low-dimensional decision domain samples, and a label branch that provides supervised labels to impose a separation constraint among decision domain samples. In addition, we characterize complex-amplitude crosstalk to obtain the receiver observation vector and formulate two detection schemes, namely single-port profile-likelihood detection and joint maximum-likelihood (ML) detection. We further design two SCI-D$^2$NN training losses called Bhattacharyya distance (BD) based loss and the ML based loss to improve decision domain separability and mitigate detection-performance degradation. Numerical results show that SCI-D$^2$NN achieves more than a 3-dB improvement in bit error rate (BER) over the conventional D$^2$NN baseline in most transmit-power regions. The BD based loss gives the lowest BER under different system parameters and provides more than a 10-dB BER improvement over the baseline in the high transmit power region.

eess.SP

Prior-Aided Iterative Channel Reconstruction with Optimized Frame Structure for DSE Mitigation in CP-OTFS-Based LEO Satellite Systems

Orthogonal time frequency space (OTFS) modulation has emerged as a promising solution to mitigate the severe Doppler shift in low Earth orbit (LEO) satellite communications. However, the frequency-dependent Doppler shift induced by the high mobility of LEO satellites leads to the Doppler squint effect (DSE). This effect compromises the channel sparsity in the delay-Doppler (DD) domain, rendering existing channel estimation methods ineffective. To overcome this challenge, this paper proposes a DSE-resilient transmission scheme for cyclic prefix OTFS (CP-OTFS)-based LEO satellite systems. Specifically, we analyze the input-output relationship of the CPOTFS- based LEO satellite communication system and derive a DSE-aware representation of the satellite-terrestrial channel in the DD domain. To efficiently capture DSE-aware channel characteristics, we propose a novel OTFS frame structure that allows the energy distribution of the received signal to serve as prior information for channel estimation. Meanwhile, this frame structure strategically allocates pilot symbols to achieve uniform energy distribution and reduce the peak-to-average power ratio (PAPR), while imposing a time-domain waveform continuity constraint to suppress out-of-band emission (OOBE) caused by rectangular pulses. Based on the frame structure, we propose a prior-aided iterative channel reconstruction (PAICR) algorithm to mitigate the severe power leakage induced by DSE. The proposed algorithm iteratively extracts and removes dominant channel components using Doppler-domain received signal energy observations, with a convergence criterion ensuring reliable termination. Furthermore, a Cramer-Rao lower bound is derived to provide a theoretical benchmark for evaluating the algorithm's performance.

eess.SP

Quantum-Limited Symbol-Blind Channel Estimation for Coherent State Discrimination

Residual dispersion breaks temporal-mode matching in photon-starved coherent links. For equiprobable $M$-ary PSK coherent states in a known spectral mode, with unknown symbols and carrier phase, we establish the quantum limit for blind joint estimation of group delay and second-order dispersion: after eliminating the common phase, it is $4N_s\mathbf{C}$, set by the covariance of the centered generators alone. A multi-output quantum pulse gate with photon-number-resolving detection locally attains it and supports reception below the standard quantum limit under turbulent fading.

quant-ph

Finite-Support Structure in i.i.d.-Constrained Capacity of Finite-Memory Poisson Channels

Discrete-time Poisson channels with finite intersymbol interference provide a natural model for direct-detection optical links in which multipath memory and signal-dependent shot noise appear simultaneously. Under peak and average optical-intensity constraints, we study the independent and identically distributed (i.i.d.)-constrained capacity problem of such channels. We prove that every input distribution maximizing the stationary mutual information rate within the i.i.d. input class has finite support. The proof is carried out directly on the entropy rate of the continuous-state hidden Markov output process induced by the finite-memory channel. We first establish a filtering-forgetting estimate whose constants are uniform over all admissible i.i.d. input laws. We then derive the entropy-rate first variation, construct a holomorphic extension of the corresponding influence function, and combine the Karush-Kuhn-Tucker condition with a supralinear growth argument.

cs.IT

GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space. However, existing studies lack a systematic evaluation that dissects these distinct competencies. To fill this gap, we introduce ChronoBench, a multidimensional benchmark that decomposes this task into four progressive cognitive levels (i.e., Land Cover Perception, Temporal Recognition, Long-Term Memory, and Spatio-Temporal Reasoning). The ChronoBench comprises 12 sub-tasks and 17,689 rigorously validated QA (Question-Answer) pairs. Extensive evaluations reveal that mainstream MLLMs fall drastically behind human experts, with Long-Term Memory emerging as the most critical bottleneck. Motivated by this finding, we further propose GeoChrono, an MLLM with enhanced capabilities for tracing, memorizing, and reasoning about long-term geographic evolution. Leveraging the physical prior that geographic parcels remain spatially fixed while their semantics evolve, we design a Temporal Trajectory Encoder~(TempEnc) that constructs per-location temporal trajectories for dedicated land cover evolution modeling, and we introduce a Coarse-to-Fine Token Compressor~(C2FComp) that adaptively preserves dynamic regions while compressing the static background. To support training, we also construct ChronoInstruct, a 104K-sample instruction-tuning dataset spanning all competency levels for training. GeoChrono achieves state-of-the-art performance on ChronoBench, surpassing the leading commercial MLLMs by over 20%, while C2FComp reduces visual tokens by over 56% while retaining GeoChrono's 94.6% performance. The code and data will be available at https://github.com/IntelliSensing/GeoChrono

cs.CV

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification. Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels. Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks. Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.

cs.CV

Graph-RHO: Critical-path-aware Heterogeneous Graph Network for Long-Horizon Flexible Job-Shop Scheduling

Long-horizon Flexible Job-Shop Scheduling~(FJSP) presents a formidable combinatorial challenge due to complex, interdependent decisions spanning extended time horizons. While learning-based Rolling Horizon Optimization~(RHO) has emerged as a promising paradigm to accelerate solving by identifying and fixing invariant operations, its effectiveness is hindered by the structural complexity of FJSP. Existing methods often fail to capture intricate graph-structured dependencies and ignore the asymmetric costs of prediction errors, in which misclassifying critical-path operations is significantly more detrimental than misclassifying non-critical ones. Furthermore, dynamic shifts in predictive confidence during the rolling process make static pruning thresholds inadequate. To address these limitations, we propose Graph-RHO, a novel critical-path-aware graph-based RHO framework. First, we introduce a topology-aware heterogeneous graph network that encodes subproblems as operation-machine graphs with multi-relational edges, leveraging edge-feature-aware message passing to predict operation stability. Second, we incorporate a critical-path-aware mechanism that injects inductive biases during training to distinguish highly sensitive bottleneck operations from robust ones. Third, we devise an adaptive thresholding strategy that dynamically calibrates decision boundaries based on online uncertainty estimation to align model predictions with the solver's search space. Extensive experiments on standard benchmarks demonstrate that \mbox{Graph-RHO} establishes a new state of the art in solution quality and computational efficiency. Remarkably, it exhibits exceptional zero-shot generalization, reducing solve time by over 30\% on large-scale instances (2000 operations) while achieving superior solution quality. Our code is available \href{https://github.com/IntelliSensing/Graph-RHO}{here}.

cs.LG

Efficient Off-Grid Near-Field Cascade Channel Estimation for XL-IRS Systems via Tucker Decomposition

Accurate cascaded channel state information is pivotal for extremely large-scale intelligent reflecting surfaces (XL-IRS) in next-generation wireless networks. However, the large XL-IRS aperture induces spherical wavefront propagation due to near-field (NF) effects, complicating cascaded channel estimation. Conventional dictionary-based methods suffer from cumulative quantization errors and high complexity, especially in uniform planar array (UPA) systems. To address these issues, we first propose a tensor modelization method for NF cascaded channels by exploiting the tensor product among the horizontal and vertical response vectors of the UPA-structured base station (BS) and the incident-reflective array response vector of the IRS. This structure leverages spatial characteristics, enabling independent estimation of factor matrices to improve efficiency. Meanwhile, to avoid quantization errors, we propose an off-grid cascaded channel estimation framework based on sparse Tucker decomposition. Specifically, we model the received signal as a Tucker tensor, where the sparse core tensor captures path gain-delay terms and three factor matrices are spanned by BS and NF IRS array responses. We then formulate a sparse core tensor minimization problem with tri-modal log-sum sparsity constraints to tackle the NP-hard challenge. Finally, the method is accelerated via higher-order singular value decomposition preprocessing, combined with majorization-minimization and a tailored tensor over-relaxation fast iterative shrinkage-thresholding technique. We derive the Cram\'er-Rao lower bound and conduct convergence analysis. Simulations show the proposed scheme achieves a 13.6 dB improvement in normalized mean square error over benchmarks with significantly reduced runtime.

eess.SP

ARIS-RSMA Enhanced ISAC System: Joint Rate Splitting and Beamforming Design

This letter proposes an active reconfigurable intelligent surface (ARIS) assisted rate-splitting multiple access (RSMA) integrated sensing and communication (ISAC) system to overcome the fairness bottleneck in multi-target sensing under obstructed line-of-sight environments. Beamforming at the transceiver and ARIS, along with rate splitting, are optimized to maximize the minimum multi-target echo signal-to-interference-plus-noise ratio under multi-user rate and power constraints. The intricate non-convex problem is decoupled into three subproblems and solved iteratively by majorization-minimization (MM) and sequential rank-one constraint relaxation (SROCR) algorithms. Simulations show our scheme outperforms nonorthogonal multiple access, space-division multiple access, and passive RIS baselines, approaching sensing-only upper bounds.

eess.SP

Accurate Network Traffic Matrix Prediction via LEAD: a Large Language Model-Enhanced Adapter-Based Conditional Diffusion Model

Driven by the evolution toward 6G and AI-native edge intelligence, network operations increasingly require predictive and risk-aware adaptation under stringent computation and latency constraints. Network Traffic Matrix (TM), which characterizes flow volumes between nodes, is a fundamental signal for proactive traffic engineering. However, accurate TM forecasting remains challenging due to the stochastic, non-linear, and bursty nature of network dynamics. Existing discriminative models often suffer from over-smoothing and provide limited uncertainty awareness, leading to poor fidelity under extreme bursts. To address these limitations, we propose LEAD, a Large Language Model (LLM)-Enhanced Adapter-based conditional Diffusion model. First, LEAD adopts a "Traffic-to-Image" paradigm to transform traffic matrices into RGB images, enabling global dependency modeling via vision backbones. Then, we design a "Frozen LLM with Trainable Adapter" model, which efficiently captures temporal semantics with limited computational cost. Moreover, we propose a Dual-Conditioning Strategy to precisely guide a diffusion model to generate complex, dynamic network traffic matrices. Experiments on the Abilene and GEANT datasets demonstrate that LEAD outperforms all baselines. On the Abilene dataset, LEAD attains a remarkable 45.2% reduction in RMSE against the best baseline, with the error margin rising only marginally from 0.1098 at one-step to 0.1134 at 20-step predictions. Meanwhile, on the GEANT dataset, LEAD achieves a 0.0258 RMSE at 20-step prediction horizon which is 27.3% lower than the best baseline.

cs.LG

Online Specific Emitter Identification via Collision-Alleviated Signal Hash

Specific Emitter Identification (SEI) has been widely studied, aiming to distinguish signals from different emitters given training samples from those emitters. However, real-world scenarios often require identifying signals from novel emitters previously unseen. Since these novel emitters only have a few or no prior samples, existing models struggle to identify signals from novel emitters online and tend to bias toward the distribution of seen emitters. To address these challenges, we propose the Online Specific Emitter Identification (OSEI) task, comprising both online \revise{few-shot and generalized zero-shot} learning tasks. It requires constructing models using signal samples from seen emitters and then identifying new samples from seen and novel emitters online during inference. We propose a novel hash-based model, Collision-Alleviated Signal Hash (CASH), providing a unified approach for addressing the OSEI task. The CASH operates in two steps: in the seen emitters identifying step, a signal encoder and a seen emitters identifier determine whether the signal sample is from seen emitters, mitigating the model from biasing toward seen emitters distribution. In the signal hash coding step, an online signal hasher assigns a hash code to each signal sample, identifying its specific emitter. Experimental results on real-world signal datasets (i.e., ADSB and ORACLE) demonstrate that our method accurately identifies signals from both seen and novel emitters online. This model outperforms existing methods by a minimum of 6.08\% and 8.55\% in accuracy for the few-shot and \revise{generalized zero-shot learning }tasks, respectively. The code will be open-sourced at \href{https://github.com/IntelliSensing/OSEI-CASH}{https://github.com/IntelliSensing/OSEI-CASH}.

eess.SP

Multi-Stage CD-Kennedy Receiver for QPSK Modulated CV-QKD in Turbulent Channels

Continuous variable-quantum key distribution (CV-QKD) protocols attract increasing attentions in recent years because they enjoy high secret key rate (SKR) and good compatibility with existing optical communication infrastructure. Classical coherent receivers are widely employed in coherent states based CV-QKD protocols, whose detection performance is bounded by the standard quantum limit (SQL). Recently, quantum receivers based on displacement operators are experimentally demonstrated with detection performance outperforming the SQL in various practical conditions. However, potential applications of quantum receivers in CV-QKD protocols under turbulent channels are still not well explored, while practical CV-QKD protocols must survive from the atmospheric turbulence in satellite-to-ground optical communication links. In this paper, we consider the possibility of using a quantum receiver called multi-stage CD-Kennedy receiver to enhance the SKR performance of a quadrature phase shift keying (QPSK) modulated CV-QKD protocol in turbulent channels. We first derive the error probability of the multi-stage CD-Kennedy receiver for detecting QPSK signals in turbulent channels and further propose three types of multi-stage CD-Kennedy receiver with different displacement choices, i.e., the Type-I, Type-II, and Type-III receivers. Then we derive the SKR of a QPSK modulated CV-QKD protocol using the multi-stage CD-Kennedy receiver and post-selection strategy in turbulent channels. Numerical results show that the multi-stage CD-Kennedy receiver can outperform the classical coherent receiver in turbulent channels in terms of both error probability and SKR performance and the Type-II receiver can tolerate worse channel conditions compared with Type-I and Type-III receivers in terms of error probability performance.

eess.SP

Channel Modeling of Satellite-to-Underwater Laser Communication Links: An Analytical-Monte Carlo Hybrid Approach

Channel modeling for satellite-to-underwater laser communication (StULC) links remains challenging due to long distances and the diversity of the channel constituents. The StULC channel is typically segmented into three isolated channels: the atmospheric channel, the air-water interface channel, and the underwater channel. Previous studies involving StULC channel modeling either focused on separated channels or neglected the combined effects of particles and turbulence on laser propagation. In this paper, we established a comprehensive StULC channel model by an analytical-Monte Carlo hybrid approach, taking into account the effects of both particles and turbulence. We first obtained the intensity distribution of the transmitted laser beam after passing through the turbulent atmosphere based on the extended Huygens-Fresnel principle. Then we derived a closed-form probability density function of the photon propagating direction after passing through the air-water interface, which greatly simplified the modeling of StULC links. At last, we employed a Monte Carlo method to model the underwater links and obtained the power distribution at the receiving plane. Based on the proposed StULC channel model, we analyzed the bit error rate and the outage probability under different environmental conditions. Numerical results demonstrated that, the influence of underwater particle concentration on the communication performance is much pronounced than those of both the atmospheric turbulence and the underwater turbulence. Notably, increasing the wind speed at the air-water interface does not significantly worsen the communication performance of the StULC links.

eess.SP

BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model

Bi-temporal satellite imagery supports critical applications such as urbanization monitoring and disaster assessment. Although powerful multimodal large language models~(MLLMs) have been applied in bi-temporal change analysis, previous methods process image pairs through direct concatenation, inadequately modeling temporal correlations and spatial semantic changes. This deficiency hampers visual-semantic alignment in change understanding, thereby constraining the overall effectiveness of current approaches. To address this gap, we propose BTCChat, a multi-temporal MLLM with advanced bi-temporal change understanding capability. BTCChat supports bi-temporal change captioning and retains single-image interpretation capability. To better capture temporal features and spatial semantic changes in image pairs, we design a Change Extraction module. Moreover, to enhance the model's attention to spatial details, we introduce a Prompt Augmentation mechanism, which incorporates contextual clues into the prompt to enhance model performance. Experimental results demonstrate that BTCChat achieves state-of-the-art performance on change captioning and visual question answering tasks. The code is available \href{https://github.com/IntelliSensing/BTCChat}{here}.

cs.CV

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial

The rapid advancement toward sixth-generation (6G) wireless networks has significantly intensified the complexity and scale of optimization problems, including resource allocation and trajectory design, often formulated as combinatorial problems in large discrete decision spaces. However, traditional optimization methods, such as heuristics and deep reinforcement learning (DRL), struggle to meet the demanding requirements of real-time adaptability, scalability, and dynamic handling of user intents in increasingly heterogeneous and resource-constrained network environments. Large language models (LLMs) present a transformative paradigm by enabling natural language-driven problem formulation, context-aware reasoning, and adaptive solution refinement through advanced semantic understanding and structured reasoning capabilities. This paper provides a systematic and comprehensive survey of LLM-enabled optimization frameworks tailored for wireless networks. We first introduce foundational design concepts and distinguish LLM-enabled methods from conventional optimization paradigms. Subsequently, we critically analyze key enabling methodologies, including natural language modeling, solver collaboration, and solution verification processes. Moreover, we explore representative case studies to demonstrate LLMs' transformative potential in practical scenarios such as optimization formulation, low-altitude economy networking, and intent networking. Finally, we discuss current research challenges, examine prominent open-source frameworks and datasets, and identify promising future directions to facilitate robust, scalable, and trustworthy LLM-enabled optimization solutions for next-generation wireless networks.

cs.NI

Uninformed-to-Informed Estimation: A Ping-Pong Positioning Method for Multi-user Wideband mmWave Systems

To enhance the positioning and tracking performance of dynamic user equipment (UE) in wideband millimeter-wave (mmWave) systems, we propose a novel positioning error lower bound (PELB)-driven ping-pong positioning framework, where the base station (BS) and UE alternately transmit and receive adaptive beamforming signals for positioning. All beam-formers are scheduled based on the locally evaluated PELB. In this framework, we exploit multi-dimensional information fusion to assist in positioning. Firstly, a multi-subcarrier collaborative positioning error lower bound (MSCPEB) is proposed to evaluate the positioning error limits of wideband mmWave systems, which quantifies the contribution of all subcarriers to positioning accuracy. Moreover, we prove that the MSCPEB does not exceed the arithmetic mean of the PELBs of the individual subcarriers. Subsequently, we develop an alternating optimization (AO) algorithm to optimize the hybrid beamformers targeted for MSCPEB minimization. By convexifying this problem, closed-form solutions of beamformers are derived. Finally, we develop a multipath collaborative positioning method that quantifies the impact of path reliability on positioning accuracy, with a closed-form solution for user position derived. The proposed method does not rely on path resolution and traditional triangular relationships. Numerical results validate that the proposed method improves estimation accuracy by at least 16% compared to potential schemes without optimized beam configurations, while requiring only approximately one-quarter of the slot resources.

eess.SP