arXiv ScienceSearch

arXiv subjects

Yuanwei Liu

Publications and source records attributed to Yuanwei Liu.

At least 19 recordsLinked to original sources

Agents in the Scene: An Agentic Framework for Resource-Efficient Site-Specific Base Station Deployment

An agentic framework is proposed for autonomous site-specific base station (BS) deployment in wireless network planning. In contrast to conventional approaches that rely on manual site surveys or extensive ray-tracing (RT) simulations with significant human intervention, the proposed framework autonomously explores and optimizes BS deployment under a limited RT evaluation budget, enabling resource-efficient network planning. To this end, a continuous, geometry-grounded deployment action space is first constructed from three-dimensional (3D) wireless digital twins. Within this action space, an agent team operates through a stateful perception--reasoning--reflection loop. Specifically, a Placement Agent first generates candidate BS deployments in two complementary modes: an experience-guided mode that refines promising solutions, and an exploration mode that avoids getting stuck in local optima. After the candidate deployments are evaluated through RT, a Reflection Agent interprets the RT results together with the scene geometry, identifies performance-limiting factors such as blockage, overlapping coverage, and uncovered areas, and converts these diagnoses into guidance for subsequent deployment. Through this iterative process, site-specific experience is accumulated and deployment plans are optimized without human intervention. Numerical results in two realistic urban scenarios show that: 1) the proposed approach substantially outperforms heuristic, learning-based, and large language model (LLM)-assisted methods; 2) it achieves highly competitive coverage against the optimal solution while requiring substantially fewer transmitter-level RT evaluations; 3) site-specific reflection effectively turns raw RT feedback into refinement guidance, whereas the dual-mode mechanism preserves diversity and facilitates escape from local optima.

eess.SY

Discrete Coupling and Localized Motion for Pinching-Antenna Systems (PASS)

The practical implementation of pinching-antenna systems (PASS) is challenging due to hardware limitations in large-scale antenna movement and continuous radiation power adjustment. This paper proposes a practical PASS-enabled downlink multi-user multiple-input multiple-output communication framework that enables discrete radiation power control and localized discrete antenna movement. Specifically, a discrete coupling strength model is exploited to tune the radiation power at each pinching antenna (PA) through quantized coupling spacing levels. Moreover, each PA can only move among discrete locations within a limited region determined by the movement speed and duration. Based on the proposed framework, a joint optimization problem of the PA positions, coupling strength, and transmit beamforming is formulated. Considering waveguide attenuation, the total average power consumption is minimized, subject to each user's minimum SINR requirement and localized motion constraints. To address this coupled mixed-integer nonconvex optimization problem, a globally optimal branch-and-bound-based algorithm is first developed for the multi-waveguide single-user scenario. To further reduce complexity, a scalable genetic algorithm-assisted particle swarm optimization (GA-PSO) method is developed for the multi-waveguide multi-user scenario, where GA operations are incorporated to preserve population diversity and alleviate premature convergence. Simulation results demonstrate that the proposed design significantly reduces the power consumption compared with the conventional PASS schemes and MIMO architectures.

eess.SP

Reliable Near-Field Multi-User Positioning Informed by Two-Stage MUSIC

Near-field localization is a promising technique for high-resolution multi-user positioning in future wireless systems, but its performance is often degraded by scattering-induced coherent propagation. Existing near-field localization methods, which require separate parameter estimation and path/source association, suffer from high computation overhead and accumulated errors, and usually do not provide any guarantee on reliability. In this paper, we propose \emph{MUSIC-Net}, an end-to-end near-field positioning deep learning (DL) framework informed by two-stage MUltiple SIgnal Classification (MUSIC) in mixed line-of-sight (LoS) and non-LoS (NLoS) multi-path scenarios, which embeds the two-stage MUSIC objects into training to isolate the LoS-related signal subspace and to identify a surrogate distance. The proposed framework directly recovers multi-user positions without the need for involved NLoS parameter estimation or path/source association. Furthermore, we introduce split conformal prediction (SCP) to move beyond point-estimation-based positioning towards statistically guaranteed (confidence) set estimation for all users. Numerical results show that the proposed MUSIC-Net achieves lower mean positioning error (MPER) than existing benchmarks and yields tighter SCP-calibrated prediction regions, demonstrating both accurate LoS localization and efficient uncertainty quantification (UQ) in coherent multi-path environments.

eess.SP

Fast Beam-Brainstorm: Few-Step Generative Site-Specific Beamforming with Flexible Probing

A novel generative site-specific beamforming (GenSSBF) approach, termed fast beam-brainstorm (F-BBS), is proposed to address the practical bottlenecks of slow beam generation and fixed channel probing lengths in existing GenSSBF. To accelerate beam generation, F-BBS utilizes a two-stage distillation strategy that learns an average velocity field, instead of an instantaneous one, to guide the beam generative process. This strategy enables larger generation steps, realizing few-step or even one-step beam generation. Furthermore, to accommodate flexible channel probing lengths, a stochastic masking mechanism and a beam index-aware masked condition encoder are proposed, enabling a single trained model to operate with variable-length channel probing observations without retraining. Therefore, FBBS achieves the fast generation of high-fidelity communication beams from coarse and variable-length channel probing feedback, i.e., reference signal received power (RSRP), from user equipments. Simulation results on accurate ray-tracing datasets show that 1) F-BBS achieves comparable performance while reducing the beam generation cost by over 90% compared with diffusion-based GenSSBF solutions, 2) F-BBS realizes robust performance across variable channel probing length, and 3) FBBS offers a desirable trade-off between beamforming gain and beam probing overhead.

eess.SP

KDGen-BF: A Generative Site-Specific Multi-User Beamforming Approach

This paper proposes knowledge-distilled generative beamforming (KDGen-BF) framework for site-specific multi-user beamforming. KDGen-BF generates a multi-user beamforming weights from low-dimensional reference signal received power (RSRP) observations without acquiring instantaneous channel state information (CSI). To address the ambiguity caused by limited RSRP observations and interference coupling, KDGen-BF formulates multi-user beamforming as a conditional generation problem and directly outputs beamforming weights beyond a finite codebook. A diffusion transformer is trained through knowledge-distillation and exponential-moving-average (KD-EMA) guidance, and multi-candidate strategy is used for online deployment. Numerical results on multiple DeepMIMO scenarios demonstrate that: 1) under limited probing budgets, KDGen-BF outperforms all baselines; 2) with larger probing budgets, KDGen-BF achieves performance comparable to exhaustive search over the discrete Fourier transform (DFT) codebook and outperforms all other baselines; and 3) under noisy RSRP observations, KDGen-BF remains robust and outperforms all compared baselines.

eess.SP

Joint Communication and Control Beamforming: A Closed-Loop Control Perspective

A joint communication and control (JCC) framework is proposed, where a base station (BS) simultaneously serves multiple communication users (CUs) and controls a physical plant in a closed loop. In the downlink, BS-generated control inputs are transmitted to and recovered at the plant, with wireless actuation distortion incorporated into the plant-state evolution. In the uplink, the plant state is reported to the BS and tracked by a Kalman filter (KF) for subsequent control-input generation. To characterize long-term control performance under communication-control interference, finite- and infinite-horizon linear quadratic Gaussian (LQG) costs are derived, directly linking beamforming design to plant-state evolution. JCC beamforming problems are then formulated for vector- and scalar-valued control inputs to minimize the infinite-horizon LQG cost subject to per-user communication signal-to-interference-plus-noise ratio (SINR) requirements. For the vector case, a second-order cone programming (SOCP)-based successive convex approximation method is developed for the resulting nonconvex problem. For the scalar case, a closed-form infinite-horizon LQG cost is derived, and the communication-control Pareto boundary is optimally characterized by an SOCP-based bisection method. Its optimality follows from the strict monotonicity of the scalar control cost with respect to the control SINR. Numerical results show that the derived costs closely match Monte Carlo simulations, the KF accurately tracks the ground-truth plant-state trajectory, and the proposed methods consistently outperform the zero-forcing benchmark. This confirms the benefit of balancing communication-control interference, especially with limited spatial degrees of freedom (DoFs).

eess.SP

A Survey of Pinching-Antenna Systems (PASS)

The pinching-antenna system (PASS), recently proposed as a flexible-antenna technology, has been regarded as a promising solution for several challenges in next-generation wireless networks. It provides large-scale antenna reconfiguration, establishes stable line-of-sight links, mitigates signal blockage, and exploits near-field advantages through its distinctive architecture. This article aims to present a comprehensive overview of the state of the art in PASS. The fundamental principles of PASS are first discussed, including its hardware architecture, circuit and physical models, and signal models. Several emerging PASS designs, such as segmented waveguide-enabled PASS (SWAN), center-fed PASS (C-PASS), and multi-mode PASS (M-PASS), are subsequently introduced, and their design features are discussed. In addition, the properties and promising applications of PASS for wireless sensing are reviewed. On this basis, recent progress in the performance analysis of PASS for both communications and sensing is surveyed, and the performance gains achieved by PASS are highlighted. Existing research contributions in optimization and machine learning are also summarized, with the practical challenges of beamforming and resource allocation being identified in relation to the unique transmission structure and propagation characteristics of PASS. Finally, several variants of PASS are presented, and key implementation challenges that remain open for future study are discussed.

eess.SP

Continuous-Aperture Array-Based ISAC Over Fading Channels

A framework of continuous-aperture array (CAPA)-based integrated sensing and communications (ISAC) under a fading communication channel is proposed. A continuous operator-based signal model is developed, and the statistics of the communication channel gain are characterized via Landau's eigenvalue theorem. On this basis, the performance of the CAPA-based ISAC system is analyzed by considering three continuous beamforming designs: i) the sensing-centric (S-C) design that optimizes sensing performance, ii) the communication-centric (C-C) design that optimizes communication performance, and iii) the Pareto-optimal design that balances the sensing-communication trade-off. For the S-C and C-C design, closed-form expressions for the sensing rate (SR), ergodic communication rate (CR), and outage probability are derived, and high-signal-to-noise ratio asymptotic analysis is conducted to obtain the multiplexing and diversity gains. For the Pareto-optimal design, the Pareto-optimal beamformer achieving the Pareto boundary is derived, and the achievable SR-CR region is characterized. Numerical results demonstrate that the proposed CAPA-ISAC scheme outperforms both conventional spatially discrete arrays-based ISAC and CAPA-based frequency-division sensing and communications.

eess.SP

Douyin Multimodal Embedding Model Technical Report

Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content, such as Douyin, Xiaohongshu, and YouTube, demand both efficiency under billion-scale indexing and fine-grained discrimination for hard matching. Existing MLLM embedding models rarely satisfy both. Contrastive models are efficient but rely on pair-level supervision too coarse for fine-grained distinctions, while CoT-based models improve discrimination through explicit generation impractical to serve online. We present Douyin Multimodal Embedding (DME), a model trained in two stages to combine both strengths. Stage 1 performs large-scale contrastive pre-training that establishes a unified multimodal embedding space with broad modality and task coverage. Stage 2 supplements semantic sufficiency, the property that an embedding is grounded in retrieval-relevant evidence and preserves fine-grained counterpart-side semantics, via two mechanisms. Evidence-Grounded Typed Latent Reasoning organizes retrieval evidence through hidden-space latent reasoning, and Cross-Conditional Reconstruction enforces counterpart-side semantics through cross-directional autoregressive reconstruction. Both act only during training and add only marginal query-side overhead, so DME serves as efficiently as a standard contrastive encoder. On MMEB-v2, DME reaches state-of-the-art results at comparable scales for its 2B and 9B variants (74.8 and 78.4), with especially strong video and visual-document tasks. In production, DME delivers a 2.92% relative gain on Douyin's in-house offline evaluation set, is deployed across Douyin scenarios such as generative, image, and AI search, and yields a 0.1% Lifetime (LT) gain in online A/B testing on Douyin search.

cs.IR

Center-Fed Pinching Antenna System (C-PASS): Modeling, Analysis, and Beamforming Design

A generalized framework for the novel center-fed pinching antenna system (C-PASS) is proposed. Within this framework, closed-form expressions for the degree of freedom (DoF) and power scaling law of the proposed C-PASS are first derived. These theoretical results reveal that the achievable DoF scales linearly with the number of input ports, $M$, and the number of receive antennas, $K$. Furthermore, the derived power scaling laws demonstrate that the C-PASS achieves a power gain of order $\mathcal{O}(P_T M)$, where $P_T$ denotes the transmit power. Based on the proposed C-PASS modeling, a sum-rate maximization problem for the joint optimization of transmit and pinching beamforming is then formulated. To solve this highly coupled non-convex problem, an efficient alternating optimization algorithm is developed. More particularly, the transmit precoding and power splitting ratios are updated via derived closed-form solutions, while the pinching antenna positions and radiation coefficients are optimized using block coordinate descent (BCD) methods. Finally, our numerical results reveal that the single-waveguide C-PASS: 1) achieves superior DoF and power scaling laws compared to the single-waveguide PASS; and 2) outperforms the multi-waveguide PASS in high-attenuation regimes, yielding a substantial gain exceeding $10$ dB.

cs.IT

Uplink Positioning for PASS in Multipath Environments

Pinching-antenna systems (PASS) enhance wireless propagation by activating or placing pinching antennas (PAs) near users. Therefore, accurate uplink positioning is essential for efficient communication. In this paper, an uplink multi-carrier positioning framework is established for PASS in multipath environments. Matrix pencil (MP)-based and low-complexity Rank-1 ranging algorithms are proposed to estimate the distances between the PAs and the user. For the MP-based ranging algorithm, the line-of-sight (LoS) component is separated from non-line-of-sight components by exploiting the shift-invariance property of the Hankel matrix, thereby enabling accurate distance estimation. For the Rank-1 ranging algorithm, the dominant LoS delay is directly isolated through truncated singular value decomposition, thereby avoiding matrix inversions. Subsequently, a two-stage weighted nonlinear least-squares (WNLS) positioning algorithm is designed to estimate the three-dimensional user position. To gain further insights, a comprehensive theoretical performance analysis of the proposed ranging and positioning algorithms is conducted. The closed-form ranging variances and position error bound (PEB) are derived to reveal the error propagation mechanism. Numerical results demonstrate that: i) The MP-based algorithm achieves higher accuracy and robustness than the Rank-1-based algorithm, while the Rank-1-based algorithm has lower computational complexity. ii) The positioning error of the MP-based algorithm follows the same trend as the derived PEB, whereas the Rank-1 algorithm exhibits an error floor due to multipath bias. iii) The positioning accuracy of the MP algorithm improves as the number of subcarriers increases.

eess.SP

Mutual Coupling in Continuous Aperture Arrays: Physical Modeling and Beamforming Design

The phenomenon of mutual coupling in continuous aperture arrays (CAPAs) is studied. First, a general physical model for the phenomenon that accounts for both polarization and surface dissipation losses is developed. Then, the unipolarized coupling kernel is characterized, revealing that polarization induces anisotropic coupling and invalidates the conventional half-wavelength spacing rule for coupling elimination. Next, the beamforming design problem for CAPAs with coupling is formulated as a functional optimization problem, leading to the derivation of optimal beamforming structures via the calculus of variations. To address the challenge of inverting the coupling kernel in the optimal structure, two methods are proposed: 1) the kernel approximation method, which yields a closed-form solution via wavenumber-domain transformation and GaussLegendre quadrature, and 2) the conjugate gradient method, which addresses an equivalent quadratic functional optimization problem iteratively. Furthermore, the optimal array gain and beampattern are analyzed at the large-aperture limit. Finally, the proposed continuous mutual coupling model is extended to spatially discrete arrays (SPDAs), and comprehensive numerical results are provided, demonstrating that: 1) coupled SPDA performance correctly converges to the CAPA limit, while uncoupled models are shown to violate physics, 2) polarization results in anisotropic array gain behavior, and 3) the coupled beampattern exhibits higher directivity than the uncoupled beampattern.

cs.IT

On the Performance of Pinching-Antenna Systems (PASS) Under Dynamic Channels with Blockages

The performance of pinching-antenna systems (PASS) is fundamentally affected by line-of-sight (LoS) blockage in practical environments. In this paper, PASS is investigated under realistic, obstacle-induced blockage by jointly considering the LoS and non-LoS (NLoS) components, rather than relying on a LoS channel or a probabilistic blockage model. A geometry-aware blockage model is adopted, where a blockage region on the waveguide is defined according to the actual locations and geometric features of obstacles, such that a pinching-antenna (PA) located within the blockage region is unable to establish a LoS link to the user equipment (UE). The channel models of PASS are developed by jointly accounting for in-waveguide attenuation and spatial propagation loss. To quantify the impact of channel factors on PASS performance, a single-PA single-UE scenario is studied under Rayleigh and Rician fading channels. Closed-form expressions for the outage probability are derived for both cases. For the ergodic rate, a closed-form expression is obtained in the Rayleigh case, while a complete analytical expression and an approximate closed-form expression are derived in the Rician case. Analytical expressions are derived for the endpoints of the blockage region, and the deployment criteria of optimal PA are provided. Simulation results validate the analysis and reveal that: i) NLoS scattering has a twofold effect on PASS performance, potentially degrading the outage performance while improving the rate performance under Rician fading; ii) Sufficiently strong NLoS scattering can still sustain communication in the presence of LoS blockage; iii) The optimal PA position is jointly determined by the environment geometry and the interplay between spatial propagation loss and in-waveguide attenuation.

eess.SP

Site-Specific Learning for Low-Overhead Multi-User MIMO Beamforming

A low-overhead site-specific multi-user multiple-input multiple-output (MU-MIMO) beamforming framework is proposed. Conventional limited-feedback MU-MIMO relies on channel state information reference signal (CSI-RS) transmission and user feedback before grouping and beamforming, which requires substantial online overhead when the antenna dimension and candidate-user pool are large. To reduce this burden, the proposed framework exploits site-specific information (SSI), which captures local radio propagation features. By learning the mapping from low-overhead beam-domain observations to effective transmit spatial subspaces of users, the BS can infer inter-user separability before high-resolution CSI acquisition and construct a compact group-level CSI acquisition subspace for the selected users. This site-specific design can be implemented within the standard limited-feedback procedure using synchronization signal block (SSB)-based reference signal received power (RSRP) fingerprints for subspace inference and CSI-RS feedback for low-dimensional CSI refinement. Extensive numerical results demonstrate that the proposed framework can identify compatible user groups before CSI-RS acquisition, preserve most scheduled-user channel energy in a compact group subspace, and achieve higher effective rates than conventional systems with significantly lower overhead and user-side processing burden.

eess.SP

Continuous Aperture Array-Assisted Integrated Communication and Navigation in LEO Satellite Constellations

This paper proposes a novel continuous aperture array (CAPA)-assisted integrated communication and navigation (ICAN) framework for low Earth orbit (LEO) satellite constellations. Within this framework, an electromagnetic-based collaborative transmission model is developed, in which multiple satellites equipped with CAPAs simultaneously radiate downlink data streams and navigation reference signals over shared spectrum. Building upon this, the achievable communication rate and the navigation Cramer-Rao bound (CRB) are derived, which explicitly characterize the intrinsic coupling between the dual-function beamformers and system performance. To improve the positioning accuracy with communication quality of service guarantee, a joint beamforming optimization problem is formulated to minimize the average CRB subject to transmit power budgets and minimum rate constraints. To tackle the inherent infinite-dimensionality of the CAPA beamformer design, an ICAN channel subspace is introduced to equivalently transform the formulation into a tractable finite-dimensional problem, which is then efficiently solved via an iterative convex optimization algorithm. Finally, numerical results demonstrate that the proposed CAPA-assisted beamforming design algorithm significantly outperforms conventional discrete phased array architectures and other benchmark schemes, yielding notable improvements in ICAN performance.

cs.IT

B2X Networks: Joint Design of Communication and Control for Embodied Intelligence

This article proposes the concept of \emph{brain-body-to-everything (B2X)} networks to facilitate the integration of wireless networks and embodied intelligence. In this framework, the \emph{brain} refers to the intelligence functions for reasoning, planning, and decision-making, the \emph{body} denotes the physical embodied agent that senses and acts in the real world, and \emph{X} represents the surrounding ecosystem involved in the brain-body interaction loop. Two B2X architectures with \emph{distributed} and \emph{centralized} brains are introduced to characterize different placements of intelligence across the body, base station, and core network. The uplink and downlink designs of B2X networks are then discussed under a representative base-station-side brain setting. For the uplink, communication is redesigned for B2X state acquisition under event urgency, sensing volume, and simultaneous multi-body access. For the downlink, communication is redesigned to coordinate command delivery and conventional service under shared radio resources. Based on these uplink and downlink considerations, a communication-control Pareto boundary is further used to characterize the loop-level trade-off between wireless transmission performance and control quality in B2X networks. Finally, several open research problems are discussed to guide future B2X network design.

eess.SP

Pinching Antennas-Assisted Sensing: A Ziv-Zakai Bound (ZZB) Perspective

The sensing capability of the pinching-antenna system (PASS) is analyzed from a Ziv-Zakai bound (ZZB) perspective, motivated by the sensing ambiguity arising from the multimodal observation model inherent to PASS. In comparison to other Bayesian sensing bounds, the ZZB provides a lower bound on the mean-squared error (MSE) across a broad range of signal-to-noise ratios (SNRs) and accounts for ambiguity in the likelihood functions. First, an observation model is developed for an uplink sensing scenario where a single sensing target transmits uplink pilots to a single-waveguide PASS receiver equipped with multiple pinching antennas (PAs). Building on this model, general ZZB expressions are derived for arbitrary prior distributions of the target's position, and are then specialized to the Gaussian and uniform cases. Second, the asymptotic ZZBs in low- and high-SNR regimes are characterized, and the relationship between the ZZBs and the conventional Bayesian Cramér-Rao bound (BCRB) is further studied by introducing the concept of an ambiguity function. Furthermore, to reduce the high computational complexity of direct evaluation of the ZZB, SNR-free and SNR-aware surrogate objective functions are proposed to facilitate ZZB-based optimization for enhancing sensing performance. Numerical results demonstrate that: i) Compared with the BCRB, the ZZB provides a tight sensing performance lower bound over a wide range of SNRs, ii) the ambiguity-awareness of the ZZB can address the multimodality-induced ambiguity in sensing, thereby yielding a reliable lower bound on the MSE, and iii) the proposed surrogate objective functions enable effective ZZB minimization with a lower computational complexity.

eess.SP

Semantic Noise Aided Secure Image Transmission over MIMO Fading Channels

Existing semantic communications have exhibited satisfactory performance in many tasks, but secure image transmission remains insufficiently explored. We propose a novel secure image semantic communication (SISC) framework over multiple-input multiple-output (MIMO) fading channels. To ensure high-quality image reconstruction for the legitimate semantic user (SU) and simultaneously interfere with the eavesdropper (Eve), we design a semantic noise generation (SNG) network. This network generates a beneficial semantic noise map based on both the source features and the SU channel state information (CSI). An efficient channel estimation enhanced network is incorporated to obtain the accurate CSI and enhance the system performance. Furthermore, to improve the secure image reconstruction quality, we develop an efficient transceiver beamformer optimization algorithm, where the formulated problem is solved using the constrained stochastic successive convex approximation method. In the proposed SISC framework, semantic noise generation and beamforming optimization work together to ensure secure and high-quality image transmission. Numerical results demonstrate that the proposed semantic noise aided transmission scheme effectively protects image information from leakage to Eve while maintaining high-fidelity image reconstruction at SU.

cs.IT