arXiv Science⌕ Search

arXiv · 2610.05888

Generative-AI for XR Content Transmission in the Metaverse: Potential Approaches, Challenges, and a Generation-Driven Transmission Framework

Abstract

How to efficiently transmit large volumes of Extended Reality (XR) content through current networks has been a major bottleneck in realizing the Metaverse. The recently emerging Generative Artificial Intelligence (GAI) has already revolutionized various technological fields and provides promising solutions to this challenge. In this article, we first demonstrate current networks' bottlenecks for supporting XR content transmission in the Metaverse. Then, we explore the potential approaches and challenges of utilizing GAI to overcome these bottlenecks. To address these challenges, we propose a GAI-based XR content transmission framework which leverages a cloud-edge collaboration architecture. The cloud servers are responsible for storing and rendering the original XR content, while edge servers utilize GAI models to generate essential parts of XR content (e.g., subsequent frames, selected objects, etc.) when network resources are insufficient to transmit them. A Deep Reinforcement Learning (DRL)-based decision module is proposed to solve the decision-making problems. Our case study demonstrates that the proposed GAI-based transmission framework achieves a 2.8-fold increase in normal frame ratio (percentage of frames that meet the quality and latency requirements for XR content transmission) over baseline approaches, underscoring the potential of GAI models to facilitate XR content transmission in the Metaverse.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhe Zhang, Yili Jiang, Xin Wei, Mingkai Chen, Haiwei Dong, Shui Yu. 2026-10-05. Generative-AI for XR Content Transmission in the Metaverse: Potential Approaches, Challenges, and a Generation-Driven Transmission Framework. https://arxiv.org/abs/2610.05888

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Graph-Based Floor Separation Using Node Embeddings and Clustering of WiFi Trajectories

Vertical localization, particularly floor separation, remains a major challenge in indoor positioning systems operating in GPS-denied multistory environments. This paper proposes a fully data-driven, graph-based framework for blind floor separation using only Wi-Fi fingerprint trajectories, without requiring prior building information or knowledge of the number of floors. In the proposed method, Wi-Fi fingerprints are represented as nodes in a trajectory graph, where edges capture both signal similarity and sequential movement context. Structural node embeddings are learned via Node2Vec, and floor-level partitions are obtained using K-Means clustering with automatic cluster number estimation. The framework is evaluated on multiple publicly available datasets, including a newly released Huawei University Challenge 2021 dataset and a restructured version of the UJIIndoorLoc benchmark. Experimental results demonstrate that the proposed approach effectively captures the intrinsic vertical structure of multistory buildings using only received signal strength data. By eliminating dependence on building-specific metadata, the proposed method provides a scalable and practical solution for vertical localization in indoor environments.

cs.NI↗

Distributed Quantum-Assisted Robust AoII Minimization in Satellite-Ground Integrated Edge Networks

Mission-critical edge applications in 6G-and-beyond networks, such as autonomous systems, disaster response, and infrastructure monitoring, require that the edge decision-maker's estimate of a monitored process remain correct, not merely up to date. Satellite-ground integrated networks (SAGIN) often provide the only connectivity in infrastructure-limited or disaster-affected regions, yet satellite handover and shadowing interrupt links, during which the process may change state several times, leaving the edge node's estimate substantially wrong. Age of information (AoI) tracks only elapsed time and cannot distinguish a harmless delay from a dangerous error. We instead adopt the age of incorrect information (AoII), which penalizes both the duration and magnitude of estimation error, and formulate, to our knowledge, the first network-level, multi-node AoII minimization problem over SAGIN under stochastic handover and shadowing. We propose SENTINEL, a distributed hybrid quantum-classical framework that jointly schedules update rates, satellite-to-base-station associations, and bandwidth allocation to minimize the worst-case time-average AoII. Because AoII is history dependent, it resists per-slot optimization; a renewal-interval decomposition that separates source dynamics from channel disruption yields a closed-form AoII cost per inter-delivery interval. The resulting robust scheduling problem, VANGUARD, is formulated as a QUBO, mapped to an Ising Hamiltonian, and solved via distributed QAOA with ADMM-based coordination across satellite and ground domains. Every returned schedule carries a certified worst-case AoII over all disruption scenarios. Simulations show that SENTINEL outperforms learning-based and random baselines, matches a state-aware threshold policy in small networks while additionally providing a worst-case guarantee, and remains close to an exact minimax reference.

cs.NI↗

LocAttMamba: A Low-Complexity Mamba Framework with Attention-Based Multi-AP Fusion for Indoor Localization

Accurate and low-complexity indoor localization is important for location-based services in fifth generation (5G) and sixth generation (6G) networks, where positioning devices operate under limited computational budgets and non conditions. Indoor localization has been studied widely using traditional signal-level localization approaches. However, these techniques often show degraded performance in on-line-of-sight (NLoS) scenarios. Recently, artificial intelligence (AI)-based techniques, including transformers, have been applied to address these challenges. While transformer-based architectures can capture the dependencies within the measurements collected from distributed access points (APs), their high computational complexity results in a large number of multiply-accumulate operations and long inference time. In contrast, lightweight recurrent and convolutional models trade this cost for degraded accuracy. In this paper, we propose LocAttMamba, a low-complexity localization framework in which the channel impulse response (CIR) and time-based features of each AP are processed by a separate Mamba encoder with near-linear complexity, and the resulting per-AP embeddings are fused through a multi-head attention layer that weights each AP according to its importance at every time step. The framework jointly predicts the user location and its per-axis uncertainty, which is refined through a post-hoc calibration step. We evaluate the proposed framework using two real-world 5G and ultra-wideband (UWB) measurement datasets. Our numerical results reveal that LocAttMamba obtains a mean two-dimensional (2-D) positioning error of 1.004 m and 0.599 m, respectively, on 5G and UWB datasets, outperforming the second-best benchmark by 9.79% and 5.82%, while requiring the fewest multiply-accumulate operations among all evaluated models and being 4.4-16 times faster than the transformer-based benchmarks.

cs.NI↗