arXiv ScienceSearch

arXiv subjects

Hao Xu

Publications and source records attributed to Hao Xu.

At least 19 recordsLinked to original sources

MAGRPO: Accelerated MARL Training for Fluid Antenna-Assisted Wireless Network Optimization

Fluid antenna systems (FASs) improve wireless links by repositioning antenna elements to exploit favorable spatial channel variations. Jointly optimizing fluid antenna (FA) positions, beamforming, and transmit power in a multi-cell network is challenging because the problem is non-convex and each base station has only local information during decentralized execution. However, the representative multi-agent reinforcement learning (MARL) algorithms, namely the on-policy multi-agent proximal policy optimization (MAPPO) and the off-policy multi-agent twin delayed deep deterministic policy gradient (MATD3), suffer from excessively long training times. To address this challenge, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) and propose multi-agent group relative policy optimization (MAGRPO) under centralized training with decentralized execution. MAGRPO constructs relative advantages from groups of joint trajectories, thereby eliminating the centralized critic and generalized advantage estimation used by MAPPO; under parameter sharing, this critic-free design reduces the per-step computational complexity by approximately half. Simulations show that joint FA optimization provides several-fold sum-rate gains over fixed-position configurations. Across the evaluated antenna settings, MAGRPO achieves test sum rates higher than those of MATD3 and comparable to or slightly higher than those of MAPPO. Meanwhile, it reduces the training time by 20%-23% compared with MAPPO and by about 35% compared with MATD3.

cs.IT

Audio for Sports Highlight Detection: A Comparative Empirical Study

Sports highlight detection aims to identify the most exciting and meaningful moments from long sports videos. While existing methods often emphasize visual or visual-language representations, sports videos contain rich audio cues, including commentator speech, crowd reactions, whistles, ball impacts, and referee calls. In this work, we revisit the role of audio in sports highlight detection and ask a simple question: how far can audio alone go? We construct lightweight audio-only baselines using pretrained audio representations and compare them with visual-only and audio-visual methods on the SV-Highlights benchmark. Surprisingly, our audio-only GRU baseline achieves strong performance and outperforms several existing audio-visual methods under our supervised evaluation setting. Furthermore, a simple audio-visual fusion baseline achieves the best performance across all metrics, indicating that audio and visual cues provide complementary information. To better understand the contribution of audio, we conduct source-separated analysis and show that vocal/commentary audio is more informative than background-only audio, while their combination performs best. We also analyze interpretable audio cues and find that highlight clips exhibit higher RMS loudness, peak loudness, and mid-frequency energy than non-highlight clips, although substantial distribution overlap indicates that loudness alone is insufficient. Our findings suggest that audio is an underexplored but highly informative modality for sports highlight detection and should be treated as a primary signal rather than merely an auxiliary cue.

cs.CV

Sensing-Aided Secure Multicast in Rotatable Antenna-Enabled ISAC Systems

Acquiring the channel state information (CSI) of passive eavesdroppers remains a fundamental challenge in physical layer security. The sensing capability of integrated sensing and communication (ISAC) systems enables estimation of a potential eavesdropper's angle of departure (AoD) before secure transmission. Accordingly, a sensing-aided secure multicast scheme is proposed using a rotatable antenna (RA) architecture that combines array-level and element-level rotations with analog beamforming. The scheme comprises eavesdropper sensing and secure communication stages. In the sensing stage, the maximum likelihood estimator (MLE) and corresponding Cramer--Rao bound (CRB) are derived for eavesdropper AoD estimation. The two rotation levels are then optimized through cyclic coordinate search to minimize the worst-case CRB. The resulting AoD estimate and CRB determine the center and width of the angular uncertainty region, respectively. In the communication stage, the constant-modulus analog beamformer and RA configuration are jointly optimized to maximize the worst-case secrecy rate over this region. After angular discretization and smooth approximation, the resulting problem is solved using a product-space joint optimization framework. Numerical simulation results validate the convergence and effectiveness of the proposed algorithms. It is demonstrated that i) the proposed RA-enabled sensing design effectively improves the eavesdropper AoD estimation accuracy; ii) a high and nearly constant secrecy rate is maintained over the uncertainty region; and iii) the joint optimization of the two rotation levels yields lower CRBs and higher secrecy rates than schemes employing either a fixed-position array or a single rotation level.

cs.IT

Morphological Decoupling-Based Skeletal Classification for Clinical Assessment of Malocclusion

Malocclusion skeletal grading is a fundamental task in orthodontics, critical for diagnosis and treatment planning. Traditionally, cone-beam computed tomography (CBCT) is used for visual measurement, and the reconstructed lateral cephalograms are handed over to expert dentists for diagnosis. However, manual review is time-consuming, labor-intensive, and subject to inter-operator variability. Therefore, an automatic CBCT-based system is needed for reliable malocclusion skeletal grading. In this case, we develop TeethGNN, a novel graph-based framework designed to combine CBCT image features with morphological information for accurate and efficient malocclusion grading. TeethGNN utilizes a decoupled learnable decoder to directly predict key morphological indicators from CBCT images, eliminating the need for manual measurements. These morphological features are then fused with image features using a graph neural network (GNN), which effectively models the relationships between the modalities. To further enhance robustness and calibration, we introduce a collaborative calibration strategy. This strategy combines multi-scale graph adversarial perturbation for explicit calibration and nonlinear topological graph calibration for implicit confidence adjustment. Extensive experiments and ablation studies on our collected clinical dataset demonstrate that our malocclusion measurement system achieves 77.08\% in accuracy and 89.61\% in AUC, outperforming the compared state-of-the-art methods. These results validate the effectiveness of graph-based multimodal fusion and collaborative calibration in improving malocclusion grading performance. Our system shows strong potential for advancing computer-aided orthodontic diagnosis, providing an accurate and reliable solution for vision-based clinical measurement and diagnosis.

eess.IV

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown objects near known-class decision boundaries. We propose CODE (Cross-Modal Calibration and Dynamic Suppression), a unified inference-time framework with three complementary components. Cross-Modal Joint Confidence Calibration injects global visual prototypes to calibrate text-driven known-class predictions. Uncertainty-Guided Universal Objectness Enhancement measures classification hesitation from local visual responses to strengthen potential unknown objects. Dynamic Outlier Suppression via Confidence Margin replaces rigid suppression with a margin-aware adjustment that preserves ambiguous out-of-distribution instances. Experiments on the Real-World Detection benchmark demonstrate that, with the OWL-ViT L/14 backbone, CODE achieves 21.7 U-mAP and 40.8 K-mAP in Task 1, surpassing the previous state of the art by 2.6 and 2.3 points, respectively.

cs.CV

Implication of the observed $ψ(3770)\to p\bar{p}π^0$ for studying the $p\bar{p}\to ψ(3770)π^0$ process

We study the charmonium $p \bar{p} \to ψ(3770) π^0$ reaction using the effective lagrangian approach where the contributions from well established $N^*$ states are considered, and all parameters are fixed in the process of $e^+e^- \to p \bar{p}π^0$ at center of mass energy $\sqrt{s} = 3.773$ GeV. The experimental data on the line shape of the mass distribution of the $e^+e^- \to p\bar{p}π^0$ can be well reproduced. Based on the studying of $e^+e^- \to p \bar{p}π^0$, the total and differential cross sections of the $p \bar{p} \to ψ(3770) π^0$ reaction are predicted. At the same time we evaluated also the cross sections of the $p \bar{p} \to ψ(3686) π^0$ reaction. It is shown that the contribution of the nucleon pole to this reaction is largest from the reaction threshold within a wide range. However, the interference between nucleon pole and the other nucleon resonance can still change the angle distributions significantly. Those theoretical results may be test by the future experiments at $\overline{\mbox{P}}$ANDA.

hep-ph

Finite-blocklength Fluid Antenna Systems

This paper investigates fluid antenna systems (FASs) subject to finite-blocklength (FBL) constraints, motivated by the strict reliability-latency and ultra-massive connectivity requirements of future wireless networks. While FAS performance has been widely studied in the asymptotic regime, its behavior under FBL remains largely unexplored. Our objective is to develop a unified set of analytical tools for evaluating FASs under FBL that remains applicable across different spatial-correlation models. First, to establish accurate benchmarks for non-orthogonal finite-length user signature design, we characterize both the average and the worst-case correlation coefficients via extreme value theory (EVT) and derive closed-form predictions of the achievable correlation levels. Second, taking block error rate (BLER) as the fundamental FBL metric, we study joint detection and decoding in FAS-assisted links and derive a closed-form BLER expression that is universally applicable across channel models. Additionally, we revisit outage probability (OP) in the FBL regime and obtain tractable OP characterizations for both FASs and conventional multiple fixed-position antenna (FPA) systems. In order to reduce the computational burden for multi-fold integrals in correlated fading models, we further propose a Taylor-expansion-assisted mean value theorem for integrals (MVTI), thus enabling efficient performance evaluation with marginal accuracy loss. Numerical results validate the analysis and reveal that even single-antenna FASs can have superior spatial diversity relative to conventional multi-FPA systems. Moreover, under both FBL and interference-limited environments, FASs provide improved energy, spectral, and hardware efficiencies, hence highlighting FAS as a promising enabler for next-generation wireless networks.

cs.IT

GraphMed-LT: Patient-Specific Graph Memory with Latent Clinical Thought Refinement for Multi-Turn Medical Conversations

Multi-turn medical question answering (QA) aims to model realistic clinical diagnosis, where a doctor gathers patient information across multiple turns of conversation. Existing multi-turn medical conversation systems have shown promising progress, but they often rely on accumulated conversation histories as memory, leaving clinical evidence fragmented across turns. We propose GraphMed-LT, a patient-specific graph memory approach with latent clinical thought refinement for multi-turn medical conversations. GraphMed-LT extracts patient-specific clinical triplets from patient responses, retrieves relevant knowledge triplets, and organises them into an incrementally updated graph memory. The graph memory is projected into graph-conditioned evidence tokens and refined inside a trainable doctor agent through hidden-state feedback, enabling the agent to update its internal clinical context before asking follow-up questions or producing the final answer. Experiments on three multi-turn medical QA benchmarks show that GraphMed-LT consistently outperforms existing multi-turn medical conversation baselines across multiple LLM backbones, achieving up to a 6.3 percentage-point absolute improvement over the strongest baseline. Further analyses show that GraphMed-LT asks more answerable follow-up questions and provides consistent gains across medical specialties.

cs.CL

SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning

Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and phenotypes. Existing agents typically rely on static retrieval workflows or coarse-grained prompt rewriting, which can lead to instruction drift when reasoning procedures need to be updated. We propose SSE-Bio, a structured self-evolving agent with an agentic retrieval policy for multi-hop biomedical reasoning. Instead of globally rewriting agent instructions, SSE-Bio maintains a structured state, selectively retrieves knowledge triplets and prior templates through a trainable proxy policy, and improves its reasoning memory through fine-grained template editing. To optimise retrieval decisions, we introduce a proxy-training strategy based on group relative policy optimization, where the proxy is improved through decision-contrastive groups over alternative retrieval choices. Experiments on three biomedical multi-hop QA benchmarks show that SSE-Bio consistently outperforms existing baselines, achieving an improvement of 6.56 absolute points over the strongest self-evolving baseline on BioHopR.

cs.CL

The Equality Cases of the Weak Simplex Conjecture

Among $n+1$ equiprobable equal-energy signals in $\R^n$ under additive white Gaussian noise with maximum-likelihood decoding, which arrangement maximizes the probability of correct decoding? The question is Shannon's, recorded by Rice in 1950. Mulgund proved in 2026 that the regular-simplex value bounds the correct-decoding probability of every signal set at every signal-to-noise ratio, leaving open whether the simplex is the only maximizer. This paper determines the equality cases in a form stronger than uniqueness. A signal set other than a regular simplex falls strictly below the bound at every positive signal-to-noise ratio. Hence a code meeting the bound at one positive operating point is already a regular simplex, up to vertex relabeling and an orthogonal map. In probabilistic form, among the correlation matrices that signal sets induce, any matrix other than the identity gives a lower-orthant probability strictly above its independent counterpart at every finite threshold, leaving no room for a nontrivial equality. No code of ambient dimension below $n$ attains the bound. Under an energy budget $E$ with unrestricted blocklength the optimal codebook is uniquely the regular simplex of circumradius $\sqrt{E}$. Every optimal codeword therefore exhausts its allowance. Equality in the Simplex Mean Width Conjecture likewise occurs only at the regular simplex. The proof strengthens the first self-convolution step of Mulgund's argument with Royen's correlation theorem. The single-parameter rigidity is machine-checked in Lean 4.

cs.IT

Spectral Localization Principle for Entanglement Harvesting

We propose a unified physical principle for entanglement harvesting: the entanglement that two localized detectors can extract from a quantum field is determined solely by how localized the field's effective spectral density is. We demonstrate this in an analytically solvable model of two qubits coupled to a leaky single-mode cavity, which in turn couples to a continuous electromagnetic bath, and derive the maximum harvestable concurrence in closed form, $\mathcal{C}_{\max}(Q)=2e^{-π/(2Q)}(1+e^{-π/(2Q)})/(1+3e^{-π/Q})$, where $Q\equiv|Δ|/κ$ is the ratio of the qubit-cavity detuning $Δ$ to the cavity linewidth $κ$. In the high-$Q$ limit, $\mathcal{C}_{\max}\simeq1-π^{2}/(16Q^{2})$, so the entanglement is robust against cavity loss; in the low-$Q$ limit it decays exponentially to zero, consistent with the irreversible-reservoir character of a continuous field, where maximal entanglement is unattainable. Since $Q$ is proportional to the inverse participation ratio (IPR) of the effective spectral density, it is the single dimensionless parameter governing the crossover from deterministic gate-based entanglement ($Q\to\infty$) to vacuum harvesting ($Q\to0$). Our framework operationalizes the Reeh-Schlieder theorem by quantifying the fraction of vacuum correlations accessible to localized detectors. It also reveals a formal correspondence of the maximal concurrence with the IPR, analogous to the conductivity-participation-ratio relation in Anderson localization. The predicted $\mathcal{C}_{\max}(Q)$ curve is, in principle, directly observable in superconducting circuit QED experiments.

quant-ph

ATOM: Geometry-Aware Microgesture towards Object-Agnostic Tangible Interaction

This paper presents ATOM, an integrated framework towards agnostic and tangible object interactions with microgestures. Our goal is to support microgesture interactions across different everyday objects, with the capability to automatically leverage the geometric affordance of each object. We formulate a fingertip-aware detection pipeline to leverage generative 2D and 3D models for geometry enhancement and refinement. We then introduce a usability-based method to prioritize the detected elements based on their ergonomic suitability for interactions. Building on this foundation, we further develop an AR system to transform everyday handheld objects into tangible user interfaces with 0D, 1D, and 2D microgesture interactions. Across transitions among everyday cooking objects of varying shapes and sizes, ATOM outperformed ablation baselines in task completion, usability (SUS), and workload (NASA-TLX). A further study with 10 objects demonstrates ATOM's generalizability across objects and grasps, highlighting its potential towards fluid, object-agnostic tangible interaction in real-world AR scenarios.

cs.HC

Secure Cooperative THz ISAC via Mamba Empowered Graph Neural Network Precoding

The terahertz (THz) band offers abundant spectrum resources for high-throughput communication and ultra high-precision localization. This paper investigates secure communication in cooperative THz orthogonal frequency-division multiplexing (OFDM) bistatic integrated sensing and communications (ISAC) systems, where multiple base stations (BSs) equipped with extremely large-scale antenna arrays (ELAAs) collaboratively serve downlink users while concurrently locating multiple targets. Malicious targets are assumed to act as potential eavesdroppers attempting to intercept confidential information intended for legitimate users. To mitigate these threats, we formulate a joint optimization problem for analog beamforming, digital precoding, true-time delayers (TTDs), and sensing signal covariance matrix design. The objective is to maximize the minimum secrecy rate subject to Cramer-Rao bound (CRB) constraints that ensure localization accuracy. This problem is highly challenging due to the non-convex CRB constraint, strongly coupled variables, high computational complexity from ELAA, and near-field channel modeling. To address these challenges, we propose a novel data-driven framework that integrates graph neural networks (GNNs) with the Mamba architecture. Our proposed framework first encodes the interactions among users, targets, and BSs into a heterogeneous graph and then employs message passing to optimize vertex features. The Mamba blocks further enhance this process through their selection mechanism and state space modeling capabilities, enabling dynamic and context-aware optimization of beamforming, TTD configurations, and sensing parameters. Numerical simulations validate that the proposed method outperforms both conventional and learning-based baselines, while offering high computational efficiency and strong generalization across different network conditions.

cs.IT

Intelligent Wiretap Code Design: Exploiting Wireless Endogenous Security via Information Theory and Deep Learning Integration

Recent advancements in wireless endogenous security have explored leveraging the inherent randomness of wireless channels to enhance communication security, providing an effective alternative to traditional encryption methods. This paper proposes a wiretap coding scheme within the semantic communication framework, which leverages discrete semantic representations compatible with conventional digital modulation to jointly enhance communication security and reliability. We investigate two eavesdropping scenarios: (i) the eavesdropper employs a maximum a posteriori (MAP) decoder, and (ii) the eavesdropper has access to a decoder identical to that of the legitimate receiver. In the first scenario, we exploit mutual information as a metric to guide the design of an optimized coding strategy, minimizing information leakage while enhancing communication reliability. In the second scenario, considering the limitations of the eavesdropper's decoding capability, we employ generalized mutual information (GMI) to characterize recoverability under the prescribed decoding rule and guide reliability-aware code optimization.

cs.IT

CSI Reconstruction in Fluid Antenna Systems Without Spatial Covariance Priors

Fluid antenna systems (FASs) exploit many candidate ports for spatial diversity, but hardware constraints allow channel observations at only a few active ports. Whether full-port CSI can be recovered without pre-acquired channel statistics remains open. Under the Clarke isotropic scattering model, we show that the channel lies in a low-dimensional spatial modal subspace determined by the scattering environment rather than the total port count. Consequently, recovery becomes feasible when the number of observed ports reaches the modal dimension (i.e., $M\geq r$), even when $M\ll N$. We further establish a sharp feasibility threshold: reliable recovery is impossible below this dimension regardless of SNR, whereas accuracy improves with additional observations above it. By decomposing the recovery error into modal truncation, estimation, and learning components, we derive explicit tradeoffs among RF chains, pilot overhead, transmit power, and training data. These results enable scalable prior-free full-port CSI recovery with few active ports.

cs.IT

WirelessOpsAgent: A Benchmark and Agent Design for Action Assurance in Wireless Networks

Large language model (LLM) agents are emerging as planners for autonomous wireless network operations. Yet a task answer that is correct at proposal time can still be unsafe at execution time if supporting telemetry is stale or inconsistent. Existing benchmarks mainly evaluate task solving from fixed observations and leave support checking at execution time untested. We introduce WirelessOptBench, a benchmark for action assurance in wireless operations. It turns wireless tasks into execution state decision episodes with controlled telemetry faults and action constraints. We further develop WirelessOpsAgent, which grounds candidate actions in current evidence and repairs recoverable support failures before execution. Across three backbone evaluations with 600 episodes each, WirelessOpsAgent achieves up to 0.983 Exact Action Accuracy. On Claude Sonnet 4.6, the Unsafe APPLY Rate decreases from 82.2% to 10.3% relative to the safest baseline. We make WirelessOptBench available at https://anonymous.4open.science/r/wirelessopsbench-artifact-D969/.

cs.NI

The 2-character theory of finite 2-groups

We generalize the notion of character for 2-representations of finite 2-groups. The properties of 2-characters bear strong similarities to those classical characters of finite groups, including conjugation invariance, additivity, multiplicativity and orthogonality. With a careful analysis using homotopy fixed points and quotients for categories with 2-group actions, we prove that the category of class functors on a 2-group $\mathcal G$ is equivalent to the Drinfeld center of the 2-group algebra $\mathrm{Vec}_{\mathcal G}$, which categorifies the Fourier transform on finite abelian groups. After transferring the canonical nondegenerate braided monoidal structure from $\mathfrak Z_1(\mathrm{Vec}_{\mathcal G})$, we discover that irreducible 2-characters of $\mathcal G$ coincide with full centers of the corresponding 2-representations, which are in a one-to-one correspondence with Lagrangian algebras in the category of class functors on $\mathcal G$. In particular, the fusion rule of $2\mathrm{Rep}(\mathcal G)$ can be calculated from the pointwise product of Lagrangian algebras as class functors. From a topological quantum field theory (TQFT) point of view, the commutative Frobenius algebra structure on a 2-character is induced from a 2D topological sigma-model with target space $\lvert \mathrm{B} \mathcal G \rvert$.

math.RT

Classification of symmetric fusion categories over $\mathbb{R}$

We show that every symmetric fusion category over $\mathbb{R}$ is equivalent to the category of finite-dimensional semi-linear representations of a $\mathbb{Z}_2$-graded finite super group. The proof uses Galois descent for tensor categories over $\mathbb{C}/\mathbb{R}$, reducing the classification to semi-linear $\mathbb{Z}_2$-actions on symmetric fusion categories over $\mathbb{C}$. As a further structural result, we establish a Tannaka-Krein type correspondence between symmetric fusion categories over $\mathbb{R}$ and finite groupoids with a $\mathbb{Z}_2 \times \mathrm{B} \mathbb{Z}_2$-action. This gives a complete real analogue of Deligne's classification result.

math.QA