arXiv ScienceSearch

arXiv subjects

Yue Xiao

Publications and source records attributed to Yue Xiao.

At least 19 recordsLinked to original sources

Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training

Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.

cs.MA

AVP-Inspect: Coordinated Cyber-Physical Testing for Privacy Analysis of COTS Apple Vision Pro Applications

XR devices introduce substantial privacy concerns due to their comprehensive data collection capabilities that surpass traditional computing platforms. While existing works have demonstrated privacy concerns on Android-based XR devices such as Meta Quest series by performing network traffic analysis, little attention has been paid to the Apple Vision Pro (AVP) devices, mainly due to the closed nature and the technical challenges associated with AVP devices. In this work, we make a bold attempt to detect privacy violations of AVP applications from network traffic through automatic testing on AVP devices. Our key insight is that effective AVP application testing requires coordinated control of both cyber (software) and physical (hardware) components, which we term Coordinated Cyber-Physical Testing. Building on this insight, we design and implement AVP-Inspect, an automatic dynamic analysis framework for AVP applications, overcoming significant challenges enforced by the closed-source nature of AVP ecosystem. AVP-Inspect consists of three components: an automatic device controller by building customized hardware devices, a 3D UI explorer by designing a new exploration engine, and a privacy violation detector by constructing a unified privacy taxonomy for AVP. We first evaluated AVP-Inspect on a manually constructed ground truth dataset, then performed a large-scale analysis on 324 AVP applications downloaded from the App Store, with each app tested for 20 minutes. We found that 188 (58.0%) of apps exhibit at least one violation, and more than 60% of the network traffic flows are not properly disclosed.

cs.CR

Breaking Network Densification Limits with Distributed Cooperative Massive Access (DCMA)

In this work, we investigate the performance of the distributed cooperative massive access (DCMA) framework in large-scale network setups by incorporating stochastic geometry modeling. A partially centralized cell-free cloud-radio access (C-RAN) architecture is considered where remote radio heads (RRHs) decode transmitted messages and cooperate with each other to enhance system performance. Specifically, they can share decoded messages via feedback links, allowing receivers to cancel inter-user interference through successive interference cancellation (SIC), thus improving the decoding capabilities of the system. For such a network, we propose a novel synergetic decoding algorithm that efficiently resolves the assignment and message sharing routing for each user while accounting for practical network constraints. Furthermore, using game theory, we develop a merge-and-split algorithm with lexicographic preference to solve the problem of minimizing the RRHs utilized without compromising the performance. Simulation results show that the proposed framework significantly outperforms systems that do not implement SIC or take advantage of the cooperation between RRHs in terms of outage probability. Finally, we evaluate the performance of the proposed algorithms and validate their efficiency.

cs.GT

Tracking the boundary between absolute/convective instability using adjoint equations

Determining absolute/convective instability boundaries conventionally requires repeated saddle searches in the complex-wavenumber plane and a subsequent scan of the physical parameter space to locate zero absolute growth. Such nested calculations become costly and sensitive to modal branch association for large non-normal eigenvalue problems. This work develops a direct continuation method for neutral stationary-saddle boundaries of frequency-affine generalised eigenvalue problems. The zero-group-velocity condition is expressed as an adjoint solvability residual and solved together with the direct and adjoint eigenproblems, complex gauge constraints and the neutral-growth condition. The resulting one-dimensional solution manifold in the combined state--parameter space is tracked by scaled pseudo-arclength continuation, allowing parameter folds to be crossed without switching the physical continuation variable. The formulation recovers the analytical Ginzburg--Landau boundary and, for a Gaussian-wake Orr--Sommerfeld problem, agrees with separately formulated finite-difference saddle corrections to approximately $10^{-8}$ in relative critical Reynolds number. Compared with nested complex-wavenumber and parameter-plane saddle scanning, the tested scans require $8.1$--$52.2$ times the wall time of the direct adjoint continuation. Extrapolation of the measured cost--accuracy trend to a boundary error of $E_H\sim10^{-6}$ suggests an estimated cost ratio of approximately $1.8\times10^{4}$ in favour of the direct continuation. Application to a coupled Oldroyd--B free-surface film reveals genuine folds of the neutral-saddle manifold and a re-entrant CI--AI--CI boundary geometry for the selected saddle family.

physics.flu-dyn

Pinching Antenna-Aided Spatial Multiplexing: Transceiver Design and Performance Analysis

In this paper, a novel pinching antenna-aided spatial multiplexing (PASM) architecture is conceived, which intrinsically amalgamates the benefits of flexible radiating element placement with radio-frequency (RF) chain transmission. Specifically, we leverage the deterministic phase variation along dielectric waveguides as a zero-power phase-control mechanism, where each waveguide fed by a single RF chain drives multiple pinching antennas (PAs) acquiring position-dependent phase shifts. Then, the PASM propagation environment is characterized by a realistic channel model encompassing Rician small-scale fading, correlated shadowing, and large-scale path loss. Based on this, a low-complexity vector approximate message passing (VAMP) detector is conceived, which exploits a waveguide-structured prior for jointly processing the signals associated with all PAs. Moreover, we derive an analytical upper bound on the bit error rate (BER) for the maximum likelihood (ML) detector to quantify the achievable performance limits. Finally, our simulation results demonstrate that the proposed PASM architecture achieves substantial signal-to-noise ratio (SNR) gain over the conventional phase-shifter-aided spatial multiplexing (PSSM), while the VAMP detector strikes an attractive trade-off between the system performance and computational complexity.

eess.SP

AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents

Autonomous AI agents extend large language models into full runtime systems that load skills, ingest external content, maintain memory, plan multi-step actions, and invoke privileged tools. In such systems, security failures rarely remain confined to a single interface; instead, they can propagate across initialization, input processing, memory, decision-making, and execution, often becoming apparent only when harmful effects materialize in the environment. This paper presents AgentWard, a lifecycle-oriented, defense-in-depth architecture that systematically organizes protection across these five stages. AgentWard integrates stage-specific, heterogeneous controls with cross-layer coordination, enabling threats to be intercepted along their propagation paths while safeguarding critical assets. We detail the design rationale and architecture of five coordinated protection layers, and implement a plugin-native prototype on OpenClaw to demonstrate practical feasibility. This perspective provides a concrete blueprint for structuring runtime security controls, managing trust propagation, and enforcing execution containment in autonomous AI agents. Our code is available at https://github.com/FIND-Lab/AgentWard .

cs.CR

Quantized Zero-Energy RIS: Residual Phase Modeling and Outage Analysis

Zero-energy reconfigurable intelligent surfaces (zeRISs) have recently emerged as a promising solution for enabling energy-efficient and scalable programmable wireless environments (PWEs) by harvesting their operational energy from impinging radio-frequency signals. However, the operation of zeRIS-assisted systems is inherently constrained by the coupling between energy harvesting and signal reflection, a dependency that becomes more intricate under practical hardware limitations such as finite-resolution phase control. In this paper, we develop a comprehensive analytical framework for zeRIS-assisted communication systems operating under quantized phase shifts and harvest-and-reflect (HaR) schemes. Specifically, we analyze the joint energy-data rate outage probability and the energy efficiency under time switching and element splitting schemes, considering both transmitter-side and user-side deployment scenarios. By explicitly modeling the residual phase error induced by quantization and incorporating its statistical properties into the analysis, we show that quantization jointly affects energy harvesting and signal reflection, thereby inducing non-trivial trade-offs. As a result, the presented framework enables accurate performance evaluation and reveals critical design trade-offs for the selection of the phase resolution, and the applied HaR scheme in zeRIS-assisted wireless networks.

cs.IT

HabitatAgent: An End-to-End Multi-Agent System for Housing Consultation

Housing selection is a high-stakes and largely irreversible decision problem. We study housing consultation as a decision-support interface for housing selection. Existing housing platforms and many LLM-based assistants often reduce this process to ranking or recommendation, resulting in opaque reasoning, brittle multi-constraint handling, and limited guarantees on factuality. We present HabitatAgent, the first LLM-powered multi-agent architecture for end-to-end housing consultation. HabitatAgent comprises four specialized agent roles: Memory, Retrieval, Generation, and Validation. The Memory Agent maintains multi-layer user memory through internal stages for constraint extraction, memory fusion, and verification-gated updates; the Retrieval Agent performs hybrid vector--graph retrieval (GraphRAG); the Generation Agent produces evidence-referenced recommendations and explanations; and the Validation Agent applies multi-tier verification and targeted remediation. Together, these agents provide an auditable and reliable workflow for end-to-end housing consultation. We evaluate HabitatAgent on 100 real user consultation scenarios (300 multi-turn question--answer pairs) under an end-to-end correctness protocol. A strong single-stage baseline (Dense+Rerank) achieves 75% accuracy, while HabitatAgent reaches 95%.

cs.LG

A Novel Low-Complexity Dual-Domain Expectation Propagation Detection Aided AFDM for Future Communications

This paper presents a dual-domain low-complexity expectation propagation (EP) detection framework for affine frequency division multiplexing (AFDM) systems. By analyzing the structural properties of the effective channel matrices in both the time and affine frequency (AF) domains, our key observation is the domain-specific quasi-banded sparsity patterns, including AF-domain sparsity under frequency-selective channels and time-domain sparsity under doubly-selective channels. Based on these observations, we develop an AF-domain EP (EP-AF) detector for frequency-selective channels and a time-domain EP (EP-T) detector for doubly-selective channels, respectively. By performing iterative inference in the time domain using the Gaussian approximation, the proposed EP-T detector avoids inverting the dense channel matrix in the AF domain. Furthermore, the proposed EP-AF and EP-T detectors leverage the aforementioned quasi-banded sparsity of the AF domain and time domain channel matrices, respectively, to reduce the complexity of matrix inversion from cubic to linear order. Simulation results demonstrate that the proposed low-complexity EP-AF detector achieves nearly identical error rate performance to its conventional counterpart, while the proposed low-complexity EP-T detector offers an attractive trade-off between detection performance and complexity.

eess.SP

Antenna Elements' Trajectory Optimization for Throughput Maximization in Continuous-Trajectory Fluid Antenna-Aided Wireless Communications

Fluid antenna (FA) systems offer novel spatial degrees of freedom (DoFs) with the potential for significant performance gains. Compared to existing works focusing solely on optimizing FA positions at discrete time instants, we introduce the concept of continuous-trajectory fluid antenna (CTFA), which explicitly considers the antenna element's movement trajectory across continuous time intervals and incorporates the inherent kinematic constraints present in practical FA implementations. Accordingly, we formulate the total throughput maximization problem in CTFA-aided wireless communication systems, addressing the joint optimization of continuous antenna trajectories in conjunction with the transmit covariance matrices under kinematic constraints. To effectively solve this non-convex problem with highly coupled optimization variables, we develop an iterative algorithm based on block coordinate descent (BCD) and majorization-minimization (MM) principles with the aid of the weighted minimum mean square error (WMMSE) method. Finally, numerical results are presented to validate the efficacy of the proposed algorithms and to quantify the substantial total throughput advantages afforded by the conceived CTFA-aided system compared to conventional fixed-position antenna (FPA) benchmarks and alternative approaches employing simplified trajectories.

eess.SP

TheBotCompany: Self-Organizing Multi-agent Systems for Continuous Software Development

Large language model (LLM)-based multi-agent systems have shown promise in automating software development tasks. However, most vibe-coding systems focus on completing small tasks and incremental code changes, leaving persistent, continuous software development largely unexplored. We present TheBotCompany, an open-source orchestration framework for continuous multi-agent software development. TheBotCompany introduces three key innovations: (1) a three-phase state machine (Strategy to Execution to Verification) for milestone-driven development, (2) self-organizing agent teams where manager agents dynamically hire, assign, and retire worker agents based on project needs, and (3) asynchronous human oversight. We evaluate TheBotCompany on real-world software projects over multiple days of continuous development, measuring team adaptation patterns, milestone completion rates, cost efficiency, and code quality. Our results demonstrate that the self-organizing approach enables effective long-term software development with measurable progress, while the verification phase catches defects that would otherwise persist.

cs.SE

How Far Should We Need to Go : Evaluate Provenance-based Intrusion Detection Systems in Industrial Scenarios

Provenance-based Intrusion Detection Systems (PIDSes) have been widely used to detect Advanced Persistent Threats (APTs). Although many studies achieve high performance in the evaluations of their original papers, their performance in industrial scenarios remains unclear. To fill this gap, we conduct the first systematic evaluation and analysis of PIDSes in industrial scenarios. We first analyze the differences between the data from DARPA datasets and that collected in industrial scenarios, identifying three main new characteristics in industry: heterogeneous multi-source inputs, more powerful attackers, and increasing benign activity complexity. We then build several datasets to evaluate five state-of-the-art PIDSes. The evaluation results reveal challenges for existing PIDSes, including poor portability across different hosts and platforms, low detection performance against real-world attacks, and high false positive rates with ever-changing benign activities. Based on the evaluation results and our industrial practices, we provide several insights to solve or explain the above problems. For example, we propose a method to mitigate the high false positives, which reduces manual effort by 2/3. Finally, we propose several research suggestions to improve PIDSes.

cs.CR

PAPR-Aware Waveform Design for Energy-Efficient MIMO-OFDM SWIPT

Simultaneous wireless information and power transfer (SWIPT) critically depends on waveform design, which governs both reliable data delivery and efficient energy harvesting. Among waveform characteristics, the peak-to-average power ratio (PAPR) plays a pivotal role: low-PAPR signals improve power amplifier (PA) efficiency, while high-PAPR signals exploit rectifier nonlinearities to boost harvested energy. This duality makes PAPR a fundamental design challenge in SWIPT systems. To tackle this issue, we establish a unified analytical framework that characterizes the PAPR-dependent behaviors of both the PA and the rectifier, thereby revealing how waveform statistics determine end-to-end energy transfer efficiency. Building on this insight, we propose a frequency-domain resource allocation strategy for power-splitting SWIPT, where spectral segments are adaptively assigned to balance communication throughput with energy harvesting performance. Here, a key contribution is to extend SWIPT to MIMO-OFDM architectures. Despite concerns over excessive PAPR in large-scale antenna-subcarrier configurations, we demonstrate that appropriate waveform adaptation and resource optimization can transform MIMO-OFDM into an energy-efficient platform for joint data and power transfer. Finally, simulation results confirm significant improvements in PA efficiency, rectifier output, and overall energy transfer, thereby validating the practical benefits of the proposed approach.

eess.SP

Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats

Autonomous Large Language Model (LLM) agents, exemplified by OpenClaw, demonstrate remarkable capabilities in executing complex, long-horizon tasks. However, their tightly coupled instant-messaging interaction paradigm and high-privilege execution capabilities substantially expand the system attack surface. In this paper, we present a comprehensive security threat analysis of OpenClaw. To structure our analysis, we introduce a five-layer lifecycle-oriented security framework that captures key stages of agent operation, i.e., initialization, input, inference, decision, and execution, and systematically examine compound threats across the agent's operational lifecycle, including indirect prompt injection, skill supply chain contamination, memory poisoning, and intent drift. Through detailed case studies on OpenClaw, we demonstrate the prevalence and severity of these threats and analyze the limitations of existing defenses. Our findings reveal critical weaknesses in current point-based defense mechanisms when addressing cross-temporal and multi-stage systemic risks, highlighting the need for holistic security architectures for autonomous LLM agents. Within this framework, we further examine representative defense strategies at each lifecycle stage, including plugin vetting frameworks, context-aware instruction filtering, memory integrity validation protocols, intent verification mechanisms, and capability enforcement architectures.

cs.CR

Understanding Human-AI Collaboration in Cybersecurity Competitions

Capture-the-Flag (CTF) competitions are increasingly becoming a testbed for evaluating AI capabilities at solving security tasks, due to the controlled environments and objective success criteria. Existing evaluations have focused on how successful AI is at solving CTF challenges in isolation from human CTF players. As AI usage increases in both academic and industrial settings, it is equally likely that human players may collaborate with AI agents to solve challenges. This possibility exposes a key knowledge gap: how do humans perceive AI CTF assistance; when assistance is provided, how do they collaborate and is it effective with respect to human performance; how do humans assisted by AI compare to the performance of fully autonomous AI agents on the same challenges. We address this gap with the first empirical study of AI assistance in a live, onsite CTF. In a study with 41 participants, we qualitatively study (i) how participants' perception, trust, and expectations shift before versus after hands-on AI use, and (ii) how participants collaborate with an instrumented AI agent. Moreover, we also (iii) benchmark four autonomous AI agents on the same fresh challenge set to compare outcomes with human teams and analyze agent trajectories. We find that, as the competition progresses, teams increasingly delegate larger subtasks to the AI, giving it more agency. Interestingly, CTF challenges solving rates are often constrained not by model's reasoning capabilities, but rather by the human players: ineffective prompting and poor context specification become the primary bottleneck. Remarkably, autonomous agents that self-direct their prompting and tool use bypass this bottleneck and outperform most human teams, coming in second overall in the competition. We conclude with implications for the future design of CTF challenges and for building effective human-in-the-loop AI systems for security.

cs.CR

"Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems

Large language model (LLM) agents are rapidly becoming trusted copilots in high-stakes domains like software development and healthcare. However, this deepening trust introduces a novel attack surface: Agent-Mediated Deception (AMD), where compromised agents are weaponized against their human users. While extensive research focuses on agent-centric threats, human susceptibility to deception by a compromised agent remains unexplored. We present the first large-scale empirical study with 303 participants to measure human susceptibility to AMD. This is based on HAT-Lab (Human-Agent Trust Laboratory), a high-fidelity research platform we develop, featuring nine carefully crafted scenarios spanning everyday and professional domains (e.g., healthcare, software development, human resources). Our 10 key findings reveal significant vulnerabilities and provide future defense perspectives. Specifically, only 8.6% of participants perceive AMD attacks, while domain experts show increased susceptibility in certain scenarios. We identify six cognitive failure modes in users and find that their risk awareness often fails to translate to protective behavior. The defense analysis reveals that effective warnings should interrupt workflows with low verification costs. With experiential learning based on HAT-Lab, over 90% of users who perceive risks report increased caution against AMD. This work provides empirical evidence and a platform for human-centric agent security research.

cs.HC

Automating Agent Hijacking via Structural Template Injection

Agent hijacking, highlighted by OWASP as a critical threat to the Large Language Model (LLM) ecosystem, enables adversaries to manipulate execution by injecting malicious instructions into retrieved content. Most existing attacks rely on manually crafted, semantics-driven prompt manipulation, which often yields low attack success rates and limited transferability to closed-source commercial models. In this paper, we propose Phantom, an automated agent hijacking framework built upon Structured Template Injection that targets the fundamental architectural mechanisms of LLM agents. Our key insight is that agents rely on specific chat template tokens to separate system, user, assistant, and tool instructions. By injecting optimized structured templates into the retrieved context, we induce role confusion and cause the agent to misinterpret the injected content as legitimate user instructions or prior tool outputs. To enhance attack transferability against black-box agents, Phantom introduces a novel attack template search framework. We first perform multi-level template augmentation to increase structural diversity and then train a Template Autoencoder (TAE) to embed discrete templates into a continuous, searchable latent space. Subsequently, we apply Bayesian optimization to efficiently identify optimal adversarial vectors that are decoded into high-potency structured templates. Extensive experiments on Qwen, GPT, and Gemini demonstrate that our framework significantly outperforms existing baselines in both Attack Success Rate (ASR) and query efficiency. Moreover, we identified over 70 vulnerabilities in real-world commercial products that have been confirmed by vendors, underscoring the practical severity of structured template-based hijacking and providing an empirical foundation for securing next-generation agentic systems.

cs.AI

How Many Pinching Antennas Are Enough?

Programmable wireless environments (PWEs) have emerged as a key paradigm for next-generation communication networks, aiming to transform wireless propagation from an uncontrollable phenomenon into a reconfigurable process that can adapt to diverse service requirements. In this framework, pinching-antenna systems (PASs) have recently been proposed as a promising enabling technology, as they allow the radiation location and effective propagation distance to be adjusted by selectively exciting radiating points along a dielectric waveguide. However, most existing studies on PASs rely on the idealized assumption that pinching-antenna (PA) positions can be continuously adjusted along the waveguide, while realistically only a finite set of pinching locations is available. Motivated by this, this paper analyzes the performance of two-state PASs, where the PA positions are fixed and only their activation state can be controlled. By explicitly accounting for the spatial discreteness of the available pinching points, closed-form analytical expressions for the outage probability and the ergodic achievable data rate are derived. In addition, we introduce the pinching discretization efficiency to quantify the performance gap between discrete and continuous pinching configurations, enabling a direct assessment of the number of PAs required to approximate the ideal continuous case. Finally, numerical results validate the analytical framework and show that near-continuous performance can be achieved with a limited number of PAs, offering useful insights for the design and deployment of PASs in PWEs.

cs.NI