arXiv ScienceSearch

arXiv subjects

Juan Wang

Publications and source records attributed to Juan Wang.

At least 19 recordsLinked to original sources

Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents

Large language model (LLM)-driven embodied agents rely on environment states to interpret scenes, generate high-level plans, and drive physical execution, making planner-visible state representations a critical security boundary. Existing attacks primarily manipulate user instructions, prompt contexts, model behavior, or perceptual inputs, while paying limited attention to whether environment-state text itself can serve as deceptive task evidence and propagate beyond planning to affect execution outcomes. Because embodied tasks are constrained by entity grounding, action preconditions, spatial relations, and environmental constraints, planning deviation alone does not guarantee adversarial execution. To address this gap, we investigate environment-state text as an independent attack surface and present the first closed-loop Environment State-Text Injection (ESTI) attack for LLM-driven embodied agents. Without modifying the original user instruction, model parameters, or executor, ESTI reformulates an adversarial objective as false state evidence compatible with the current environment and influences planning and execution through object properties, spatial relations, affordances, task-stage rules, and execution feedback. We further develop ESTI-Bench to evaluate attack propagation across the planning-to-execution closed loop and compare ESTI with Vanilla IPI, EIRAD, and BADROBOT across ProgPrompt/VirtualHome, VoxPoser/RLBench, and AI2-THOR/iTHOR. ESTI consistently outperforms existing baselines, improving planning-level and execution-level attack success rates by up to 89.32\% and 43.69\%, respectively. Further analysis shows that grounding, consistency, and executability jointly determine whether manipulated state evidence can propagate through the embodied closed loop and produce verifiable environmental changes.

cs.RO

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks, prompt injection, backdoors, poisoning, or adversarial examples, but these categories do not consistently identify where an adversary first enters the embodied control loop. We present a trust-boundary-centric survey of foundation-model-powered embodied-agent security. Using a first-compromised-trust-boundary principle, we separate attack surface from attack mechanism and organize the system into five layers and twelve attack surfaces spanning the model supply chain, user instructions, context and memory, physical semantic environments, multimodal perception, world state, internal reasoning, task planning, action interfaces, middleware, multi-agent communication, and execution control. Based on 58 attack records and 61 defense records collected through August 15, 2026, we analyze representative attacks, cross-layer propagation, defense placement, and evaluation practices. Our quantitative analysis shows that attack research is concentrated on multimodal perception and action interfaces, while defenses are especially concentrated on action-level and runtime protection. Context and long-term memory, middleware and networking, world-state integrity, and multi-agent trust remain comparatively underexplored. We conclude with open challenges in state provenance, compositional defenses, long-horizon attack propagation, physical realizability, Byzantine multi-robot behavior, and unified closed-loop evaluation.

cs.RO

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing ``verification'' signals are not diagnostic: answer matching observes only the outcome, LLM-as-judge provides subjective and non-verifiable critiques, and scalar rewards (e.g., PRMs/RMs) offer little insight into where a multi-step derivation fails.We propose \textbf{SymDiag}, a neuro-symbolic framework that \textbf{reframes reasoning verification as structured failure diagnosis}. SymDiag translates natural-language CoT into symbolic constraints and performs step-level satisfiability/entailment checks to (i) localize failing steps and (ii) produce verifiable diagnostic evidence, including counterexamples, inconsistency witnesses, and missing-premise indicators. A central challenge is that apparent ``logic violations'' can be caused either by genuine reasoning defects or by neural-to-symbolic translation noise. SymDiag therefore incorporates a Self-Auditor that disentangles TranslationError from ReasoningError via dual symbolic encodings consistency checks, enabling robust diagnosis under partial observability. Across diverse mathematical, logical, scientific, and general reasoning benchmarks, SymDiag improves detection of unfaithful reasoning and provides substantially more effective feedback for multi-round reasoning repair than outcome-only verification and LLM-based judging, offering a principled foundation for trustworthy and scalable reasoning diagnosis.

cs.AI

LLM agents security duality: a comprehensive survey of self-security and empowered cybersecurity

Large language model (LLM) agents are rapidly being integrated into real-world systems. Their autonomy and tool-use capabilities generate substantial value while simultaneously expanding the security attack surface. This survey provides a comprehensive overview of the opportunities and challenges of LLM agents in security, focusing on two core areas: (1) threats to LLM agents themselves and corresponding mitigation strategies (LLM agents self-security), and (2) the role of LLM agents in empowering the cybersecurity lifecycle across offense and defense (LLM agents empowered cybersecurity). We first examine the internal and external attack surfaces of agents, propose a taxonomy organized by threat sources, and analyze associated mitigations and evaluation frameworks. We then investigate how agent capabilities are applied in cybersecurity practice and present, to our knowledge, the first agent-empowerment framework aligned with the full cyber offense-defense lifecycle. By systematically surveying these two areas, we are the first to highlight a positive feedback synergy between LLM agents self-security and empowered cybersecurity, offering new insights for the advancement of both. We further identify current limitations and outline promising directions for future research. The insights provided aim to catalyze the coordinated development of LLM agents self-security and agent empowered cybersecurity, paving the way for more capable and robust agent applications.

cs.CR

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding

World Action Models (WAMs) extend robot policy learning by incorporating future prediction as an additional training objective, encouraging the policy to encode task-relevant temporal structure in its representations. Current WAMs often rely on large-scale generative architectures that incur high training costs and inference latency, making them difficult to deploy as efficient closed-loop policies. We propose Light-WAM, a lightweight World Action Model for efficient robot manipulation. Specifically, it is built with a compact video backbone and performs future-video supervision in a downsampled latent space, reducing the cost of video co-training while retaining its benefits for representation learning. For action prediction, Light-WAM introduces the StateFusionActionExpert, which reads adapted states from multiple backbone layers, fuses them through learned-query pooling, and directly predicts action chunks in a single forward pass. This design provides an efficient interface between video backbone representations and robot actions, avoiding the need for heavy generative action experts. Experiments demonstrate that Light-WAM maintains strong performance on LIBERO and achieves usable multi-task performance on RoboTwin 2.0, while using only 0.44B trainable parameters. It also achieves 72.03ms inference latency with 4.1GiB peak GPU memory and improved training throughput.

cs.CV

Impact of $N^*$ and $\Lambda^*$ resonances on $CP$ violation in $\Lambda_b^0$ decays

The four-body decay $\Lambda_b^0\to pK^-\pi^+\pi^-$ has led to the first observation of baryonic $CP$ violation. However, the underlying subprocesses $\Lambda_b^0\to N^* M$ and $\Lambda_b^0\to \Lambda^* M$, as well as the roles of excited nucleon ($N^*$) and hyperon ($\Lambda^*$) resonances, remain largely unexplored. Within the constituent quark model, we identify the relevant resonant states contributing to these underlying two-body transitions, including $N(1535)$, $N(1520)$, $\Lambda(1670)$, $\Lambda(1690)$, together with the remaining $1P$-wave baryon states. We obtain the resonant branching fraction ${\cal B}(\Lambda_b^0\to pK^-\pi^+\pi^-) =(30.0^{+2.8+4.0}_{-1.3-3.4}\pm1.8)\times10^{-6}$, while the resulting ${\cal A}_{CP}(\Lambda_b^0\to pK^-\pi^+\pi^-)=(3.18\pm0.11\pm0.13\pm0.11)\%$ provides a natural interpretation of the first observed baryonic $CP$ asymmetry. Our analysis establishes the first comprehensive framework for quantifying the impact of excited baryon resonances in multi-body beauty-baryon decays, with the associated mechanism generally applicable to baryonic $CP$ asymmetries.

hep-ph

IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts

Current image quality assessment methods are heavily biased towards global distortions (e.g., noise, blur), neglecting local perceptual artifacts such as ghosting, lens flare, and moire effects. Although significant progress has been made in artifact removal, the fundamental problem of automatic artifact detection remains largely unexplored. In this paper, we formalize the Image Perceptual Artifact Detection (IPAD) task to address this gap. We contribute a benchmark dataset comprising 3,520 artifact images, including 520 real-captured and 3,000 synthetic samples, each paired with pixel-level masks across three representative artifact categories. The core challenge of IPAD lies in the localized, subtle, and semantically weak nature of these artifacts, which makes them prone to missed detection. To overcome this, we introduce IPAD-CLIP, a novel framework built upon CLIP that enhances artifact discrimination in both textual and visual spaces while preserving generalization capabilities. Our key insight is that local artifacts often exhibit strong correlations with specific semantic contexts. Accordingly, we learn artifact-aware text embeddings to explicitly model the object-artifact relationships, resulting in enhanced representations that clear differentiate between clean and artifact prompts. These text embeddings are then used as anchors to shift the visual encoder's attention from high-level semantics to subtle, low-level artifacts. Extensive experiments demonstrate that IPAD-CLIP offers a resource-efficient adaptation of CLIP for detection, significantly outperforming advanced image anomaly detection and manipulation detection methods on our benchmark. To the best of our knowledge, this is the first study addressing multi-class local perceptual artifact detection in terms of both dataset and model.

cs.CV

Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

Multi-Hop Fact Verification requires complex reasoning across disparate evidence, posing significant challenges for Large Language Models , which may suffer from hallucinations and fractured logical chains. Existing methods, while improving transparency via Chain-of-Thought , often lack explicit modeling of the structural dependencies between evidence and claims. In this work, we introduce an SCM-inspired framework that grounds reasoning in explicit directed dependency graphs, treating verification as a constructive structural reasoning process rather than full causal inference with interventions or counterfactual semantics. We empirically identify an "inverted U-shaped" correlation between reasoning-chain length and accuracy, revealing that excessive structural complexity can degrade performance. To address this, we propose a rule-based reinforcement learning strategy using Group Relative Policy Optimization. This approach dynamically optimizes the trade-off between structural depth and conciseness. Extensive experiments on HoVer and EX-FEVER demonstrate that our SCM-GRPO framework outperforms strong baselines while producing more traceable reasoning structures for complex fact verification.

cs.AI

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Professional Image Quality Assessment (Track 1)

In this paper, we present an overview of the NTIRE 2026 challenge on the 3rd Restore Any Image Model in the Wild, specifically focusing on Track 1: Professional Image Quality Assessment. Conventional Image Quality Assessment (IQA) typically relies on scalar scores. By compressing complex visual characteristics into a single number, these methods fundamentally struggle to distinguish subtle differences among uniformly high-quality images. Furthermore, they fail to articulate why one image is superior, lacking the reasoning capabilities required to provide guidance for vision tasks. To bridge this gap, recent advancements in Multimodal Large Language Models (MLLMs) offer a promising paradigm. Inspired by this potential, our challenge establishes a novel benchmark exploring the ability of MLLMs to mimic human expert cognition in evaluating high-quality image pairs. Participants were tasked with overcoming critical bottlenecks in professional scenarios, centering on two primary objectives: (1) Comparative Quality Selection: reliably identifying the visually superior image within a high-quality pair; and (2) Interpretative Reasoning: generating grounded, expert-level explanations that detail the rationale behind the selection. In total, the challenge attracted nearly 200 registrations and over 2,500 submissions. The top-performing methods significantly advanced the state of the art in professional IQA. The challenge dataset is available at https://github.com/narthchin/RAIM-PIQA, and the official homepage is accessible at https://www.codabench.org/competitions/12789/.

cs.CV

Nonvolatile single-ion memory with picosecond switching

The rapid development of artificial intelligence (AI), Internet of Things (IoT), and edge computing applications has posed severe challenges to conventional memory technologies in terms of density, speed, and energy consumption. Herein, a single-ion transport mechanism is proposed to achieve picosecond (ps) switching capability. For monolayer hexagonal boron nitride (h-BN) with single-atom vacancy defects, first-principles calculations reveal that single-ion penetration across the BN plane dominates the resistive switching. The trapping and release of a single ion correspond to different states of the memory device for one bit of information. Experimentally fabricated single-ion memory exhibits nonvolatile resistive switching with ultra-fast switching speed of 20 ps and ultra-low energy consumption of 310 aJ/bit. This high performance is attributed to the extremely short distance for the single ion to travel through. Such devices pave the way for the realization of high-performance nonvolatile memory with ultra-fast speed, ultra-low energy consumption, and high storage density, that is called the "Unified Memory" long desired by the whole industry.

physics.app-ph

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Multi-Exposure Image Fusion in Dynamic Scenes (Track 2)

This paper presents NTIRE 2026, the 3rd Restore Any Image Model (RAIM) challenge on multi-exposure image fusion in dynamic scenes. We introduce a benchmark that targets a practical yet difficult HDR imaging setting, where exposure bracketing must be fused under scene motion, illumination variation, and handheld camera jitter. The challenge data contains 100 training sequences with 7 exposure levels and 100 test sequences with 5 exposure levels, reflecting real-world scenarios that frequently cause misalignment and ghosting artefacts. We evaluate submissions with a leaderboard score derived from PSNR, SSIM, and LPIPS, while also considering perceptual quality, efficiency, and reproducibility during the final review. This track attracted 114 participating teams and received 987 submissions. The winning methods significantly improved the ability to remove artifacts from multi-exposure fusion and recover fine details. The dataset and the code of each team can be found at the repository: https://github.com/qulishen/RAIM-HDR.

cs.CV

Investigating $\Omega_c$ spectroscopy in two-body $\Omega_b$ decays

We investigate the $1S$-, $1P$-, $2S$-, and $1D$-wave $\Omega_c(css)$ spectroscopy through the non-leptonic decays of $\Omega_b^-$ baryon within the constituent quark model. For the lowest-lying $1S$-wave $\Omega_c$ state, we obtain the branching fractions ${\cal B}(\Omega_b^- \to \Omega_c^0 \pi^-,\Omega_c^0 \rho^-)=(1.2,6.3) \times 10^{-3}$, which are consistent with existing model predictions. For $\Omega_c(3000)$, $\Omega_c(3050)$, $\Omega_c(3065)$, and $\Omega_c(3090)$, observed in the proton-proton and $e^+ e^-$ collisions and interpreted as members of the $1P$-wave multiplet (collectively denoted as $\Omega_c^{**}$), we predict the branching fractions of $\Omega_b^-\to\Omega_c^{**}\pi^-,\Omega_c^{**}\rho^-$ at the level of $10^{-3}$. Assigning the newly observed $\Omega_c(3327)$ baryon to a $1D$-wave excitation with $J^P=5/2^+$ or $7/2^+$, we obtain ${\cal B}[\Omega_b^-\to \Omega_c(3327)^0\pi^-,\Omega_c(3327)^0\rho^-] =(2.0,5.7)\times 10^{-3}$ or $(4.6,0.8)\times 10^{-3}$, respectively. The pronounced differences between these two scenarios provide a clear discriminant that can be tested in future measurements at LHCb.

hep-ph

AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification

Large language model (LLM) agents increasingly rely on external tools and retrieval systems to autonomously complete complex tasks. However, this design exposes agents to indirect prompt injection (IPI), where attacker-controlled context embedded in tool outputs or retrieved content silently steers agent actions away from user intent. Unlike prompt-based attacks, IPI unfolds over multi-turn trajectories, making malicious control difficult to disentangle from legitimate task execution. Existing inference-time defenses primarily rely on heuristic detection and conservative blocking of high-risk actions, which can prematurely terminate workflows or broadly suppress tool usage under ambiguous multi-turn scenarios. We propose AgentSentry, a novel inference-time detection and mitigation framework for tool-augmented LLM agents. To the best of our knowledge, AgentSentry is the first inference-time defense to model multi-turn IPI as a temporal causal takeover. It localizes takeover points via controlled counterfactual re-executions at tool-return boundaries and enables safe continuation through causally guided context purification that removes attack-induced deviations while preserving task-relevant evidence. We evaluate AgentSentry on the \textsc{AgentDojo} benchmark across four task suites, three IPI attack families, and multiple black-box LLMs. AgentSentry eliminates successful attacks and maintains strong utility under attack, achieving an average Utility Under Attack (UA) of 74.55 %, improving UA by 20.8 to 33.6 percentage points over the strongest baselines without degrading benign performance.

cs.CR

Spin Faraday Waves in Periodically Modulated Spin-Orbit-Coupled Bose Gases

This paper investigates the formation of Spin Faraday waves in spin-orbit-coupled Bose-Einstein condensate under the stripe phase and explores the dispersion relation under three different phases. We discover that the SFW exhibit temporal and spatial patterns when the interaction is modulated periodically, and appear with resonant waves and higher order harmonics. SFW can be excited even when the modulation frequency resonates with the trap frequency. Furthermore, we study the dispersion relation of these Faraday modes through periodic modulation, which agrees well with our theoretical results under three quantum phases. Our work indicates novel physical phenomena originating from the introduction of spin-orbit coupling and provides a possible method for studying the dispersion of Bose gases.

cond-mat.quant-gas

Quenching dynamics of vortex in spin-orbit coupled Bose-Einstein condensates

We investigate the ground states and rich dynamics of vortices in spin-orbit coupled Bose-Einstein condensates (BEC) subject to position-dependent detuning. Such a detuning plays the role of an effective rotational frequency, causing the generation of a synthetic magnetic field. Through scanning the detuning gradient, we numerically obtain static vortex lattice structures containing 1 to 6 vortices using the coupled Gross-Pitaevskii equations. When quenching detuning gradient below its initial value, the vortex lattices exhibit interesting periodic rotation motion, and their dynamical stability can persist for up to 1000ms. In particular, depending on the detuning gradient, the twin vortices exhibit either a scissors-like rotational oscillation or a clockwise periodic rotation, reflecting the response to the magnetic field gradient experienced by the condensates. We fit the numerical results to quantitatively analyze the relation between rotation dynamics and magnetic field gradients. When quenching the detuning gradient beyond its initial value, additional vortices appear. Our findings may motivate further experimental studies of vortex dynamics in synthetic magnetic fields and offer insights for engineering a BEC-based magnetic field gradiometer.

cond-mat.quant-gas

Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking

3D instance segmentation is an important task for real-world applications. To avoid costly manual annotations, existing methods have explored generating pseudo labels by transferring 2D masks from foundation models to 3D. However, this approach is often suboptimal since the video frames are processed independently. This causes inconsistent segmentation granularity and conflicting 3D pseudo labels, which degrades the accuracy of final segmentation. To address this, we introduce a Granularity-Consistent automatic 2D Mask Tracking approach that maintains temporal correspondences across frames, eliminating conflicting pseudo labels. Combined with a three-stage curriculum learning framework, our approach progressively trains from fragmented single-view data to unified multi-view annotations, ultimately globally coherent full-scene supervision. This structured learning pipeline enables the model to progressively expose to pseudo-labels of increasing consistency. Thus, we can robustly distill a consistent 3D representation from initially fragmented and contradictory 2D priors. Experimental results demonstrated that our method effectively generated consistent and accurate 3D segmentations. Furthermore, the proposed method achieved state-of-the-art results on standard benchmarks and open-vocabulary ability.

cs.CV

COCONUT: A coronal model with an energy decomposition strategy

In this paper, we propose an energy decomposition method combined with an HLL Riemann solver that includes an additional dissipation term in the energy equation to improve the numerical stability of the fully implicit, time-evolving coronal model COCONUT and extend its applicability to solar-maximum phases. In MHD simulations that evolve conservative variables in time, the thermal pressure is typically computed by subtracting the magnetic and kinetic energies from the total energy. In low-beta (the ratio of thermal to magnetic pressure; $< 10^{-3}$) regions, discretization errors of magnetic energy can be comparable to the thermal pressure, potentially leading to negative thermal pressure and causing the simulation to crash. Therefore, we update the decomposed energy, excluding the magnetic energy, at each time step. It avoids subtracting a large magnetic energy from the total energy to obtain a very small thermal pressure in low-$\beta$ regions, thereby improving the numerical stability of MHD models. We validate the algorithm using a time-evolving solar-maximum Carrington rotation simulation in 2025, which the previous code failed to run to completion. We also perform quasi-steady-state coronal simulations and 2D benchmark tests to further assess the algorithm's performance. The simulation results show that the algorithm produces results nearly identical to those obtained using the traditional full energy equation during solar minimum, while significantly improving COCONUT's ability to simulate coronal evolution under strong magnetic fields, even including fields exceeding 100 Gauss with $\beta<10^{-3}$. This method provides a promising approach for performing quasi-realistic coronal simulations during solar maxima.

astro-ph.SR