arXiv ScienceSearch

arXiv subjects

Fei Sun

Publications and source records attributed to Fei Sun.

At least 19 recordsLinked to original sources

Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems

Large language model (LLM)-based agents have shown strong potential in solving complex tasks through multi-step reasoning, yet they remain vulnerable to execution failures. Accurate failure attribution is therefore critical for improving agent reliability. Existing topology- and spectrum-based methods exploit trajectory structures but often overlook fine-grained semantics, while LLM-based attribution methods capture semantic cues but suffer from long-context degradation over lengthy trajectories. To address these challenges, we propose DUOTRACE, a plug-and-play detection filter for LLM-based failure attribution. DUOTRACE follows a detect-before-attribute paradigm: it first detects anomalous executions and then supplies focused trajectory evidence to downstream LLM-based attribution methods. For effective VAE-based anomaly detection on agent trajectories, DUOTRACE integrates dual-view semantic-structural node representations, a Tree-LSTM-based trajectory encoder, and prefix-chain- and LLM-based data augmentation to handle heterogeneous nodes, hierarchical execution structures, and limited failure data. Experiments with six LLM-based attribution baselines show that DUOTRACE improves agent-level and step-level attribution accuracy by 8.7% and 7.0%, respectively.

cs.AI

Towards Faithful Simulation of Human Shopping Behavior

Simulating realistic user shopping behavior underpins offline evaluation and reinforcement learning in e-commerce scenarios. While recent LLM- and VLM-based simulators have made encouraging progress, reproducing a real browsing session remains difficult for two reasons. (i) Memory Challenge: a shopping session spans dozens of pages, yet existing agents either discard long-range observation histories, losing the evolving user state, or naively concatenate them, overwhelming the context window and even degrading simulation quality. (ii) Optimization Challenge: current user simulators are typically supervised to match each logged action via imitation or step-level rewards; the resulting sessions often display unrealistic patterns, such as over-exploration or excessive passivity, which per-step supervision can neither detect nor correct. To address the above challenges, we present RecVerse, a GUI-grounded simulation agent that perceives pages through screenshots and produces faithful multi-turn trajectories. For the memory challenge, RecVerse adopts a cognitive-inspired hierarchical memory: Working Memory for short-term focus, Episodic Memory for in-session traces, and Preference Memory for high-level intent, with memory updates treated as actions so that the agent adaptively learns when and what to memorize. For the optimization challenge, RecVerse is optimized with a trajectory-level RL objective that scores entire sessions, aligning both macro-level action-type distributions and micro-level shopping intent with real users. We further release USB (User Simulation Benchmark), an interactive e-commerce GUI trajectory dataset for multi-turn user simulation. Experiments show that RecVerse significantly outperforms existing baselines in both behavioral fidelity and intent consistency.

cs.IR

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of recorded steps underexplored. We introduce LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors. It provides independently scored human labels for the responsible role and earliest decisive root-cause step. The median trajectory contains 145 steps, and the strongest baseline reaches only 13.2% exact root-step accuracy. We further present Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions. Using the same backbone, benchmark instances, and scoring protocol, RCTA reaches 51.1% responsible-role accuracy and 24.1% exact root-step accuracy. These results highlight the need to evaluate responsible-role attribution and exact root-step localization as separate targets in long-trajectory failure diagnosis.

cs.AI

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.

cs.AI

Critical behavior and critical exponents of rotating QCD matter

We investigate the thermodynamic properties and critical behavior of rotating strongly interacting matter within the two-flavor Nambu--Jona-Lasinio (NJL) model in the mean-field approximation. The phase structure and the critical endpoint (CEP) are determined in the temperature--angular velocity \((T,\omega)\) plane. By analyzing the singular behavior of thermodynamic observables near the CEP, we extract the corresponding effective critical exponents characterizing the scaling behavior of the specific heat density, the rotational polarization discontinuity, the rotational susceptibility, and the critical-isotherm behavior of the rotational polarization. The obtained exponents approach the expected mean-field values and satisfy the corresponding scaling relations, indicating that the rotational degree of freedom does not alter the underlying mean-field critical scaling behavior within the present framework. These results provide a systematic characterization of rotation-induced critical phenomena and establish a basis for further studies of rotating QCD matter beyond the mean-field approximation.

hep-ph

Breaking the trade-off between invisibility and sensitivity in electromagnetic sensing

Weak electromagnetic signals demand highly sensitive sensors, yet increasing a sensor's sensitivity inevitably strengthens its interaction with the surrounding field, producing scattering that perturbs the very signals being measured. Conversely, existing cloaking strategies suppress scattering only by isolating the sensor from incident waves, thereby compromising signal reception. Resolving this long-standing trade-off between invisibility and sensitivity has remained an outstanding challenge. Here we overcome this dilemma through an integrated transformation-optical architecture that co-designs the entire sensing system, including the electrically large sensor body, the subwavelength sensing probe, and their electrical interconnection. The proposed multifunctional core-shell structure guides incident waves around the sensor body while simultaneously concentrating them into the sensing region without disturbing the external electromagnetic field. A deep-subwavelength aperture preserves electrical connectivity without degrading either cloaking or field concentration, enabling invisible sensing within a single platform. A microwave prototype based on practical optic-null-medium metamaterials experimentally demonstrates broadband scattering suppression exceeding 3 dB together with an average sixfold enhancement of the detected signal over 4.9-5.1 GHz. By simultaneously eliminating measurement-induced field perturbation and amplifying the local sensing field, our approach establishes a general framework for invisible yet highly responsive electromagnetic sensors, opening new opportunities for weak-signal detection in biomedical diagnostics, secure communications, quantum technologies, and deep-space exploration.

physics.optics

Topological-Charge-Enabled Photonic Doping in ENZ Media

Conventional photonic doping schemes predominantly employ circular or rectangular dielectric dopants with zero topological charge, where the effective permeability can only be tuned through material selection and geometric scaling, resulting in limited design flexibility. In this work, topological structures are introduced into dielectric dopants by embedding internal holes to generate nonzero topological charge. Based on this concept, a theoretical model is established to describe the effective permeability of photonic doping systems with nonzero topological charge, and the underlying mechanisms governing topological-charge-dependent transmission are systematically elucidated. The results demonstrate that engineering nonzero topological charge through the number, shape, size and position of internal holes within dielectric dopants enables flexible manipulation of the internal magnetic field distributions, thereby providing precisely control over the effective permeability, as well as the resonance frequency and spectral linewidth of the transmission spectrum. The proposed multi-dimensional photonic doping strategy, integrating topological-charge engineering with geometric design, substantially enriches the available degrees of freedom for dispersion engineering and provides a versatile platform for advanced functional photonic devices.

physics.optics

optimal credit portfolio and consumption with regime switching and default contagion

We study optimal portfolio and consumption in a regime-switching multi-name credit market with default contagion. Defaults generate portfolio losses and alter the intensities of surviving securities. Under Cobb--Douglas utility, homogeneity reduces the HJB equation to a recursive ODE system indexed by the default states. Solving it backward from the all-default state, we establish existence and uniqueness of positive classical solutions, characterize the optimal feedback controls, and prove a verification theorem.

q-fin.MF

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical steps, misuse methods, or overstate conclusions. We evaluated whether autonomous access to a medical research skill package was associated with higher-quality AI-generated transcriptomic research-analysis outputs compared with native AI without skills. Methods. We conducted an exploratory multi-model human evaluation using a non-small cell lung cancer immunotherapy biomarker task. Six model backbones were tested. The evaluation included 21 anonymized outputs: 9 native-AI outputs and 12 skill-augmented outputs generated through an AI agent implementation represented by OpenClaw. Four non-expert biomedical reviewers and two blinded experts evaluated each output, with two ratings from each reviewer type. The primary outcome was expert-rated overall quality. Results. Skill-augmented outputs showed directionally higher expert overall quality than native-AI outputs (mean 5.50 vs 5.11; difference=0.39; bootstrap 95\% CI, -0.04 to 0.90; Welch p=0.156). Non-expert reviewer quality showed the same direction (mean 4.72 vs 4.47; difference=0.26; bootstrap 95\% CI, -0.25 to 0.80; Welch p=0.373). Expert agreement was limited (single-rating ICC=-0.15), and model-specific effects were descriptive and heterogeneous. Conclusions. Autonomous skill access showed a directional quality signal in this exploratory sample, but the signal was smaller than expert-rating noise and should not be interpreted as confirmatory evidence. The findings primarily motivate larger evaluations of skill-augmented AI agents with stronger reliability controls, platform replication, and biological-validity assessment.

cs.AI

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity

Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framework for mitigating shortcut learning in reward model training. Unlike static shortcut heuristics, DynaCF measures shortcut sensitivity online during optimization by applying semantics-preserving counterfactual perturbations and tracking the resulting margin shifts and preference flips under the current model. Samples with higher shortcut sensitivity are dynamically downweighted in the Bradley-Terry objective, encouraging the model to rely less on superficial patterns and more on task-relevant preference signals. Extensive experiments show that DynaCF consistently improves robustness in preference modeling.

cs.LG

SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization

Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features still remains a central challenge. Current explanation methods, however, typically operate within an open-loop paradigm, failing to leverage mechanistic feedback for further refinement. In this paper, we propose SAEExplainer, a training framework utilizes activation scores as an objective reward signal to train the model for self-correction and iterative bootstrapping. By iteratively verifying and correcting foundational explanations through a two-round optimization process, SAEExplainer achieves continuous improvement in its explanatory capabilities. This mechanism significantly reduces explanation hallucinations and reinforces causal triggering patterns. Extensive experiments demonstrate our approach improves upon established baselines across most metrics, especially in causal triggering and discriminative activation.

cs.CL

RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning

Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations, often treating a single expert trajectory as the target behavior. However, reasoning is not simple path imitation: rigidly following one demonstrated solution may overfit to surface forms and suppress the model's own reasoning distribution. We propose Rollout-Adaptive Supervised Fine-Tuning (RASFT), a policy-aware SFT framework that calibrates expert supervision according to problem-level solvability estimated from verified on-policy rollouts. For each problem, RASFT strengthens expert guidance when the current policy struggles, while relaxing rigid imitation and incorporating correct self-generated trajectories when the model already exhibits reliable reasoning behavior. To preserve useful reasoning priors, RASFT further introduces a clipped inverse ratio between the frozen reference model and the current policy to constrain excessive policy drift. Experiments across multiple models on six mathematical reasoning benchmarks and two code reasoning benchmarks show that RASFT achieves better overall performance than SFT, SFT variants, and representative RL methods. The code is available at https://github.com/zjd1sq/RASFT.

cs.LG

Agent System Operations: Categorization, Challenges, and Future Directions

As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional systems, garnering increasing attention. However, despite the widespread research interest and industrial application of agent systems, these systems, like their traditional counterparts, frequently encounter anomalies. These anomalies lead to instability and insecurity, hindering their further development. Therefore, a comprehensive and systematic approach to the operation and maintenance of agent systems is urgently needed. Unfortunately, current research on the operations of agent systems is sparse. To address this gap, we have undertaken a survey on agent system operations with the aim of establishing a clear framework for the field, defining the challenges, and facilitating further development. Specifically, this paper begins by systematically defining anomalies within agent systems, categorizing them into intra-agent anomalies and inter-agent anomalies. Next, we introduce a novel and comprehensive operational framework for agent systems, dubbed Agent System Operations (AgentOps). We provide detailed definitions and explanations of its four key stages: monitoring, anomaly detection, root cause localization, and resolution.

cs.MA

Cloaking of Arbitrarily Shaped Large-Scale Objects Through the Injection of Electromagnetic Invisibility Genes

Full-space electromagnetic invisibility mainly includes light-bending and scattering-cancellation cloaking. Light-bending cloaking causes double-blind phenomenon and is incompatible with sensing, while scattering-cancellation cloaking allows signal interaction and is more suitable for sensors and communication systems. However, traditional scattering-cancellation cloaking depends highly on target shape and size, making it difficult to realize cloaking for irregular, inhomogeneous and electrically large objects. To solve these problems, this work proposes an electromagnetic invisibility gene injection strategy inspired by biological camouflage. Objects are decomposed into subwavelength units, and customized invisibility genes are injected into each unit according to electromagnetic parameters to achieve overall scattering cancellation. Simulations and microwave experiments verify that this method can realize efficient cloaking for objects with arbitrary shapes, dielectric constants from 2 to 10, and different unit morphologies. This strategy breaks the limits of traditional cloaking and provides a universal, flexible scheme for practical applications such as antenna supports and electromagnetic transparent covers.

physics.optics

Decoupling heat and electricity: A thermal invisible gateway

The Wiedemann-Franz law couples electrical and thermal conductivity, making high electrical conduction with low thermal conduction a major challenge. To overcome this, we designed an active thermal metasurface (ATMS) - based thermal invisible gateway that decouples thermal and electrical paths. Built on a copper substrate with a dumbbell-shaped bridge, the structure suppresses heat flow via directional compensation while allowing unimpeded electrical conduction. Room-temperature experiments show an effective thermal conductivity below 10^-3 W m^-1 K^-1 (near zero, air-like insulation) and an electrical conductivity up to 2.8x10^7 S m^-1 (metal-level). Unlike conventional material-modification approaches, our work uses macroscopic structural design to break the intrinsic coupling, offering a promising solution for applications like on-chip interconnects and wearable electronics.

physics.app-ph

Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs

Reinforcement learning (RL) has achieved remarkable success in LLM reasoning, but whether it can also improve direct recall of parametric knowledge remains an open question. We study this question in a controlled zero-shot, one-hop, closed-book QA setting with no chain-of-thought, training only on binary correctness rewards and applying fact-level train-test deduplication to ensure gains reflect improved recall rather than reasoning or memorization. Across three model families and multiple factual QA benchmarks, RL yields ~27% average relative gains, surpassing both training- and inference-time baselines alike. Mechanistically, RL primarily redistributes probability mass over existing knowledge rather than acquiring new facts, moving correct answers from the low-probability tail into reliable greedy generations. Our data-attribution study reveals that the hardest examples are the most informative: those whose answers never appear in 128 pre-RL samples (only ~18% of training data) drive ~83% of the gain, since rare correct rollouts still emerge during training and get reinforced. Together, these findings broaden the role of RL beyond reasoning, repositioning it as a tool for unlocking rather than acquiring latent parametric knowledge.

cs.CL

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

Background: Agent skills are increasingly deployed as modular, reusable capability units in AI agent systems. Medical research agent skills require safeguards beyond general-purpose evaluation, including scientific integrity, methodological validity, reproducibility, and boundary safety. This study developed and preliminarily evaluated a domain-specific audit framework for medical research agent skills, with a focus on reliability against expert review. Methods: We developed MedSkillAudit (skill-auditor@1.0), a layered framework assessing skill release readiness before deployment. We evaluated 75 skills across five medical research categories (15 per category). Two experts independently assigned a quality score (0-100), an ordinal release disposition (Production Ready / Limited Release / Beta Only / Reject), and a high-risk failure flag. System-expert agreement was quantified using ICC(2,1) and linearly weighted Cohen's kappa, benchmarked against the human inter-rater baseline. Results: The mean consensus quality score was 72.4 (SD = 13.0); 57.3% of skills fell below the Limited Release threshold. MedSkillAudit achieved ICC(2,1) = 0.449 (95% CI: 0.250-0.610), exceeding the human inter-rater ICC of 0.300. System-consensus score divergence (SD = 9.5) was smaller than inter-expert divergence (SD = 12.4), with no directional bias (Wilcoxon p = 0.613). Protocol Design showed the strongest category-level agreement (ICC = 0.551); Academic Writing showed a negative ICC (-0.567), reflecting a structural rubric-expert mismatch. Conclusions: Domain-specific pre-deployment audit may provide a practical foundation for governing medical research agent skills, complementing general-purpose quality checks with structured audit workflows tailored to scientific use cases.

cs.AI

Heavy-quark transport across the QCD crossover driven by a lattice-constrained in-medium potential

We present a self-consistent framework for heavy-quark transport in the quark-gluon plasma across the QCD crossover region. By synthesizing perturbative and nonperturbative interactions into a unified interaction kernel, we circumvent the traditional reliance on arbitrary soft-hard momentum separation scales. The interaction is governed by an in-medium effective potential, incorporating short-range Yukawa screening and long-range confining string contributions, both rigorously constrained by the latest lattice QCD data. Our results reveal that the nonperturbative string tension is indispensable for capturing the extreme opacity of the medium near the critical temperature $T_c$. Specifically, our model predicts a spatial diffusion coefficient of $2\pi T D_s \approx 0.5 \sim 1.7$, demonstrating a striking quantitative agreement with the recent lattice QCD extractions. Ultimately, our results provide a robust dynamical interpretation of the strong heavy-quark coupling near the QCD crossover and offer a unified framework for describing heavy-flavor transport in hot and dense QCD matter.

hep-ph