arXiv ScienceSearch

arXiv subjects

Yi Huang

Publications and source records attributed to Yi Huang.

At least 19 recordsLinked to original sources

Robust and Generalizable Background Subtraction on Images of Calorimeter Jets using Unsupervised Generative Learning

Accurate separation of signal from background is one of the main challenges for precision measurements across high-energy and nuclear physics. Conventional supervised learning methods are insufficient here because the required paired signal and background examples are impossible to acquire in real experiments. Here, we introduce an unsupervised unpaired image-to-image translation neural network that learns to separate the signal and background from the input experimental data using cycle-consistency principles. We demonstrate the efficacy of this approach using images composed of simulated calorimeter data from the sPHENIX experiment, where physics signals (jets) are immersed in the extremely dense and fluctuating heavy-ion collision environment. Our method outperforms conventional subtraction algorithms in fidelity and overcomes the limitations of supervised methods. Furthermore, we evaluated the model's robustness in an out-of-distribution test scenario designed to emulate modified jets as in real experimental data. The model, trained on a simpler dataset, maintained its high fidelity on a more realistic, highly modified jet signal. This work represents the first use of unsupervised unpaired generative models for full detector jet background subtraction and offers a path for novel applications in real experimental data, enabling high-precision analyses across a wide range of imaging-based experiments.

nucl-ex

Leveraging Imperfect Restoration for Data Availability Attack

The abundance of online data is at risk of unauthorized usage in training deep learning models. To counter this, various Data Availability Attacks (DAAs) have been devised to make data unlearnable for such models by subtly perturbing the training data. However, existing attacks often excel against either Supervised Learning (SL) or Self-Supervised Learning (SSL) scenarios. Among these, a model-free approach that generates a Convolution-based Unlearnable Dataset (CUDA) stands out as the most robust DAA across both SSL and SL. Nonetheless, CUDA's effectiveness against SSL is underwhelming and it faces a severe trade-off between image quality and its poisoning effect. In this paper, we conduct a theoretical analysis of CUDA, uncovering the sub-optimal gradients it introduces and elucidating the strategy it employs to induce class-wise bias for data poisoning. Building on this, we propose a novel poisoning method named Imperfect Restoration Poisoning (IRP), aiming to preserve high image quality while achieving strong poisoning effects. Through extensive comparisons of IRP with eight baselines across SL and SSL, coupled with evaluations alongside five representative defense methods, we showcase the superiority of IRP. Code: https://github.com/lyumingzhi/IRP

cs.AI

Flippered hyperbolic surfaces and renormalized volumes of their moduli spaces I

We introduce a natural class of hyperbolic surfaces called flippered surfaces that generalize crowned hyperbolic surfaces (i.e.: worldsheets for open strings). We develop their Teichmüller and moduli-space theory, construct generalized Weil-Petersson volume forms and Chekhov's action, and prove that the resulting generalized Mirzakhani volumes are finite. We establish three geometric recursion formulae-neck chopping, disk excision, and crown extraction-which express these volumes in terms of those of topologically simpler surfaces. For the fundamental polygonal and annular cases, we derive integral representations involving conical Legendre functions, as well as explicit formulae in terms of elliptic integrals and polylogarithms. We further describe the arithmetic structure of the Taylor coefficients of these volumes, show that suitable specializations are Kontsevich-Zagier periods, and prove identities at the imaginary boundary length $2π\sqrt{-1}$ that generalize the Do-Norbury paraphrasing of string and dilaton-type equations.

math.GT

No Data? No Problem: Synthesizing Security Graphs for Better Intrusion Detection

Provenance graph analysis plays a vital role in intrusion detection, particularly against Advanced Persistent Threats (APTs), by exposing complex attack patterns. While recent systems combine graph neural networks (GNNs) with natural language processing (NLP) to capture structural and semantic features, their effectiveness is limited by class imbalance in real-world data. To address this, we introduce PROVSYN, a novel hybrid provenance graph synthesis framework, which comprises three components: (1) graph structure synthesis via heterogeneous graph generation models, (2) textual attribute synthesis via fine-tuned Large Language Models (LLMs), and (3) five-dimensional fidelity evaluation. Experiments on six benchmark datasets demonstrate that PROVSYN consistently produces higher-fidelity graphs across the five evaluation dimensions compared to four strong baselines. To further demonstrate the practical utility of PROVSYN, we utilize the synthesized graphs to augment training datasets for downstream APT detection models. The results show that PROVSYN effectively mitigates data imbalance, improving normalized entropy by up to 0.35 in absolute terms, and enhances the generalizability of downstream detection models, yielding an absolute increase of up to 0.38 in balanced accuracy.

cs.CR

Pandora: Leveraging Code-driven Knowledge Transfer for Unified Structured Knowledge Reasoning

Unified Structured Knowledge Reasoning (USKR) aims to answer natural language questions by using structured sources such as tables, databases, and knowledge graphs in a unified way. Existing USKR methods rely on task-specific strategies or bespoke representations, which hinder their ability to dismantle barriers between different SKR tasks, thereby constraining their overall performance in cross-task scenarios. In this paper, we introduce \textsc{Pandora}, a novel USKR framework that addresses the limitations of existing methods by leveraging two key innovations. First, we propose a code-based unified knowledge representation using \textsc{Python}'s \textsc{Pandas} API, which aligns seamlessly with the pre-training of LLMs. This representation facilitates a cohesive approach to handling different structured knowledge sources. Building on this foundation, we employ knowledge transfer to bolster the unified reasoning process of LLMs by automatically building cross-task memory. By adaptively correcting reasoning using feedback from code execution, \textsc{Pandora} showcases impressive unified reasoning capabilities. Extensive experiments on six widely used benchmarks across three SKR tasks demonstrate that \textsc{Pandora} outperforms existing unified reasoning frameworks and competes effectively with task-specific methods.

cs.CL

Relay, Don't Route: Adaptive Population Handoff for Cost-Efficient LLM-Driven Evolution

Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly. A natural alternative is to combine cheap and strong models under a fixed inference budget. However, existing approaches typically allocate models at the level of individual queries or mutation steps, overlooking that evolutionary search is \textit{stateful}: each generated candidate changes the population from which subsequent mutations are produced. We empirically analyze LLM-driven evolutionary trajectories and find that search progress is strongly front-loaded, early trajectory performance is informative but noisy, and cheap models recover much of the early progress achieved by strong models at lower cost. Motivated by these findings, we propose \textbf{\model}, a training-free framework that shifts budget allocation from individual calls to evolving populations through adaptive \textit{population handoff}. A cheap model explores multiple trajectories in short blocks allocated by a bandit scheduler. Relay Gain, defined as the marginal improvement of a compact, quality-diverse candidate bank constructed for handoff, serves as the scheduler reward and determines when to hand off. The curated candidates initialize a shared strong model population for refinement. Across four benchmarks and three budgets, \model achieves the highest mean score in 11 of 12 settings, outperforming competitive baselines. Our results suggest that in stateful search, budget allocation should be organized around the population, not the individual call.

cs.CL

Stripe-tuned superconductivity in single-flavor metals with nontrivial quantum geometry

We study how the interplay between nontrivial quantum geometry and an applied stripe potential affects superconductivity in a two-dimensional single-flavor metal. Assuming a weak contact attractive interaction and focusing on the lowest subband in the presence of a strong stripe potential, we analytically derive two possible pairing states in the quasi-one-dimensional limit. In addition to the conventional longitudinal $p_y$-wave order (with the stripes along the $y$ direction), we find that an exotic transverse $p_x$-wave order can be stabilized. The competition between these two orders is controlled by the electron density of each stripe and the Berry-curvature-dressed interaction. Notably, the transverse $p_x$ wave order develops a nodal line at $k_x=0$, while the longitudinal $p_y$ order is fully gapped. We discuss the possible experimental probes distinguishing these orders. Our results establish a way of controlling the pairing symmetry through a stripe potential, predicting superconductivity with nontrivial quantum geometry.

cond-mat.supr-con

Coupling dynamical accretion and chemical differentiation: A unified framework for the diversity of Earth and Mars

The physical and geochemical differences between Earth and Mars provide fundamental constraints on terrestrial planet formation, yet a self-consistent framework linking dynamical and chemical aspects remains elusive. Here we present an integrated modeling framework that couples high-resolution N-body simulations with impact-driven metal-silicate equilibration to track the dynamical accretion history and chemical differentiation for Earth and Mars. Using a narrow ring planetesimal accretion scenario, we show that Earth and Mars analogs naturally sample systematically different solid reservoirs within the protoplanetary disk. Earth analogs preferentially accrete reduced material around the planetesimal ring center, whereas Mars analogs acquire a larger fraction of oxidized material exterior to the ring. This leads to diverse bulk redox states, with composition further modified by impact-dependent pressure-temperature equilibration conditions during core formation. As a result, Earth analogs experience deeper equilibration and more efficient transfer of iron into the core, producing mantles with low iron oxide contents and larger core mass fractions. In contrast, Mars analogs equilibrate at shallower conditions, retain more iron in their mantles, and develop smaller cores. Our results demonstrate that the dynamical and geochemical differences between Earth and Mars emerge from the coupled effects of accretion pathways, the disk's radial redox structure, and impact-controlled differentiation rather than from any single process. Our unified framework physically explains the geochemical diversity of terrestrial planets and offers a potential pathway to interpret compositions of rocky planets in exoplanetary systems.

astro-ph.EP

Cognitive Alpha Mining via LLM-Driven Code-Based Evolution

Discovering effective predictive signals, or "alphas," from financial data with high dimensionality and extremely low signal-to-noise ratio remains a difficult open problem. Despite progress in deep learning, genetic programming, and, more recently, large language model (LLM)-based factor generation, existing approaches still explore only a narrow region of the vast alpha search space. Neural models tend to produce opaque and fragile patterns, while symbolic or formula-based methods often yield redundant or economically ungrounded expressions that generalize poorly. Although different in form, these paradigms share a key limitation: none can conduct broad, structured, and human-like exploration that balances logical consistency with creative leaps. To address this gap, we introduce the Cognitive Alpha Mining Framework (CogAlpha), which combines code-level alpha representation with LLM-driven reasoning and evolutionary search. Treating LLMs as adaptive cognitive agents, our framework iteratively refines, mutates, and recombines alpha candidates through multi-stage prompts and financial feedback. This synergistic design enables deeper thinking, richer structural diversity, and economically interpretable alpha discovery, while greatly expanding the effective search space. Experiments on 5 stock datasets from 3 stock markets demonstrate that CogAlpha consistently discovers alphas with superior predictive accuracy, robustness, and generalization over existing methods. Our results highlight the promise of aligning evolutionary optimization with LLM-based reasoning for automated and explainable alpha discovery.

cs.CL

A Self-Evolving Agentic Framework for Metasurface Inverse Design

Metasurface inverse design can realize complex optical functionality, but turning a target optical response into executable optimization code still requires substantial expertise in computational electromagnetics and solver-specific software engineering. We present a self-evolving agentic framework that lowers this barrier by coupling a coding agent, explicit human-readable skill files, and a deterministic physics-based evaluator. Rather than updating model weights, it revises the skill files from solver-grounded feedback, while the base model and differentiable solver, which provides the physics simulation and gradients, stay fixed. On a multi-type benchmark, skill evolution raises same-type task success from 38\% to 74\%, the fraction of physical criteria met from 0.51 to 0.87, and reduces average attempts from 4.10 to 2.30. On two new-type families, success holds near ceiling on one (0.92 to 0.90) and rises from 0.20 to 0.90 on the other. Skill evolution offers a practical path toward autonomous and accessible inverse-design workflows.

cs.AI

Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions

Self-evolving frameworks usually optimize task solutions while treating the surrounding harness as fixed. We introduce Harness-Aware Self-Evolving (HASE), an agentic reinforcement-learning framework in which a single model can generate task solutions or edit selected harness components in a multi-turn action space. HASE enables a single Qwen3-8B model to match the text-classification performance of a GPT-OSS-120B model that uses Claude Code as the harness proposer. In alpha factor mining, HASE outperforms the reported GPT-OSS-120B baseline. HASE also repairs imperfect evaluation components and converges to state-of-the-art performance in circle-packing algorithm discovery. These results show that HASE improves the harness and the solution through one unified agentic process.

cs.AI

Transferable Attack against Face Swapping in an Extended Space

Although deep Face Swapping (FS) models may benefit the entertainment industry, they pose severe threats to privacy and security. Existing protections, including deepfake detection and adversarial perturbation, are either passive responses or ineffective to unseen subject-agnostic FS models. In this paper, we propose a transferable attack against subject-agnostic FS models named Additive Identity attack based on a Relighting function (AIR). AIR leverages reillumination and additive perturbations to mislead the identity extraction modules in subject-agnostic FS models. By using these two types of perturbations simultaneously, the attack space is extended such that stronger but more visually natural adversarial examples can be identified. To further enhance the visual quality while preserving the effectiveness of the attack, an adaptive translation-invariant operation and an illumination control scheme are designed for AIR. Unlike other methods, AIR does not require a surrogate FS model to achieve high transferability. In addition, a mathematical proof is given for the extension of the attack space. Extensive experiments using 1000 image pairs across various state-of-the-art subject-agnostic FS models, including GAN and diffusion-based FS models, show that AIR surpasses all existing attacks in terms of both attack success rate and image quality.

cs.CV

AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments

Autonomous AI agents have driven the transition from conversation to task execution, shifting security failures from textual deception to system compromise. Although security evaluation is crucial for proactive risk prevention, prior work is constrained by fundamental bottlenecks, including fragmented risk coverage, static or low-fidelity execution environments, and single-dimensional and coarse-grained assessment metrics. To address these challenges, we propose AgentCanary, a comprehensive security evaluation framework for autonomous AI agents. AgentCanary provides a systematic solution along three contributions. First, comprehensive risk coverage: we introduce an orthogonal Entry $\times$ Impact risk taxonomy that decouples how adversarial influence enters the agent from what harm it ultimately causes, and instantiate it as a scenario-aligned task suite spanning realistic deployment workflows. Second, a high-fidelity real executable environment: rather than static Q&A or mocked tool responses, agents interact with real tools against dynamically provisioned task artifacts, with persistent state across multi-step interactions that naturally supports long-horizon attack evaluation. Third, trajectory-grounded multi-dimensional evaluation: evaluation consumes the full agent trajectory rather than the reply text or a single tool call, enabling decomposed scoring along three orthogonal dimensions, Outcome Safety, Security Awareness, and Task Utility. We evaluate a broad set of frontier models on AgentCanary against multiple established adversarial attack methods across three agent frameworks. The results reveal that current agents often fail to recognize the attacks they face, particularly under compromised skills, persistent state, and long-horizon execution attacks, and provide a systematic baseline for developing more reliable and secure agent systems.

cs.CR

DiffSight-Former: Modeling Structural Differences and Temporal Dynamics for Glaucoma Progression Prediction

Glaucoma is a leading cause of irreversible blindness worldwide, and early detection from fundus images is critical for effective disease management. While deep learning has achieved promising performance in fundus image analysis, most existing methods rely on single time-point images and fail to capture longitudinal structural and vascular changes associated with disease progression. Sequential fundus images acquired during clinical follow-up provide valuable temporal information; however, current sequential models often struggle to detect subtle early progression signals and commonly depend on fixed-length inputs or diagnostic cues from already glaucomatous images, limiting their clinical utility for early prediction. To address these limitations, we propose DiffSight-Former, a framework for glaucoma progression prediction from sequential fundus images. It incorporates a time-variant feature extraction module based on a fundus-specific foundation model to obtain robust anatomical representations. A multi-structure difference modeling module is introduced to quantify progression-related changes in the optic disc/cup region and retinal vasculature. These representations are integrated with temporal interval embeddings and processed by a time-aware Transformer to model disease progression and estimate the probability of future glaucoma onset. Experiments were conducted on two longitudinal datasets, SIGF (405 sequences) and GRAPE (263 sequences). On SIGF, DiffSight-Former achieved an AUC of 91.54% and a sensitivity of 92.16% for progression prediction. On GRAPE, it achieved an average accuracy of 87.48% across three clinical visual-field progression criteria. Compared with existing approaches, DiffSight-Former demonstrates strong performance and robustness across different temporal settings, highlighting its potential for longitudinal glaucoma monitoring and early risk prediction.

cs.CV

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression

The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when they process long contexts. Current KV Cache eviction has become an important research direction; however, existing methods based on fixed Soft Tokens (e.g., Judge Q) rely on a static parameter set as the query to evaluate the importance of KV pairs, so they cannot adapt dynamically to different input prompts, and they cannot precisely capture complex and changing task relevance. Also, evicted KV pairs are discarded permanently, so this causes irreversible information loss and context breaks. To address this problem, we propose Meta-Soft, a dynamic compression framework based on probe-driven context integration. Specifically, we build a meta-library with a learnable orthogonal basis matrix $\mathcal{L}$, and we use a selector network with Gumbel-Softmax to produce differentiable sparse combination weights, so we dynamically synthesize the most targeted $k$ Soft Tokens from the input prompt features. We append these Soft Tokens to the end of the input sequence to probe key information. We also introduce an attention-flow based integration mechanism, which redistributes the semantic information of removed tokens into retained tokens, and this keeps the dropped context information effectively. Experiments on multiple datasets show that our method outperforms existing state-of-the-art eviction methods and provides a new solution for KV Cache compression.

cs.AI

Is Fixing Schema Graphs Necessary? Full-Resolution Graph Structure Learning for Relational Deep Learning

Relational prediction tasks are fundamental in many real-world applications, where data are naturally stored in relational databases (RDBs). Relational Deep Learning (RDL) addresses this problem by modeling RDBs as graphs and applying graph neural networks (GNNs) for end-to-end learning. However, the full-resolution property is commonly adopted as a design principle in graph construction for RDBs to preserve relational semantics, which leads most existing methods to rely on fixed graph structures. In this paper, we propose FROG, a Full-Resolution and Optimizable Graph Structure Learning} framework for RDL that formulates relational structure learning as a learnable table role modeling problem, allowing tables to contribute as nodes and edges in message passing. We further design role-driven message passing mechanisms to capture relational semantics, enabling joint optimization of graph structure and GNN representations. To ensure semantic consistency, we introduce functional dependency constraints that regularize representations across table and entity levels. Extensive experiments demonstrate that our method outperforms existing approaches and reveal how table roles impact downstream tasks, offering new insights into graph construction for RDL

cs.LG

Singular band Induced by Long-Range Interaction Enables Unsplit Spreading of Localized Excitations

In conventional lattice models, the dispersion relation $ω(k)$ is assumed to be a smooth function which is periodic over the first Brillouin Zone. However, in subwavelength atom arrays the dispersion of the light-mediated long-range interaction is singular at the light cone. This observation prompts us to ask what effect arises from such band singularity. Here we demonstrate that, due to the topology of smooth functions defined over the periodic Brillouin zone, smoothness implies the splitting of an initially localized excitation into counter-propagating wave packets. Consequently, unsplit spreading can occur only when $ω(k)$ develops singular features, precisely what long-range interactions enable. We identify unsplit spreading in 1D toy tight-bounding models and the realistic models of 1D and 2D subwavelength atomic arrays. Our work establishes unsplit spreading as an experimentally accessible, smoking-gun signature of singular band structure.

quant-ph

The Ekeland--Nirenberg Variational Problem:A Sharp Positivity Threshold and Extensions

We study the Ekeland--Nirenberg variational problem in the two-dimensional diagonal family \[ J_{a,c,d}(u)=\int_{\Rp^2}\bigl(u_{xy}^2+a u_x^2+c u_y^2+d u^2\bigr)\dd x\dd y, \qquad a,c,d>0, \] under the constraint $u(0,0)=1$. If $u_{a,c,d}$ is the unique minimizer and $K_{a,c,d}$ is its cosine kernel, we prove the sharp classification \[ K_{a,c,d}>0 \hbox{ on } \Rp^2\quad\Longleftrightarrow\quad u_{a,c,d}>0 \hbox{ on } \Rp^2\quad\Longleftrightarrow\quad d\le ac . \] Thus every supercritical triple $d>ac$ produces sign change. We also prove local sign-change stability under small two-dimensional non-diagonal perturbations and a sharp product-type $n$-dimensional diagonal threshold. The domain and evolution results are stated in precise auxiliary settings: a free-boundary capacity formulation for domains and a selected decaying branch of the second-order evolution equation.

math.AP