arXiv ScienceSearch

arXiv subjects

Yi Fan

Publications and source records attributed to Yi Fan.

At least 19 recordsLinked to original sources

Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence

Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly and does not scale. The ATT&CK taxonomy comprises several hundred attack techniques, and a single CTI passage may describe multiple techniques, making accurate and complete extraction challenging. Existing automated approaches fall short in different ways: multi-label classifiers struggle with severe class imbalance and the large label space, while LLM-based methods--retrieval pipelines and fine-tuned generators--optimize token-level objectives that treat technique annotation as sequence generation rather than set prediction, lacking direct supervision on whether the predicted technique set is correct and complete. We propose TTP-R1, a two-stage framework that combines retrieval-augmented supervised fine-tuning (SFT) with reinforcement learning using verifiable rewards (RLVR). A hybrid retriever first narrows the large label space to a candidate set, and a fine-tuned LLM learns to select the correct techniques. We then apply Group Relative Policy Optimization with a decomposed reward that directly supervises the precision, recall, and output format of the predicted technique set. Across four CTI benchmarks, TTP-R1 achieves the best average F1, improving sub-technique-level F1 by 7.4 percentage points over Claude Sonnet 4.5 with retrieval augmentation, while running 28x faster when served as an 8B-parameter model on a single GPU.

cs.CR

Embedded quantum computing for many-body surface reaction

Predictive simulations of catalytic interfaces require correlated electronic-structure treatments that describe localized chemical transformations while retaining the influence of the extended metallic environment. We introduce QC-DFET, a quantum-computing density-functional embedding framework that maps surface-reaction active spaces to compact, environment-aware qubit Hamiltonians. A reaction-consistent active-space protocol preserves orbital continuity along reaction coordinates, while quantum-selected configuration interaction based on measurements from the Zuchongzhi superconducting quantum processor and strongly contracted perturbation theory capture static and dynamic correlation. On Cu(111), QC-DFET treats active spaces up to 28 qubits and is validated through a hierarchy of experimentally constrained surface-chemistry challenges. H2 dissociation/desorption tests balanced bond breaking and recombination barriers, CO adsorption tests site selectivity and metal-adsorbate bonding, and formate hydrogenation tests competing hydrogenation branches with different kinetic and thermodynamic signatures. Across these cases, QC-DFET reproduces bidirectional H2 barriers, recovers the observed top-site preference and adsorption strength of CO, and reconciles the experimentally benchmarked H2COO* reverse barrier with the lower forward barrier to HCOOH*. These results establish embedded quantum computing as a practical route to correlated surface-reaction energetics.

quant-ph

The Price of Quietness: How a Pandemic Affects City Dwellers' Response to Road Traffic Noise

Using the outbreak of COVID-19 in Singapore as a quasi-natural experiment, we investigate tenants' changing responses to road traffic noise in the rental housing market, using 46,980 transaction records between 2006 and 2022. Our difference-in-differences estimates show that road traffic noise decreases housing rents by 3.8% immediately after the pandemic outbreak and further declines by 12.7% in the subsequent year-equivalent to 186.7 US dollars per month. The results are robust to parallel trend analysis, permutation placebo tests, and tests using alternative distance thresholds or distance to the nearest main road. Then, we adopt a machine learning text analysis of 10,425 rental housing advertisements, showing that tenants' preference for quietness increases by approximately 10% from 2019 into 2020. The new work-from-home business model and rising traffic from delivery services can explain for this pattern. To the best of our knowledge, this is the first paper using a large volume of transaction records to quantify city dwellers' willingness to pay for quietness in the COVID-19 context. Our results have policy implications for other nations and post-pandemic era on the interaction among urban planning, transport networks, and human settlements, and shed light on the pathway to achieve sustainable development goals.

econ.GN

Noise Pollution and Household Sustainability: An Economic Approach

Examining the economic impact of noise pollution from a lens of household is a burgeoning field in the study of environmental sustainability. Economics studies cover the source, measure, consequence of noise pollution, as well as the econometric methods used to identify the causal impact of noise pollution on socioeconomic welfare. There are broadly four major noise origins along with the industrial growth and urban development, which are airport, railway, urban traffic, and neighborhood. Four general kinds of measures or data sources are used in economics studies to capture the noise variations, namely, proximity to noise origins, real-time noise monitor records, household surveys, and administrative records on noise complaints. The socioeconomic consequences of noise pollution span from physical or mental health to happiness, violence and suicide, housing market capitalization, and inequality. In economics studies, generally three types of econometric methods are used to identify causal impact of noise pollution on the household's welfare, which are instrumental variable estimation, difference-in-difference estimation, randomized and quasi-natural experiments. The causal impact of noise pollution on household's socioeconomic welfare derived from economics studies can help guide policy efforts in allocating resources for noise elimination and conduct cost-benefit analysis. The economics research contributes to the general noise research from both conceptual and methodological perspectives: It expands the scope of research from sound-poof technology or site layout planning to human welfare, and endeavors to isolate the causal impact of noise pollution from other confounding factors. Future studies are warranted along the lines of environmental injustice of noise pollution and socioeconomic consequences in less developed countries when the data become more available.

econ.GN

Social Integration and Housing Behaviours of Immigrants: Evidence from Singapore's Public Housing Market

This study investigates the impact of social integration on immigrants' housing behaviours from a temporal perspective, using Singapore's differential public housing policies on immigrants as a quasi-natural experiment. With the support of a local town council, we conducted a survey on social integration among 1,128 immigrant and local households living in public housing estates. In the public open rental housing market - primarily accommodating yet-to-integrate immigrants - we find immigrant renters live up to 3.04% farther from their workplace and pay lower rents up to 0.67% per additional year of residency. Such impacts are more substantial among minority ethnic groups. The results remain robust when using alternative subjective or objective measures of social integration. However, in the public resale housing market - primarily accommodating native and well-integrated naturalised citizens - we find that naturalised citizens face no price premiums relative to native homebuyers, implying no further effect of integration on housing prices after well-integration. This study extends the literature of spatial assimilation focusing on ethnic residential segregations and is generalizable to cities with few ethnic enclaves.

econ.GN

Ageing in which place? Spatial analytical framework for evaluating ageing-in-place practices

Over the past decade, governments around the world have made significant investments in creating elderly-friendly urban environments within local neighborhoods. However, the lack of a standardized evaluation framework for Ageing-in-Place (AIP) practices makes it challenging to generalize these experiences. First, we compare the AIP models of the U.S.- San Francisco, Japan-Tokyo, and Singapore using a cost-benefit analysis, demonstrating the comparative advantage of the Singapore model in terms of low cost and high accessibility for the independent ageing population. Second, we propose a spatial analytics framework to visualize and quantify the degree of alignment between a basket of ageing facilities and the active ageing population, enabling a data-driven, timely evaluation of the effectiveness of Singapore's AIP policies. Singapore's AIP model, either in its entirety or as a hybrid with other models, can be generalized to other global cities, providing valuable insights for optimal elderly-friendly urban planning.

econ.GN

Testing Black Holes with Interstellar Missions: II. Flyby Probes

Recently, we demonstrated that while an interstellar mission to the nearest black hole remains highly speculative and extraordinarily challenging, it is not entirely implausible within the coming decades. Given that such a mission would likely take about a hundred years and require substantial financial and human investment, it is essential to assess whether it could investigate black holes and test General Relativity to a degree that cannot be achieved by Solar System observatories for the foreseeable future. In Paper I, we assumed the capability to decelerate the spacecraft and presented a preliminary study of how orbiting probes could test the nature of the compact object. In this second paper, we study how the black hole can be tested without decelerating the spacecraft, using flyby probes.

gr-qc

A High-Performance Pauli-Algebra Framework for Large-Scale Quantum Simulations

Efficient manipulation of Pauli-algebraic objects is a key bottleneck in the classical emulation and benchmarking of quantum algorithms for chemistry and many-body physics. This bottleneck appears in Hamiltonian construction, variational ansatz preparation, expectation-value and gradient evaluation, and real-time propagation, all of which require repeated Pauli-algebra operations. Here, we present a high-performance Pauli-algebra framework tailored to quantum many-body and quantum-chemical simulations. The framework combines compact binary symplectic encoding, canonical coefficient reduction, and grouped sparse operator representations that exploit shared bit-flip patterns among Pauli strings. The resulting Julia/C\texttt{++} implementation accelerates Pauli multiplication, Hamiltonian construction, and operator--state multiplication in sparse and symmetry-adapted many-electron spaces. Benchmarks demonstrate efficient Hamiltonian construction, large-active-space VQE and ADAPT-VQE calculations, and real-time variational dynamics on modern multicore CPU and GPU architectures. These results show that structure-aware Pauli-algebra engines provide a scalable classical backend for developing and benchmarking quantum algorithms in quantum chemistry and many-body simulation.

quant-ph

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines. Existing benchmarks are limited to small or simplified examples far below real-world scale and complexity. We introduce FrontierOR, among the first benchmarks to systematically evaluate LLM-based efficient algorithm design for realistic large-scale optimization problems. FrontierOR includes 180 tasks derived from methodologically diverse papers published in top-tier operations research venues, each with standardized instances and a hidden, expert-verified evaluation suite. We evaluate seven LLMs spanning frontier, cost-effective, and open-source models both in one-shot and test-time evolution settings. The results reveal that frontier models still struggle to move from executable formulations to efficient optimization algorithms: the strongest one-shot model outperforms Gurobi in only 31% of cases in both solution quality and computational efficiency, and even strong coding agents with test-time evolution achieve only 50% on selected hard tasks. FrontierOR establishes a practical evaluation platform for LLM-based optimization algorithm design, which enables future LLMs and agents to be systematically tested on whether they can move beyond correct formulation toward a feasible, high-quality, and efficient algorithm. Code and data are publicly released at https://github.com/Minw913/FrontierOR.

cs.AI

Transformer refined quantum sampling for strongly correlated electronic structure

Although quantum computing offers a promising solution for strongly correlated system simulation, existing algorithms face significant bottlenecks on current noisy intermediate-scale quantum (NISQ) devices. Here, we introduce QiankunNet-QSCI, a hybrid quantum-classical framework that addresses this challenge by combining efficient quantum-sampling with a transformer neural network. An efficient unitary selected configuration Interaction (USCI) ansatz especially designed for quantum sampling is proposed to identify the most chemically significant electronic configurations on the Zuchongzhi 3.1 quantum processor. Subsequently, the transformer model QiankunNet learns from these sparse yet critical quantum data to infer and reconstruct the complete electronic wavefunction with high fidelity. Simulation of the challenging 40-qubit [2Fe-2S] ferredoxin active center achieves chemical accuracy. Simulation of the nitrogenase P-cluster in a 114-electron 73-orbital active space also reaches 12 milli-Hartree-level agreement with the best density matrix renormalization group (DMRG) result. QiankunNet-QSCI thus offers a practical route to accurate quantum-assisted electronic structure calculations on current devices.

quant-ph

Testing Black Holes with Interstellar Missions: I. Orbiting Probes

Recently, we showed that the possibility of an interstellar mission to the closest black hole, while highly speculative and extremely challenging, is not completely unrealistic within the next few decades. Since such a mission might last around a century and require significant financial and human resources, it is crucial to assess whether it can truly study black holes and test General Relativity at levels unattainable by observational facilities in the Solar System for many years. In this manuscript, we assume the capability to decelerate the spacecraft and present a preliminary study of how probes orbiting a black hole could test the nature of the compact object.

gr-qc

SAT: Sequential Agent Tuning for Coordinator Free Plug and Play Multi-LLM Training with Monotonic Improvement Guarantees

Large language models (LLMs) with a large number of parameters achieve strong performance but are often prohibitively expensive to deploy. Recent work explores using teams of smaller, more efficient LLMs that collectively match or even outperform a single large model. However, jointly updating multiple agents introduces compounding distribution shifts, making coordination and stability during training difficult. We address this by introducing Sequential Agent Tuning (SAT), a coordinator-free training paradigm. SAT represents the team as a factorized policy and employs block-coordinate updates over agents, enabling scalable, decentralized training without a central controller. Specifically, we develop a sequence-aware, on-policy advantage estimator that conditions on the evolving team policy, coupled with per-agent KL trust regions that isolate occupancy drift. Theoretically, this framework provides two critical guarantees. First, it ensures monotonic improvement, stabilizing the training process. Second, it establishes provable plug-and-play invariance: any agent can be upgraded to a stronger model without retraining the rest of the team, with a formal guarantee that the performance bound improves. Empirically, a team of three 4B agents (12B total) trained with SAT surpasses the much larger Qwen3-32B on AIME24/25 benchmarks by 3.9\% on average. We validate our plug-and-play theory by swapping in two 8B agents, which boosts the composite score by 10.4\%. We provide code and appendix of proof at https://github.com/Yydc/SAT-AAMAS

cs.LG

FraudFox: Adaptable Fraud Detection in the Real World

The proposed method (FraudFox) provides solutions to adversarial attacks in a resource constrained environment. We focus on questions like the following: How suspicious is `Smith', trying to buy \$500 shoes, on Monday 3am? How to merge the risk scores, from a handful of risk-assessment modules (`oracles') in an adversarial environment? More importantly, given historical data (orders, prices, and what-happened afterwards), and business goals/restrictions, which transactions, like the `Smith' transaction above, which ones should we `pass', versus send to human investigators? The business restrictions could be: `at most $x$ investigations are feasible', or `at most \$$y$ lost due to fraud'. These are the two research problems we focus on, in this work. One approach to address the first problem (`oracle-weighting'), is by using Extended Kalman Filters with dynamic importance weights, to automatically and continuously update our weights for each 'oracle'. For the second problem, we show how to derive an optimal decision surface, and how to compute the Pareto optimal set, to allow what-if questions. An important consideration is adaptation: Fraudsters will change their behavior, according to our past decisions; thus, we need to adapt accordingly. The resulting system, \method, is scalable, adaptable to changing fraudster behavior, effective, and already in \textbf{production} at Amazon. FraudFox augments a fraud prevention sub-system and has led to significant performance gains.

cs.CR

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration

Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong barrier for PbRL. In this work, we address the problem of low query efficiency in offline PbRL, pinpointing two primary reasons: inefficient exploration and overoptimization of learned reward functions. In response to these challenges, we propose a novel algorithm, \textbf{O}ffline \textbf{P}b\textbf{R}L via \textbf{I}n-\textbf{D}ataset \textbf{E}xploration (OPRIDE), designed to enhance the query efficiency of offline PbRL. OPRIDE consists of two key features: a principled exploration strategy that maximizes the informativeness of the queries and a discount scheduling mechanism aimed at mitigating overoptimization of the learned reward functions. Through empirical evaluations, we demonstrate that OPRIDE significantly outperforms prior methods, achieving strong performance with notably fewer queries. Moreover, we provide theoretical guarantees of the algorithm's efficiency. Experimental results across various locomotion, manipulation, and navigation tasks underscore the efficacy and versatility of our approach.

cs.LG

Machine learning modularity

Based on a transformer based sequence-to-sequence architecture combined with a dynamic batching algorithm, this work introduces a machine learning framework for automatically simplifying complex expressions involving multiple elliptic Gamma functions, including the $q$-$\theta$ function and the elliptic Gamma function. The model learns to apply algebraic identities, particularly the SL$(2,\mathbb{Z})$ and SL$(3,\mathbb{Z})$ modular transformations, to reduce heavily scrambled expressions to their canonical forms. Experimental results show that the model achieves over 99\% accuracy on in-distribution tests and maintains robust performance (exceeding 90\% accuracy) under significant extrapolation, such as with deeper scrambling depths. This demonstrates that the model has internalized the underlying algebraic rules of modular transformations rather than merely memorizing training patterns. Our work presents the first successful application of machine learning to perform symbolic simplification using modular identities, offering a new automated tool for computations with special functions in quantum field theory and the string theory.

hep-th

False Positives Raised by Quantum Readout Error Mitigation

Quantum readout error mitigation is essential for noisy intermediate-scale quantum devices to achieve reliable data. The conventional approaches, conflating initialization errors with measurement errors, not only suppress the influence of measurement errors, but also strengthen that of initialization errors, which is a systematic bias grows exponentially with the qubit number. Here, we have proved that this effect causes severe fidelity overestimation for all stabilizer states and might lead to false positives in large-scale entangled state characterization. Similarly, the results from algorithms like the variational quantum eigensolver and time evolution also deviate negatively, and cover up other errors in the quantum circuit. These findings highlight the critical need for rigorous benchmarking and careful management of initialization errors. Consequently, we establish an upper bound for the tolerable initialization error rate to ensure effective error mitigation at a given system scale.

quant-ph

Scalable parallel simulation of quantum circuits on CPU and GPU systems

Quantum computing enables parallelism through superposition and entanglement and offers advantages over classical computing architectures. However, due to the limitations of current quantum hardware in the noisy intermediate-scale quantum (NISQ) era, classical simulation remains a critical tool for developing quantum algorithms. In this research, we present a comprehensive parallelization solution for the Q$^2$Chemistry software package, delivering significant performance improvements for the full-amplitude simulator on both CPU and GPU platforms. By incorporating batch-buffered overlap processing, dependency-aware gate contraction and staggered multi-gate parallelism, our optimizations significantly enhance the simulation speed compared to unoptimized baselines, demonstrating the effectiveness of hybrid-level parallelism in HPC systems. Benchmark results show that Q$^2$Chemistry consistently outperforms current state-of-the-art open-source simulators across various circuit types. These benchmarks highlight the capability of Q$^2$Chemistry to effectively handle large-scale quantum simulations with high efficiency and high portability.

quant-ph

The Three-Dimensional Velocity Field of Kinesin-Driven Microtubules in Torroidal Channels

We study two regimes of flow in multiple three-dimensional toroidal channels by tracking the fluorescent spherical particles in the kinesin-driven microtubule systems: ``chaotic'' flow and ``coherent'' flow. In the smallest aspect ratio torus, where the channel height $h$ is a quarter of the width $w$, the active system shows zero mean velocity, small-scale isotropy in fluctuation and no persistent flow structure. In other tori with higher aspect ratios $h/w$ close to 1, we find faster coherent flows along the azimuthal direction and increasing fluctuation strengths with growing confinement geometries. Regardless of flow regimes, the flow profiles at $r-z$ cross-section and $r-\theta$ plane are symmetric. The ``coherent'' profiles show two criteria: ``Poiseuille-like'' profiles, which have the peak velocities near the centers of channels; a ``peak-separated'' profile, which has four peak velocities near a certain distance to four confining surfaces. These flow profiles, after scaled by the local isotropic fluctuation strength, reveal universal three-dimensional flow structures among the ``Poiseuille-like'' criterion and the same level of scaled peak velocity at the ``peak-separated'' one. These results illustrate scalable flow structures in this kinesin-driven microtubule active system.

physics.flu-dyn