arXiv Science⌕ Search

arXiv · 2609.36807

XRepoSkill: Learning Transferable Skills for Software Engineering Agents

Abstract

Software engineering agents increasingly use reusable skills distilled from prior experience to resolve repository-level issues, yet such skills often fail to transfer across repositories. A central challenge is that a behavior appearing in a successful trajectory is not necessarily responsible for the successful outcome: it may be genuinely useful, merely incidental, or simply a recurring habit of the model. We introduce XRepoSkill, a trajectory-based approach for learning transferable skills. We represent a skill as a collection of rules, each specifying what action to take and when to take it during issue resolution. XRepoSkill first contrasts successful and failed trajectories of the same agent on the same issue and derives candidate rules from where their execution paths diverge. Each rule is paired with an executable predicate that enables its prescribed behavior to be evaluated systematically on other trajectories. A rule is verified based on its association with successful issue resolution and retained only when its prescribed behavior recurs across multiple repositories; repository-specific variants of the same behavior are then consolidated into transferable rules. For a new issue, XRepoSkill selects relevant rules to guide the agent. We learn skills from publicly released trajectories on the official SWE-bench Verified leaderboard and evaluate them on SWE-bench Pro and DeepSWE using three backbone LLMs from different vendors; none of the evaluation repositories appears in the skill-learning trajectory pool. Against three recent skill learning methods, XRepoSkill achieves the highest issue resolution rate in all six benchmark--LLM combinations. In particular, on the challenging long-horizon DeepSWE benchmark, XRepoSkill improves issue resolution by 10.3 percentage points over the same agent without learned skills and by 5.0 points over the strongest skill-learning baseline.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yaoqi Guo, Haoyang Zhou, Jiayi Zhang, Yiran Zhang, Yang Liu, Qiuyuan Chen, Qiang Lin, Hande Dong, Jie M. Zhang, Zhenpeng Chen. 2026-09-29. XRepoSkill: Learning Transferable Skills for Software Engineering Agents. https://arxiv.org/abs/2609.36807

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Specification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks

Code generated by large language models passes security checks at a rate that has barely moved in four years. In regulated backends, the defect classes that matter most are money arithmetic, time handling, retry safety, and access control. Teams answer with instruction files, yet the largest controlled study of instruction files we are aware of found no general benefit. This paper tests a narrower idea: generated code improves when the prompt carries a specification, a fixed preamble stating what must be true of the result. We pre-registered hypotheses, refuters, analysis code, and a one-shot generation rule, then ran 50 realistic backend tasks from finance, healthcare, and insurance practice through five frontier models from five vendor lineages, each task twice: bare, and preceded by a 267-word filled specification frame. Nine deterministic AST-based checkers scored the outputs. The Bandit security scanner, which knows nothing of the frame, scored them independently. The frame reduced defects in all five models (mean reduction 0.16 to 0.70 findings per task, every Holm-adjusted sign test significant, every bootstrap confidence interval excluding zero). Where the arms differed, the frame arm won 95 of 100 times. It never made any model worse in any domain. Bandit found 53 medium-or-high issues in the bare arm and 11 in the frame arm, in the same direction for every model. The effect was largest where a model's unprompted defaults were weakest: the frame supplies the discipline a model lacks. All 500 outputs, prompts, checkers, scoring code, and the pre-registration are published with a DOI, so any team can re-derive the result without trusting the author.

cs.SE↗

Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests

Reflective optimizers such as GEPA improve language model prompts from execution traces and evaluator feedback; full-program extensions can also rewrite tools and control flow. In practice, a user hands the same endpoint heterogeneous requests whose effective solutions require different tools, reasoning modes, and control flow. Optimizing one shared program leaves this division of work implicit in source-code search, while optimizing a separate program per request family fixes it beforehand. We introduce Adaptive-GEPA, which learns both how to divide requests and how to solve them. It evolves a router and a library of specialist programs under one search budget. The router's instructions, each specialist's description, and its program code are plain, human-readable text, edited from feedback. To combine branches, it aligns specialists by the requests they handle and inherits descriptions together with programs. On a fixed mixture of four task families, the reported Qwen3-8B run evolves four experts without supplying family labels to the router or reflection model; its routing matches the task partition on all 651 test requests. Its family-mean test score (x100) rises from 52.6 to 70.6, compared with 62.5 for GEPA's full-program adapter and 54.0 for GRPO at a nominal budget of 18,000 scored calls. These counts do not equate total compute. Figure 1 summarizes the learning curves, final test scores, and routing agreement.

cs.SE↗

Toward Quantum Software Automation: A Quantum-Aware Harness for LLM-Guided Evolution

Quantum software is critical for improving the efficiency and reliability of scarce quantum hardware. However, its design still relies heavily on ad-hoc, handcrafted heuristics that are often suboptimal and quickly become obsolete as quantum hardware evolves. LLM-guided evolutionary search offers a promising way to automatically explore complex software designs, but existing search frameworks lack the quantum-specific support needed for efficient evolution: verification is expensive, feedback is sparse, and heterogeneous quantum programs require different optimization objectives. In this paper, we present QSA, a quantum-aware harness for LLM-guided evolutionary search toward automating quantum software design. QSA equips the search with three forms of quantum-specific guidance: an evolution-hardness-guided coreset and approximate scoring to reduce verification cost, static and snapshot analyses to provide fine-grained execution context, and task-specific rewards for compiler passes and runtime policies. We evaluate QSA on the IBM Quantum platform across three benchmark suites. For multiprogramming, QSA improves QPU utilization by 4.2%-9.5% and Hellinger fidelity by 15.2%-19.5% over the state of the art. For error mitigation, QSA reduces mitigation time by at least 96.8% while achieving comparable or better fidelity. These gains require only $6.9 in LLM API cost over 11.3 hours.

cs.SE↗