arXiv · 2609.33153
What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review
Abstract
Large language model (LLM) agents increasingly draw on external tools and reusable skills selected at run time from libraries that hold thousands of entries. Reports that a retriever, router, or skill library "improves" an agent may refer to retrieval recall, the success change from enabling a library, a paired contrast restricted to triggered tasks, or a gain under an approximate budget constraint. This critical review asks what each design compares and under which assumptions. Building on estimand-based approaches to agent evaluation, we describe tool and skill designs along six axes: treatment contrast, target population, outcome, budget constraint, summary measure, and identification assumptions. Thirteen core empirical studies anchor the evidence synthesis, supplemented by related methodological work and design-level reading of the wider literature. Our contribution is to make explicit distinctions that some source authors already acknowledge through a decomposition of trigger-conditioned pairing, analytic counterexamples, and comparisons across studies. Pairing on the task does not by itself identify an invocation effect; paired gain and regression counts measure protocol-specific discordance rather than the share of tasks whose expected outcomes worsen; and total effects of module deployment answer a different question from budget-constrained efficiency. We compare curated skill provision with retriever replacement, triggered subsets with all-task outcomes, and observed cost reductions with budget-constrained comparisons. A reporting checklist and worked examples connect these distinctions to information that studies can report. The review runs no new experiments; empirical results come from the cited studies, and numerical toy examples are analytical illustrations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shuyang Zhang. 2026-09-27. What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review. https://arxiv.org/abs/2609.33153
Cite the original work for its findings. Save a collection to share your selection of sources.