arXiv · 2610.09489
Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness
Abstract
Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definition. Goldsmith treats this definition as a trainable textual object. Candidate definitions are run on the same gold examples and scored with an executable structured loss, while the output schema, formatting, retrieval, repair, judging, and human review remain in an external harness. A large language model (LLM) editor converts the highest-loss failures into textual-gradient revisions, which are accepted only when the measured loss decreases. In prompt-optimization comparisons, Goldsmith improves over direct rewriting, OPRO, APE, and PromptBreeder under matched evaluation protocols. The resulting definition also improves downstream annotation when combined with retrieval, score-based routing, and human review across typed span, pair-level relation, and fixed-trigger event-argument tasks. These results show that scarce expert supervision can support both task-definition learning and scalable annotation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yihan Li, Hanyi Zhang, Xiaoxi Jiang, Man Guo. 2026-10-07. Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness. https://arxiv.org/abs/2610.09489
Cite the original work for its findings. Save a collection to share your selection of sources.