Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of absent annotation guidelines and label scarcity in early-stage data labeling by proposing an agent-based pipeline that treats textual definitions as trainable objects. The method introduces the concept of "textual gradients," utilizing execution-based loss as a scoring signal to drive an LLM editor in iteratively refining structured definitions. Revisions are accepted only when the loss decreases, effectively decoupling definition learning from formatting, retrieval, and human review processes. Experimental results demonstrate that the proposed approach outperforms baselines such as OPRO under matching evaluations. Furthermore, when integrated with routing mechanisms and human review, it significantly enhances both the quality and scalability of multi-task annotation.
📝 Abstract
Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definition. Goldsmith treats this definition as a trainable textual object. Candidate definitions are run on the same gold examples and scored with an executable structured loss, while the output schema, formatting, retrieval, repair, judging, and human review remain in an external harness. A large language model (LLM) editor converts the highest-loss failures into textual-gradient revisions, which are accepted only when the measured loss decreases. In prompt-optimization comparisons, Goldsmith improves over direct rewriting, OPRO, APE, and PromptBreeder under matched evaluation protocols. The resulting definition also improves downstream annotation when combined with retrieval, score-based routing, and human review across typed span, pair-level relation, and fixed-trigger event-argument tasks. These results show that scarce expert supervision can support both task-definition learning and scalable annotation.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Annotation Pipeline
Textual Gradient Optimization
Gold-Loss-Guided Definition
Prompt Optimization
Structured Loss
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yihan Li
Sun Yat-sen University, Guangzhou, China
Hanyi Zhang
Hanyi Zhang
University of Heidelberg
Deep LearningBiomedical Imaging
X
Xiaoxi Jiang
Sun Yat-sen University, Guangzhou, China
M
Man Guo
Sun Yat-sen University, Guangzhou, China