Verb-ICL: Rethinking In-Context Learning for Structured Prediction

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of sentence-level in-context learning (ICL) for structured prediction, which struggles to capture fine-grained patterns and acquire human annotation conventions. To this end, we propose Verb-ICL, a framework that employs a token-level coverage strategy to select representative examples for capturing local semantic patterns. Furthermore, it generates error feedback encoding annotation guidelines, which is then integrated into the demonstration context. Experiments across six datasets demonstrate that Verb-ICL significantly outperforms existing baselines. It achieves optimal performance particularly in low-resource scenarios and yields consistent improvements as the budget increases. Notably, the generated error feedback exhibits strong cross-task generalizability as instructional guidance.
📝 Abstract
Structured prediction tasks pose unique challenges for in-context learning (ICL): their compositional outputs require modeling fine-grained, token-level patterns that sentence-level approaches fail to capture, and their task-specific annotation conventions are human-defined artifacts that cannot be acquired through pretraining alone. We propose Verb-ICL, a selective annotation framework for ICL-based structured prediction that addresses both challenges. Verb-ICL first selects representative examples using a token-level coverage strategy that captures local semantic patterns critical for structured prediction, then generates actionable error feedback that codifies task-specific annotation guidelines and incorporates this feedback into ICL demonstrations. We evaluate Verb-ICL on six structured prediction datasets spanning information extraction and semantic parsing. Experiments with recent LLMs show that Verb-ICL consistently outperforms strong selective annotation baselines under low-resource settings and continues to provide gains as the annotation budget increases. Extended analyses demonstrate that the generated feedback is predominantly useful across a four-category quality taxonomy, generalizes as task-level guidance beyond instance-specific corrections, and improves performance regardless of the underlying selection strategy.
Problem

Research questions and friction points this paper is trying to address.

Structured Prediction
In-Context Learning
Token-level Patterns
Annotation Conventions
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Structured Prediction
Selective Annotation
Token-level Coverage
Error Feedback
🔎 Similar Papers
No similar papers found.