You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost and reliance on heuristic proxies in demonstration set selection for in-context learning (ICL) with large language models. We propose Local Demonstration Editing (LDE), which reformulates demonstration selection as a constrained local search problem. Introducing a novel "edit-only-once" paradigm, LDE employs reinforcement learning to train a Jev-LDE model that executes retain, delete, or replace actions guided by verifiable rewards. A single structured edit thereby replaces conventional iterative scoring and subset search, substantially reducing inference overhead. Experimental results demonstrate that LDE consistently improves ICL performance across multiple benchmarks while exhibiting plug-and-play transferability across different models and tasks.
📝 Abstract
In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likelihood proxies to implicitly assess ICL quality. Making repeated queries to the target LLM with these strategies can incur substantial costs. This work simplifies selection by framing it as a constrained local search problem and presents local demonstration editing (LDE). Starting with an initially retrieved set of demonstrations, LDE employs a single structured edit to explore its surrounding neighborhood while balancing performance gains with search costs. Technically, LDE is reduced to a policy search problem, for which we train a small LLM, referred to as Jev-LDE. This model as the System-1 modifies the retrieved demonstration set by performing actions such as \texttt{Keep}, \texttt{Delete}, or \texttt{Replace} elements, all within a framework of reinforcement learning with verifiable rewards. At test time, Jev-LDE executes a single edit of the retrieved demonstration set, followed by one inference from the target LLM, avoiding the need for iterative context scoring or subset searches. Across standard classification benchmarks, various target LLMs with Jev-LDE as the plug-and-play module consistently improve ICL performance, and Jev-LDE shows transferability to held-out benchmarks and models without retraining. These findings indicate that the LDE approach offers an efficient and adaptable method for harnessing the ICL capabilities of target LLMs.
Problem

Research questions and friction points this paper is trying to address.

In-context learning
Large language models
Demonstration selection
Search cost
Local search
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Local Demonstration Editing
Reinforcement Learning
Policy Search
Large Language Models
🔎 Similar Papers
No similar papers found.
J
Jiarong Wen
Department of Automation, Tsinghua University
Qi Wang
Qi Wang
Tsinghua University
Operation researchReinforcement learning
Y
Yun Qu
Department of Automation, Tsinghua University
Y
Yixiu Mao
Department of Automation, Tsinghua University
Heming Zou
Heming Zou
Tsinghua University
Machine Learning
H
Haoang Chi
Department of Automation, Tsinghua University
L
Lizhou Cai
Department of Automation, Tsinghua University
Y
Yiqin Lv
Department of Automation, Tsinghua University
Kaiyu Zhang
Kaiyu Zhang
ML Researcher, University of Washington
Vascular ImagingMagnetic Resonance Imaging
Yuhang Jiang
Yuhang Jiang
Tsinghua University
Reinforcement LearningMachine Learning
X
Xiangyang Ji
Department of Automation, Tsinghua University