Streamlined Reflective Evolution for Task-Adaptive Self-Refinement Pipelines

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of fixed architectures and instruction redundancy caused by repetitive revisions in large language model self-correction. To overcome these challenges, we propose WDA, a framework that jointly evolves stage-wise instructions and structures to generate task-adaptive self-correction pipelines. Methodologically, WDA introduces a SPLIT mechanism to disperse redundant instructions, integrating three-example reflection, local filtering, and Pareto admission strategies. It further combines prompt optimization, workflow design agents, and evolutionary search to achieve efficient calibration. Experimental results demonstrate that this framework improves average scores by 8.63 and 5.62 percentage points on Qwen3.5-9B and GPT-4.1-mini, respectively, significantly enhancing the self-correction capabilities of large language models.
📝 Abstract
Reflective prompt optimization improves large language model (LLM) systems without updating model weights, but fixed architectures constrain how self-refinement is organized. We introduce Workflow-Designing Agents (WDA), a framework for streamlined reflective evolution of task-adaptive self-refinement pipelines. Starting from a minimal prompt, WDA jointly evolves stage instructions and their sequential structure. During evolution, we find that repeated revisions can accumulate redundant instructions in a single prompt. In WDA, we propose to address this problem with SPLIT, which redistributes these instructions across specialized stages. Three-example reflection and local screening guide selective search, while calibration scores guide Pareto admission and rollback of unhelpful trailing updates. The resulting pipelines are task-adaptive: their instructions and depth are learned from task data, then fixed for all test inputs within that task. We evaluate WDA on five benchmarks spanning knowledge, mathematical reasoning, multi-hop question answering, and instruction following. On Qwen3.5-9B, WDA achieves an average score of 51.24%, improving over the initial solver by 8.63 percentage points and the variant without SPLIT by 2.60 points. On GPT-4.1-mini, it achieves 49.00%, with corresponding gains of 5.62 and 3.69 points. These results support task-adaptive self-refinement as a complementary direction to broader agentic workflow search.
Problem

Research questions and friction points this paper is trying to address.

reflective prompt optimization
self-refinement pipelines
large language models
workflow evolution
redundant instructions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Workflow-Designing Agents
Reflective Evolution
Task-Adaptive Self-Refinement
SPLIT Strategy
Prompt Optimization
🔎 Similar Papers