APEX: Adaptive Principle EXtraction A Three-Layer Self-Evolution Framework for Production AI Agents

📅 2026-06-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current approaches to self-improvement in AI agents are largely confined to prompt template optimization, overlooking the co-evolution of behavioral principles and workflow structures, thereby limiting overall performance gains. This work proposes APEX, a three-tiered self-evolution framework that, for the first time, enables joint optimization of prompts, behavioral principles, and workflow topology. The framework drives a lightweight LLM evolution process through failure-mode repair, successful-trajectory distillation, and a structural fitness-based selection mechanism. Experimental results demonstrate that a single evolutionary cycle improves the health score by 90% to 0.570, uncovers six reusable new principles, and yields an optimal workflow topology with a score of 0.900—all achieved within approximately 270 seconds of local computation time.
📝 Abstract
Self-improvement in AI agents has emerged as a key research frontier: systems that modify their own prompts, workflows, and decision rules based on accumulated operational experience. The state-of-the-art Self-Harness framework [1] achieves 14--21% improvement on Terminal-Bench-2.0 by mining failure clusters and patching the agent harness. However, Self-Harness optimises only one dimension -- the prompt harness -- leaving behavioural principles and workflow topology unchanged. We propose APEX (Adaptive Principle EXtraction), a three-layer co-evolution framework that simultaneously evolves: (L1) the harness via failure-mode patching, (L2) behavioural principles via success-trace distillation [2], and (L3) the agent workflow topology via structural fitness-based selection [6]. We implement APEX on Joe [13], a production-grade super AI Agent built on NVIDIA Nemotron and designed as an Edge AI Agent Factory for the NVIDIA Agent Challenge 2026, managing a 15-node compute fleet using 114 real task traces collected over 18 days. APEX achieves an APEX Health Score of 0.570 (+90% vs. baseline 0.300) in a single evolutionary run, distilling 6 novel reusable principles and selecting a research-first workflow topology scoring 0.900 (+20%). Our results demonstrate that multi-dimensional co-evolution substantially outperforms single-axis harness optimisation, at a cost of only 4 LLM calls (~270 s) on a local qwen2.5-coder:32b instance.
Problem

Research questions and friction points this paper is trying to address.

AI agent self-improvement
prompt harness
behavioural principles
workflow topology
multi-dimensional co-evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive principle extraction
self-evolution framework
multi-dimensional co-evolution
behavioral principle distillation
workflow topology optimization