Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that relying solely on final outputs fails to capture process-level behavioral drift during the skill evolution of enterprise AI agents. To this end, it proposes a continuous evaluation framework integrating both outcome and process assessments. The method independently computes ground-truth references and designs reusable test templates, combining programmatic checks with constrained LLM judges to enable fine-grained monitoring of tool selection, parameter configuration, and execution order. Furthermore, dependency attribution techniques are introduced to substantially reduce false-positive noise. Experimental results demonstrate that 92.6% of runs passing final numerical checks still exhibit process deviations, while dependency attribution reduces the average number of failed checks from 6.34 to 2.65, effectively revealing differences in specification sensitivity.
📝 Abstract
Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing all applicable final numerical checks, 162 (92.6 percent; Wilson 95 percent CI: 87.7-95.6 percent) contained another evaluator-detected deviation. Under a broader seven-check final-state definition, 151 of 164 passing runs (92.1 percent; 95 percent CI: 86.9-95.3 percent) still violated a trajectory check. Dependency attribution reduced a mean of 6.34 failed checks per run to 2.65 roots. Specification sensitivity varied by model and harness, with exploratory bootstrap interaction intervals excluding zero for all three Revenue comparisons and one of three Productivity comparisons. Runtime-resolved templates provided reusable regression coverage across the evaluated configurations; longitudinal validation under actual API evolution remains future work.
Problem

Research questions and friction points this paper is trying to address.

Enterprise AI Agents
Process-Level Evaluation
Behavioral Drift
Continuous Evaluation
Agent Skills
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous Evaluation
Process-Level Evaluation
AI Agent Skills
Dependency Attribution
LLM Judge
🔎 Similar Papers
No similar papers found.
N
Ngoc Phuoc An Vo
IBM
A
Aarya Doshi
Georgia Institute of Technology
V
Vadim Sheinin
IBM