🤖 AI Summary
This study addresses the challenge that relying solely on final outputs fails to capture process-level behavioral drift during the skill evolution of enterprise AI agents. To this end, it proposes a continuous evaluation framework integrating both outcome and process assessments. The method independently computes ground-truth references and designs reusable test templates, combining programmatic checks with constrained LLM judges to enable fine-grained monitoring of tool selection, parameter configuration, and execution order. Furthermore, dependency attribution techniques are introduced to substantially reduce false-positive noise. Experimental results demonstrate that 92.6% of runs passing final numerical checks still exhibit process deviations, while dependency attribution reduces the average number of failed checks from 6.34 to 2.65, effectively revealing differences in specification sensitivity.
📝 Abstract
Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing all applicable final numerical checks, 162 (92.6 percent; Wilson 95 percent CI: 87.7-95.6 percent) contained another evaluator-detected deviation. Under a broader seven-check final-state definition, 151 of 164 passing runs (92.1 percent; 95 percent CI: 86.9-95.3 percent) still violated a trajectory check. Dependency attribution reduced a mean of 6.34 failed checks per run to 2.65 roots. Specification sensitivity varied by model and harness, with exploratory bootstrap interaction intervals excluding zero for all three Revenue comparisons and one of three Productivity comparisons. Runtime-resolved templates provided reusable regression coverage across the evaluated configurations; longitudinal validation under actual API evolution remains future work.