Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of LLM agents to unapproved persistent side effects during action execution, which can cause erroneous outcomes to be misidentified as successful and propagated downstream. To mitigate this, we propose EffectMatch, a runtime mechanism that enables real-time verification of persistent results for the first time. Specifically, it captures state changes within controlled boundaries and employs a logical comparison engine to match them against approved states before determining whether to commit and proceed with subsequent execution, thereby bridging the gap left by existing methods in intercepting inconsistent states. Evaluated across 206 tasks, EffectMatch preserves normal execution without degradation while successfully blocking all erroneous commits and invalid propagations observed in testing.
📝 Abstract
Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notification because execution can produce persistent effects beyond the requested change. Current safeguards can approve an action or record its aftermath, but without checking the persistent result before continuation, an unapproved outcome can be accepted as success and propagated to later steps. We present EffectMatch, a runtime that collects persistent changes within a controlled execution boundary and compares them with what the application approved for the current state and execution. The comparison governs commit and dependent execution. In comparative evaluation on 206 public business tasks, EffectMatch preserved all clean executions and prevented all tested incorrect commits. Six 20-run ablations exposed the failure caused by each removed mechanism, while 80 task-topology cases preserved truthful handoffs and blocked invalid continuation. Together, these results show that EffectMatch blocks the silent acceptance and downstream propagation of persistent outcomes inconsistent with application approval.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
runtime validation
persistent outcomes
agent workflows
unapproved side effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Runtime Validation
Agent Workflows
Persistent Outcomes
EffectMatch
Controlled Execution Boundary
🔎 Similar Papers
Haoran Zhang
Haoran Zhang
PhD Student, Massachusetts Institute of Technology
Machine LearningHealthcare
H
Hengtong Zhang
Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China
Zhiyu Liang
Zhiyu Liang
Harbin Institute of Technology
Time SeriesMachine LearningFederated LearningDatabase
Y
Yu Yan
Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China
D
Decheng Zuo
Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China
Hongzhi Wang
Hongzhi Wang
IBM Almaden Research Center
Medical Image Analysis