🤖 AI Summary
This study addresses the challenges of policy compliance, latent side effects, and long-horizon state tracking faced by LLM agents operating in partially observable enterprise workflows. We propose E-Ledger, a multi-agent framework that achieves secure and persistent execution control through a code review layer and a dynamic world ledger. Furthermore, we introduce WorldAbduct, an abductive evolution mechanism that leverages multi-agent collaboration, evidence-driven state tracking, and four-view trajectory diagnostics to infer latent rules from execution traces for verified integration, thereby transcending conventional static constraints. Experimental results demonstrate that our approach improves safe task completion rates by 5–15% over the strongest baselines on enterprise benchmarks, while exhibiting strong generalization capabilities in scientific scenarios.
📝 Abstract
Large language model agents can invoke tools fluently, but enterprise workflows demand more than selecting the right tools: actions must strictly comply with organizational policies, tool feedback often conceals hidden side effects under partial observability, and long-horizon tasks require persistent state tracking across multiple records. To address these challenges, we introduce E-Ledger, a multi-agent harness for safe and persistent execution. E-Ledger employs a code approval layer that checks every proposed action against policy before execution, and maintains a world ledger of verified hidden rules alongside evidence-backed dynamic state. Because hidden rules are typically unknown a priori, we further propose WorldAbduct, an abductive, world-model-driven harness evolution framework. WorldAbduct diagnoses execution trajectories across four complementary views (state consistency, world-observation gap, policy-gate correctness, and goal judgment) to hypothesize latent rules, and verifies them through targeted abductive interactions before integrating them into the ledger. On the enterprise benchmark World of Workflows, E-Ledger with WorldAbduct improves safe task completion across four LLM backbones, outperforming the strongest evolution baseline by 5--15 percentage points. Experiments in ScienceWorld and DiscoveryWorld further show that abductive harness evolution carries over to scientific environments. Our code is available at https://github.com/HKUST-KnowComp/E-LEDGER-WorldAbduct.