🤖 AI Summary
This study addresses the challenges of fragmented evidence, the absence of unified compositional logic, and unquantified residual risks in evaluating large language model agents. We propose a systematic framework grounded in two-level temporal logic and assurance ledgers. This approach establishes a unified formalism to integrate heterogeneous checking methods by encoding rule violations as formal formulas, synthesizes multi-source evidence via information fragment theory, and introduces the concept of residuals to quantify uncovered risks. Experiments on 456 trajectories demonstrate that combining multiple methods achieves a detection rate exceeding 90%, effectively identifying and quantifying risk residuals overlooked by conventional single-method approaches. Ultimately, this framework enables the composability of evidence and facilitates comprehensive completeness auditing for LLM agent evaluation.
📝 Abstract
Methods for assessing language-model agents include rule checkers over logs, analyses of skill coverage and composition, support checkers, prefix monitors, execution gates, and rare-event estimators. Each observes a different part of a run and makes a claim of a different strength, and no common account says how these claims combine or what they leave unchecked. We give one, built from a logic, information fragments, and an assurance ledger. Requirements are formulas of a two-tier logic: an outer finite-trace temporal logic over recorded events, and an inner logic of standing over the argument structure that the agent's recorded context supports at a decision. A rule violation and an action taken on withdrawn support are thus formulas of one language. The information available to an assessor, to the agent, and to an execution gate defines fragments of that language; each checking method decides one fragment and returns a typed claim: exact, on a named test suite, a risk bound, or descriptive. The ledger merges the evidence and classifies every obligation as established, addressed but not established, or unaddressed. On 456 released $\tau^2$-bench telecom trajectories, the benchmark oracle flags 231 runs; adding a rule checker, a support stand-in, and a prefix-monitor baseline raises the union to 391, 401, and 407, each contributing flags the others miss, and a gate makes one prohibition exact on a gated deployment. The 49 unflagged runs and the unmet or unaddressed obligations form the residual. Discovery makes unknown requirements explicit, and new or stronger checkers then reduce it.