🤖 AI Summary
Existing evaluations of LLM-based agents focus solely on tool output correctness, failing to detect "functional forgery" risks where malicious side effects are concealed. This work proposes SINGED, a benchmark that systematically audits agents' awareness of covert side effects and quantifies execution risks across strategies using controlled experiments, cross-candidate comparison, and process oracle verification. Results reveal that 45% of top-ranked candidates involve forged executions, while cross-comparison reduces the tiered failure rate to 4.2%, confirming that most models execute forged code when no alternative sources are available. By exposing the non-identifiability gap between outcomes and execution paths, this study highlights critical security blind spots inherent in single-output evaluation paradigms for autonomous agents.
📝 Abstract
Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the requested output but add an effect forbidden by the task contract. We introduce SINGED (Source Integrity and the Nonidentifiability Gap in Execution Decisions for LLM Agents), a controlled benchmark covering five primary and two held-out task families. It varies displayed rank, evidence depth, decision policy, model release, and agent configuration, while task and process oracles verify the artifact and execution path. Across 7,549 audited trials, the randomized-rank study finds counterfeit execution in 45% (27/60) of rank-one trials and none at later ranks. Cross-candidate comparison eliminates shallow failures and reduces layered failures from 15.7% to 4.2%, but leaves dependency failures; its benefit is uncertain on unseen effects and public-package structures. Moreover, seven releases with no counterfeit executions when benign alternatives are available execute the counterfeit in 55/175 single-source cells after alternatives are removed. SINGED thus exposes a rank-, evidence-, and choice-sensitive outcome-to-execution gap: evaluation must connect correct outputs to execution paths.