Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of credible evidence in auditing self-generated execution logs of AI agents. It proposes a "second author" mechanism, advocating for external independent logging rather than internal hash chains to ensure record integrity. The research systematically evaluates the log credibility of 16 frameworks through experiments employing double-blind LLM review, human expert panels, mechanical citation verification, and external append-only logging techniques. Results demonstrate that external logging successfully detected all 28 instances of tampering and omission, whereas internal hash chains failed to identify any anomalies. This work substantiates the necessity of external independent logging as a second author, providing a more reliable tamper-proof solution for agent auditing.
📝 Abstract
An agent harness, the code that turns a model into an agent, writes its own record of each run, and that record is all a later reader gets when a run is disputed, investigated or audited. We call a record evidentiary when a reader who was not there can check it without trusting the writer. Across sixteen deployed frameworks, none writes one in full. Hearsay examines the record, not the task: five harnesses run fourteen tasks, three blinded LLM examiners and a human panel read the records, and every excerpt an examiner quotes is checked mechanically for who wrote it. First, the record lets a reader name the fault but not prove how the run went. Examiners name the right fault in 74 to 91% of 140 runs, but the fault can be proved only from two files the benchmark adds; for what happened in between, fewer than one citation in ten lands on anything the harness did not write, and the examiner with the fewest false alarms catches half of the entries we delete, rewrite or fabricate. Second, the remedy is a second author, not a stronger seal on the first. An append-only log of what passes between harness and model, kept outside the harness and read against the record in both directions, reports all 28 omissions and fabrications we made a harness commit as it ran, where a hash chain over the harness's own record passes all 28. Handed the log, examiners keep their fault verdicts but rest more of their citations on what the harness did not write. What makes a record evidence is who writes it, not what is captured.
Problem

Research questions and friction points this paper is trying to address.

agent harness
auditability
evidentiary record
trustworthiness
log integrity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Harness Auditing
Evidentiary Record
Append-only Log
Tamper Detection
LLM Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiahong Dai
Nanyang Technological University
Z
Zhuochen Yang
Nanyang Technological University
Pengyang Shao
Pengyang Shao
Hefei University of Technology
Recommender SystemsCognitive Diagnosis
K
Kelvin Ng
Nanyang Technological University
Zhongyi Liu
Zhongyi Liu
Ant Group
Information RetrievalRecommender SystemsNatural Language Processing
C
Chengquan Ju
Nanyang Technological University
Yuting He
Yuting He
Foundation Medicine Inc.
Precision MedicineBiomarker and CDxCancer GenomicsMachine LearningData Mining
B
Bo Hu
Nanyang Technological University