🤖 AI Summary
This study addresses the limited observability and lack of reproducibility in human analysis caused by massive outputs during long-horizon agent evaluation. To overcome these challenges, this work proposes Transect, an open-source tool built upon Inspect Scout that introduces a flexible, customizable pipeline for transcript analysis. By aligning events, token consumption, and behavioral labels along a unified timeline, and integrating multi-agent systems, large language model judges, and structured visualization techniques, it generates navigable analytical reports. This approach preserves the flexibility of human analysis while ensuring interpretive traceability and scientific rigor. Demonstrating its efficacy, the tool successfully processed nearly 13 million tokens of AI research and development evaluation data, precisely identifying a behavioral pattern wherein agents prioritize operational execution over hypothesis generation.
📝 Abstract
Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.