Coding-Agent Benchmarks Should Match Their Users' Task Flows

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between existing coding agent benchmarks and real-world software engineering interaction workflows, which leads to distorted evaluations. By analyzing task flow discrepancies across 4,782 authentic sessions, this work proposes SWE-TaskFlow, a method that calibrates benchmarks to target scenarios through prompt decomposition and verifiable repository question answering. Furthermore, it formally defines and quantifies the concept of "task flow," introducing the Task Flow Alignment Score (TFAS) to demonstrate that interaction protocols themselves constitute a critical evaluation dimension. In pilot experiments on SWE-Bench Pro, multi-step sequential solving doubled computational costs without significantly improving resolution rates, thereby confirming the necessity of benchmark calibration.
📝 Abstract
The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is interactive agents, we study the sessions with at least three user messages (33% of the sample). These long sessions differ from issue-derived benchmark tasks in two ways: (i) user requests span a far wider mix of task types - questions about the project's code, planning, review, refactoring, execution - and (ii) users switch between types throughout a session. Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-to-type transitions), so no single interaction distribution is universally realistic: benchmarks should name a target use case and calibrate to measurements from it. We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories. In a pilot on 700 SWE-Bench Pro tasks, solving the task sequentially in several steps approximately doubles agent cost without a stable change in resolve rate: the interaction protocol itself is an important dimension of evaluation.
Problem

Research questions and friction points this paper is trying to address.

coding agents
benchmark evaluation
task flows
interactive agents
realistic evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Coding-Agent Benchmarks
Task Flows
SWE-TaskFlow
Prompt Splitting
TaskFlow Alignment Score
🔎 Similar Papers
No similar papers found.