SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant misalignment between existing coding assistant benchmarks and real-world scenarios, particularly regarding task long-horizon complexity and the authenticity of multi-turn interactions. To bridge this gap, we propose a weak-to-strong automated synthesis pipeline for constructing long-horizon tasks and derive four user personas from authentic interaction logs to establish user-simulating agents that replicate realistic multi-turn behaviors. Evaluations using this benchmark reveal that current models provide critically insufficient support for non-expert users, achieving pass rates below 25%. Furthermore, our analysis identifies precise questioning, accurate localization, and effective repair as key capability bottlenecks. By systematically addressing these limitations, this work narrows the divide between benchmark evaluation and practical deployment in AI-assisted software engineering.
📝 Abstract
Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions. On average, models pass over 75% of tests for requested functionality with software architects, but fewer than 25% with non-coders. These results show that current coding assistants still fall short of enabling reliable coding for non-coders. We further analyze the reasons for this gap and identify asking right, finding right, and fixing right as key capabilities during interaction.
Problem

Research questions and friction points this paper is trying to address.

coding assistants
benchmark evaluation
long-horizon tasks
multi-turn interaction
LLM agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-Horizon Tasks
Multi-Turn Interaction
Weak-to-Strong Synthesis
User Simulation Agent
Coding Assistants
🔎 Similar Papers
No similar papers found.
H
Hexuan Deng
Tencent Hy AI Data
Y
Yue Wang
Beijing Zhongguancun Academy
W
Wenyu Jiang
Tencent Hy AI Data
C
Cheng Yang
Tencent Hy AI Data
Haolin Yang
Haolin Yang
University of Chicago
large language modelsnatural language processing
Z
Zhaohua Zhang
Tencent Hy AI Data
C
Chenchen Zhao
Tencent Hy AI Data
Beiduo Chen
Beiduo Chen
ELLIS PhD Student, Ludwig-Maximilians-Universität München
LinguisticsNatural Language Processing
M
Muxi Chen
Tencent Hy AI Data
S
Sa Zhu
Tencent Hy AI Data
G
Geyuan Zhu
Tencent Hy AI Data
Jianhuan Zhuo
Jianhuan Zhuo
Institute of Information Engineering, Chinese Academy of Sciences
Representation LearningRecommendation System
Q
Qiuyong Xiao
Tencent Hy AI Data
Tianwen Jiang
Tianwen Jiang
Harbin Institute of Technology
Knowledge GraphInformation ExtractionNatural Language Processing
J
Jihong Zhang
Tencent Hy AI Data
X
Xuebo Liu
Tencent Hy AI Data