🤖 AI Summary
This work addresses the lack of systematic evaluation of large language models as intelligent IDE agents in realistic multilingual, full-stack development environments. We propose a Docker-based, IDE-native benchmarking framework that simulates authentic development workflows through structured tool interfaces—including code search, structured editing, and full-stack testing—and constructs 80 real-world tasks across eight private, unpublished codebases spanning C/C++, Java, and MERN stacks. These tasks encompass feature implementation, bug fixing, refactoring, and performance optimization. Our framework establishes, for the first time under contamination-free conditions, a systematic linkage between agent intent and project-level modification outcomes, thereby introducing an evaluation paradigm that closely mirrors real-world engineering collaboration and enables comprehensive, reliable assessment of AI-powered IDE agents.
📝 Abstract
IDE-Bench is a comprehensive framework for evaluating AI IDE agents on real-world software engineering tasks through an IDE-native tool interface. We present a Dockerized test harness that goes beyond raw terminal execution, granting models a structured tool ecosystem that represents AI-native IDEs like Cursor and Windsurf. By providing high-level abstractions for codebase search, structured file editing, and tools for testing full-stack applications, IDE-Bench evaluates an agent's ability to act as a true engineering collaborator. For evaluation and to prevent training data contamination, we created 80 tasks across eight never-published repositories spanning C/C++, Java, and MERN stacks, representing modern tech stack production scenarios, including feature implementation, bug fixing, refactoring, and performance optimization that mirror daily developer workflows in private codebases. Our benchmark is the first to systematically correlate agent-reported intent with successful project-level modifications in a multi-language, full-stack environment on completely uncontaminated code. We release IDE-Bench and a public leaderboard at: https://ide-bench.com.