IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks

📅 2026-01-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of systematic evaluation of large language models as intelligent IDE agents in realistic multilingual, full-stack development environments. We propose a Docker-based, IDE-native benchmarking framework that simulates authentic development workflows through structured tool interfaces—including code search, structured editing, and full-stack testing—and constructs 80 real-world tasks across eight private, unpublished codebases spanning C/C++, Java, and MERN stacks. These tasks encompass feature implementation, bug fixing, refactoring, and performance optimization. Our framework establishes, for the first time under contamination-free conditions, a systematic linkage between agent intent and project-level modification outcomes, thereby introducing an evaluation paradigm that closely mirrors real-world engineering collaboration and enables comprehensive, reliable assessment of AI-powered IDE agents.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
IDE-Bench is a comprehensive framework for evaluating AI IDE agents on real-world software engineering tasks through an IDE-native tool interface. We present a Dockerized test harness that goes beyond raw terminal execution, granting models a structured tool ecosystem that represents AI-native IDEs like Cursor and Windsurf. By providing high-level abstractions for codebase search, structured file editing, and tools for testing full-stack applications, IDE-Bench evaluates an agent's ability to act as a true engineering collaborator. For evaluation and to prevent training data contamination, we created 80 tasks across eight never-published repositories spanning C/C++, Java, and MERN stacks, representing modern tech stack production scenarios, including feature implementation, bug fixing, refactoring, and performance optimization that mirror daily developer workflows in private codebases. Our benchmark is the first to systematically correlate agent-reported intent with successful project-level modifications in a multi-language, full-stack environment on completely uncontaminated code. We release IDE-Bench and a public leaderboard at: https://ide-bench.com.
Problem

Research questions and friction points this paper is trying to address.

IDE agents
large language models
software engineering tasks
full-stack development
codebase evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

IDE-Bench
AI IDE agents
tool-augmented LLMs
uncontaminated evaluation
full-stack software engineering
🔎 Similar Papers
No similar papers found.
S
Spencer Mateega
AfterQuery, San Francisco, CA, US
J
Jeff Yang
AfterQuery, San Francisco, CA, US
T
Tiana Costello
AfterQuery, San Francisco, CA, US
S
Shaurya Jadhav
AfterQuery, San Francisco, CA, US
N
Nicole Tian
AfterQuery, San Francisco, CA, US
A
Agustin Garcinuno
AfterQuery, San Francisco, CA, US