🤖 AI Summary
This work investigates the feasibility and economic efficacy of large language models (LLMs) as autonomous agents in real-world freelance software development and data analysis tasks. To this end, we introduce the first programmatically verifiable, price-annotated synthetic benchmark—synthesized from Kaggle outsourcing data—with standardized monetary labels. We propose an automated evaluation framework centered on monetized revenue (USD), integrating structured I/O correctness verification, price prediction modeling, and cross-model comparison across Claude 3.5 Haiku, GPT-4o-mini, Qwen 2.5, and Mistral. Experimental results show that Claude 3.5 Haiku achieves the highest total revenue ($1.52M), substantially outperforming competitors; inter-model revenue disparities reveal a reliability hierarchy for complex, open-ended tasks. This study pioneers the integration of economic value quantification into LLM agent evaluation—jointly optimizing task success rate and real-world financial return—thereby significantly enhancing scalability, reproducibility, and practical relevance of agent benchmarks.
📝 Abstract
This study explores Large Language Models (LLMs) as autonomous agents for real-world tasks, including freelance software development. This work presents a new benchmark that evaluates LLMs on freelance programming and data analysis tasks derived from economic data. We construct the benchmark using synthetic tasks created from a Kaggle Freelancer dataset of job postings, with all job prices standardized to USD (median fixed-project price around $250, and an average of $306). Each task is accompanied by structured input-output test cases and an estimated price tag, enabling automated correctness checking and a monetary performance valuation. This approach is inspired by OpenAI's recent SWE-Lancer benchmark (1,400 real Upwork tasks worth $1M total). Still, our framework simplifies evaluation using programmatically testable tasks and predicted price values, making it highly scalable and repeatable. On this benchmark, we evaluate four modern LLMs - Claude 3.5 Haiku, GPT-4o-mini, Qwen 2.5, and Mistral. We report each model's accuracy (task success rate and test-case pass rate) and the total"freelance earnings"it achieves (sum of prices of solved tasks). Our results show that Claude 3.5 Haiku performs best, earning approximately $1.52 million USD, followed closely by GPT-4o-mini at $1.49 million, then Qwen 2.5 ($1.33M) and Mistral ($0.70M). We analyze the distribution of errors per task and observe that the strongest models solve the most tasks and rarely fail completely on any project. We discuss the implications of these results for the feasibility of AI as a freelance developer, the advantages and limitations of our automated benchmark approach, and the gap between performance on structured tasks versus the true complexity of real-world freelance jobs.