Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows
This study addresses the high energy consumption of agentic large language models (LLMs) in software engineering by systematically quantifying the performance trade-offs between multi-agent architectures and single-agent baselines. Through a large-scale empirical evaluation involving six open-source LLMs, two prompting strategies, and three hardware platforms, we compare accuracy, latency, and energy consumption across five task categories. Results indicate that multi-agent systems consume 6.36 times more energy on average than baselines while yielding only marginal accuracy improvements. Furthermore, 59 of the 66 optimal configurations are non-agentic or single-agent, with lightweight architectures dominating the Pareto frontier. This work reveals the energy bottlenecks inherent in multi-agent designs and proposes task-aware guidelines for sustainable architecture selection.