Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high energy consumption of agentic large language models (LLMs) in software engineering by systematically quantifying the performance trade-offs between multi-agent architectures and single-agent baselines. Through a large-scale empirical evaluation involving six open-source LLMs, two prompting strategies, and three hardware platforms, we compare accuracy, latency, and energy consumption across five task categories. Results indicate that multi-agent systems consume 6.36 times more energy on average than baselines while yielding only marginal accuracy improvements. Furthermore, 59 of the 66 optimal configurations are non-agentic or single-agent, with lightweight architectures dominating the Pareto frontier. This work reveals the energy bottlenecks inherent in multi-agent designs and proposes task-aware guidelines for sustainable architecture selection.
πŸ“ Abstract
Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across five software engineering tasks: code generation, technical debt identification, code vulnerability detection, log parsing, and log analysis. For each task, we compare LLM configurations that range from a non-agentic single-query baseline to multi-agent workflows, using six open-weight LLMs, two prompt strategies, and three hardware platforms. We assess each configuration in terms of accuracy, inference latency, and energy consumption. Our results reveal substantial trade-offs between agentic complexity and energy efficiency: multi-agent designs consume on average 6.36$\times$ as much energy and run 6.07$\times$ as long as the non-agentic baseline, with worst-case slowdowns of up to 160$\times$ for individual task--hardware pairs. Accuracy gains from additional agents are limited and task-specific: multi-agent improves average vulnerability-detection accuracy, but lightweight non-agentic and single-agent configurations still dominate the Pareto front, accounting for 59 of 66 Pareto-optimal configurations. Model and prompt choice act as task-specific levers whose effective direction varies between tasks rather than as global defaults. We translate these findings into design guidelines for sustainable, task-aware LLM-based development tools.
Problem

Research questions and friction points this paper is trying to address.

Agentic LLMs
Software Engineering
Energy Consumption
Multi-agent Systems
Sustainability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic LLMs
Energy Efficiency
Software Engineering
Multi-agent Systems
Pareto Optimization
Merve Astekin
Merve Astekin
Research Scientist, SINTEF Digital
Y
Yan Naing Tun
Singapore Management University, Singapore
A
Arda Goknil
SINTEF, Norway
E
Erik Johannes Husom
SINTEF, Norway
L
Lwin Khin Shar
Singapore Management University, Singapore
H
Hasan SΓΆzer
Ozyegin University, Turkiye
Ratnadira Widyasari
Ratnadira Widyasari
Singapore Management University
Computer science
Hui Song
Hui Song
SINTEF, Norway