π€ AI Summary
This study addresses the challenges of hardware heterogeneity, system complexity, and the absence of unified evaluation benchmarks in edge-to-cloud deployments of agent applications. To this end, it proposes AgenticOps, a framework that enables automated management across the entire agent lifecycle. The framework establishes an end-to-end pipeline integrating distributed deployment with telemetry collection, and introduces an LLM-as-a-Judge mechanism to facilitate semantic-level automated evaluation and reproducible report generation. Experimental results demonstrate that this approach significantly reduces manual operational overhead while providing a standardized experimental paradigm and an efficient evaluation methodology for agent systems.
π Abstract
Deploying LLM-enabled agentic applications across the Edge-to-Cloud continuum remains challenging due to hardware heterogeneity, deployment complexity, limited observability, and the lack of systematic evaluation methods. Existing solutions address agent development, observability, or benchmarking separately, offering limited support for the full lifecycle of distributed agentic applications. This paper presents AgentWare, an AgenticOps framework that automates the provisioning, deployment, observability, and evaluation of agentic applications across Edge-to-Cloud infrastructures. AgentWare introduces an end-to-end lifecycle pipeline that automatically prepares heterogeneous execution environments, transforms user-defined agent implementations into distributed applications, deploys agent components across the continuum, and performs unified collection of execution traces, infrastructure telemetry, and evaluation metrics. The framework further supports automated semantic evaluation through LLM-as-a-Judge workflows and generates reproducible reports covering correctness, performance, resource utilization, and energy consumption. We demonstrate the applicability of AgentWare through a distributed book assistant agent deployed across real Edge-to-Cloud infrastructure under multiple deployment and model configurations. The results show that AgentWare enables systematic experimentation and evaluation of distributed agentic applications while significantly reducing the manual effort required for deployment, instrumentation, and analysis.