🤖 AI Summary
Multi-agent creative systems often suffer from output homogenization due to LLMs’ inherent uniformity, hindering their ability to meet brand- and art-specific diversity requirements. To address this, we propose “Spark”, a persona-driven agent architecture that employs a persona-conditioned system prompt library to induce behavioral differentiation among agents. Spark integrates a multi-agent collaborative workflow with an LLM-as-a-judge automated evaluation framework, calibrated against human expert annotations (gold-standard references). Its core innovation lies in the explicit embedding of persona modeling into system prompt design, enabling controllable and scalable regulation of creative diversity. Experimental results demonstrate that Spark achieves an average improvement of 4.1 points on a 1–10 diversity scale, narrowing the gap with human expert ratings to only 1.0 point—significantly bridging the performance divide between AI-generated and professional creative outputs.
📝 Abstract
Creative services teams increasingly rely on large language models (LLMs) to accelerate ideation, yet production systems often converge on homogeneous outputs that fail to meet brand or artistic expectations. Art of X developed persona-conditioned LLM agents -- internally branded as "Sparks" and instantiated through a library of role-inspired system prompts -- to intentionally diversify agent behaviour within a multi-agent workflow. This white paper documents the problem framing, experimental design, and quantitative evidence behind the Spark agent programme. Using an LLM-as-a-judge protocol calibrated against human gold standards, we observe a mean diversity gain of +4.1 points (on a 1-10 scale) when persona-conditioned Spark agents replace a uniform system prompt, narrowing the gap to human experts to 1.0 point. We also surface evaluator bias and procedural considerations for future deployments.