🤖 AI Summary
This study addresses the escalating storage costs driven by the proliferation of AI-generated content (AIGC), investigating the economic trade-off between persistent storage and on-demand regeneration. To this end, the authors develop a system-level cost model and propose an intermediate representation (IR) regeneration strategy operating within the latent space of diffusion models, thereby overcoming the substantial computational overhead inherent in conventional prompt-based regeneration. Furthermore, the framework integrates multi-tier caching with heterogeneous storage technologies to optimize large-scale AIGC management. Experimental evaluations demonstrate that IR regeneration reduces costs by at least twofold compared to both full storage and prompt-based approaches. When evaluated against billion-scale request traces, the proposed system halves total operational overhead while maintaining low latency.
📝 Abstract
AI-generated content is becoming a rapidly growing class of digital artifacts. Because these artifacts accumulate over time, their exponential growth creates a substantial storage, energy, and infrastructure cost for operators and society. At the same time, GPU compute cost continues to fall rapidly with each hardware generation. This divergence raises a fundamental question: when does on-demand regeneration become cheaper than persistent storage?
This paper develops a cost model for comparing persistent storage and on-demand regeneration for AI-generated artifacts. The model accounts for corpus growth, HDD and tape price trends, drive replacement, electricity, request skew, caching, generator FLOPs, and future GPU price-performance improvements. For image generation, our analysis shows that prompt-based regeneration does not become cheaper than storage until around 2040, because every cache miss must still rerun the full prompt-to-artifact generation pipeline.
We observe that widely used diffusion-based generation models operate in latent space, creating an alternative point in the cost tradeoff: instead of storing the final artifact or only the prompt, operators can store a compact intermediate representation (IR) and perform cheap on-demand decoding. Our analysis shows that caching combined with IR-based regeneration substantially reduces both storage and compute cost, making it at least 2x cheaper than both full-object storage and prompt-based regeneration even today. On a production image trace with 2.07 billion requests, the same conclusion holds: prompt-based regeneration is over 100x more expensive than storage, while IR-based regeneration reduces total cost to roughly half that of full-object storage while preserving interactive miss latency.