🤖 AI Summary
This study addresses the frequent failures of large language models (LLMs) in generating visualization domain-specific languages (DSLs), which stem from a fundamental mismatch between human-designed constraints and model characteristics. We systematically evaluate ten JSON-based DSLs across 41 tasks, integrating multi-model benchmarking, rendering validation, and qualitative coding analysis to identify four recurring failure modes strongly correlated with specific DSL features. Our findings reveal the mechanisms through which existing DSLs prove unfriendly to LLMs. Building on these insights, we propose design guidelines for next-generation, LLM-oriented DSLs. This work provides both empirical evidence and actionable design principles for optimizing AI-assisted visualization programming, ultimately bridging the gap between DSL engineering and the generative capabilities of contemporary language models.
📝 Abstract
As LLMs take up the role of authoring charts using visualization domain-specific languages (DSLs), the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model. To understand how LLMs might work better with DSLs, we explore where and how they fail with current DSL designs. We evaluate 10 JSON-style visualization DSLs with 41 tasks across 3 LLMs, then assess the generated specifications with JSON and rendering checks, and qualitative coding of failed cases. Analyzing how this specification generation process fails, we identify four recurring failure patterns, link each to specific DSL features, and discuss design considerations for future DSL designs.