🤖 AI Summary
This work addresses spatial occlusion and geometric inconsistency issues in educational animation code generated by large language models (LLMs) by introducing Symbolic Geometry Agents (SGA), a lightweight, plug-and-play post-processing module. SGA intercepts and partially executes LLM-generated code to construct a symbolic scene graph, enabling real-time detection and correction of spatial conflicts. Its key innovation lies in integrating a render-free Manim Visual Quality Score (MVQS) with a pluggable geometric validation mechanism, facilitating efficient assessment of spatial integrity. Evaluated on the MMMC-Code benchmark, SGA significantly improves MVQS in seven out of eight configurations, achieving a peak score of 73.11—a relative improvement of 16.1% over the baseline.
📝 Abstract
Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim. However, ensuring spatial correctness and visual legibility remains challenging, as existing frameworks emphasize pedagogical content while overlooking geometric occlusions. We propose the Symbolic Geometric Agent (SGA), a plug-and-play module for code-centric animation pipelines that intercepts LLM-generated code, performs partial execution to extract symbolic scene graphs, and applies targeted refinement when spatial conflicts are detected. We further introduce the Manim Visual Quality Score (MVQS), a deterministic rendering-free proxy for spatial integrity. Experiments on the MMMC-Code benchmark across four LLM backbones and two agentic pipelines show that SGA achieves a peak MVQS of 73.11 (Code2Video + GPT-5.1), corresponding to a 16.1% relative improvement over the raw baseline, and improves MVQS in 7 of 8 backbone x pipeline configurations.