🤖 AI Summary
This study addresses the lack of editable construction rules in existing material generation methods by proposing a hierarchical procedural language that unifies PBR material appearance and construction logic into compact, executable code. The approach defines material channels through compositing alpha-masked layers and leverages pretrained language models for controllable text-to-material generation. Optimization is achieved via parser-based repair, a preview-critic mechanism, and noise seed search. The generated programs average only 21 lines while retaining named parameters, enabling direct editing without fine-tuning. Experiments demonstrate that this method comprehensively outperforms diffusion-based baselines across benchmarks; notably, even the initial programs surpass competing approaches, achieving a 59.2% preference rate in blind user studies.
📝 Abstract
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physically based rendering (PBR) channels, making dependencies between patterns, color, and relief explicit. A standalone interpreter evaluates the program into material maps, while the source retains named fields and layer parameters for subsequent authoring. Without task-specific fine-tuning, our pipeline uses parser-guided repair and preview-based critique to revise material designs, then searches noise seeds while keeping each candidate's remaining source fixed. On a curated benchmark of 141 prompts evaluated with six backbones, our best-performing configuration achieves higher mean scores than three diffusion baselines on all four flat-layout prompt-alignment metrics. Its initial programs already exceed all three baselines on mean BLIPScore, before critique or seed search. Retained programs have a median length of 21 lines when pooled across backbones. In a blind four-way comparison involving 30 participants and 20 prompts, our renders receive 59.2% of choices, compared with 19.3% for the most-preferred baseline. Compact executable programs thus offer a way to generate prompt-aligned materials while retaining their construction as part of the asset.