THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the bottleneck in analyzing analog integrated circuit GDSII layout files through natural language interaction by proposing the THEIA dataset and evaluation benchmark. Methodologically, it introduces the first multimodal question-answering dataset tailored for analog IC physical layouts, integrating GDSII parsing with multimodal annotation techniques to perform domain-specific fine-tuning of vision-language models (VLMs). Experimental results demonstrate that the fine-tuned model outperforms state-of-the-art general-purpose VLMs by up to 73% across five real-world tasks, revealing substantial comprehension gaps of general models in specialized domains. Ultimately, this work significantly enhances designers’ capability to intuitively query and analyze circuit layouts.
πŸ“ Abstract
The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDSII file represents the industry-standard database containing the ultimate and most accurate source of information of the analog circuit, encapsulating the complex physical geometries and parasitic realities that define tape out performance. This paper proposes THEIA, a novel dataset containing thousands of layout images paired with question-answer conversations, along with a benchmark that employs a fine-tuned vision-language model (VLM) to analyze GDSII files of analog circuits, enabling designers to interact with and query physical layouts as intuitive, meaningful entities. Experimental results using thousands of analog designs across five realistic tasks demonstrate that the proposed fine-tuned VLM outperforms state-of-the-art general-purpose VLMs by a significant margin (up to 73%), highlighting a fundamental gap between general-purpose multimodal reasoning and domain-specific layout understanding.
Problem

Research questions and friction points this paper is trying to address.

analog integrated circuits
layout analysis
vision-language model
GDSII
multimodal reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Model
Analog Integrated Circuits
GDSII
Multimodal Dataset
Layout Analysis