2d-fet-bench: from spatial reasoning to fet design on flakes

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of evaluation benchmarks for language model agents in two-dimensional material flake field-effect transistor (FET) layout assessment by constructing the first benchmark comprising 128 tasks. Agents are required to generate GDSII-formatted layouts from microscopic image contours, with a deterministic verifier introduced to ensure geometric compliance and contour integrity. Methodologically, models such as GPT5.6-Luna are integrated with ReAct-3 and Plan-and-Execute strategies, utilizing Python scripts to generate polygon paths rendered into GDSII format, thereby supporting multi-flake and hole-containing complex tasks for the first time. Experimental results demonstrate that the optimal configuration achieves an 80.5% solution rate, 43.8% consistency, and approximately 60% expert audit acceptance, significantly outperforming single-pass planning baselines. These findings validate both the solvability of this task and the effectiveness of the proposed benchmark.
📝 Abstract
Field-effect transistor (FET) layouts on exfoliated two-dimensional flakes are typically drawn by hand for each flake, placing contacts and gates to match its position and outline in optical micrographs. To our knowledge, no executable benchmark tests whether language-model agents can perform this flake-specific construction reliably. We introduce 2D-FET-Bench V2, a benchmark of 128 layout tasks built from microscopy-derived flake contours, including hole-containing flakes and multi-flake tasks. Each task supplies a textual device specification and contour coordinates. An agent generates typed polygon and path operations rendered to GDSII. A deterministic verifier checks geometric and structural requirements, and a separate integrity check verifies that the supplied contours remain unchanged. Scripted reference layouts pass all 128 tasks, showing that every task is solvable. We evaluate six models and seven workflow and scaffold variants of GPT5.6-Luna, with five attempts per task. The best-performing configuration in the six-model panel, GPT5.6-Luna with ReAct-3, passes 62.3% of attempts and solves 80.5% of tasks at least once (coverage) and 43.8% in all five attempts (consistency). ReAct-3 exceeds the one-pass Plan-and-Execute by 27.0 pass@1 points at 2.46 times the tokens. An expert audit of one sampled verifier-passing layout per covered task, across five ReAct-3 configurations, accepts 56.4% to 63.5% of them. The benchmark evaluates geometric and structural FET layout construction.
Problem

Research questions and friction points this paper is trying to address.

2D-FET layout
spatial reasoning
language-model agents
benchmark
flake-specific design
Innovation

Methods, ideas, or system contributions that make the work stand out.

2D-FET-Bench
LLM agents
spatial reasoning
GDSII layout generation
ReAct framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dunzhi Zhou
University of Minnesota, Minneapolis, MN, USA
C
Chengyu Zhu
University of Minnesota, Minneapolis, MN, USA
G
Gang Qiu
University of Minnesota, Minneapolis, MN, USA
Caiwen Ding
Caiwen Ding
Associate Professor, University of Minnesota - Twin Cities
Efficient Machine LearningML for EDAComputer Architecture