🤖 AI Summary
This study addresses the absence of evaluation benchmarks for language model agents in two-dimensional material flake field-effect transistor (FET) layout assessment by constructing the first benchmark comprising 128 tasks. Agents are required to generate GDSII-formatted layouts from microscopic image contours, with a deterministic verifier introduced to ensure geometric compliance and contour integrity. Methodologically, models such as GPT5.6-Luna are integrated with ReAct-3 and Plan-and-Execute strategies, utilizing Python scripts to generate polygon paths rendered into GDSII format, thereby supporting multi-flake and hole-containing complex tasks for the first time. Experimental results demonstrate that the optimal configuration achieves an 80.5% solution rate, 43.8% consistency, and approximately 60% expert audit acceptance, significantly outperforming single-pass planning baselines. These findings validate both the solvability of this task and the effectiveness of the proposed benchmark.
📝 Abstract
Field-effect transistor (FET) layouts on exfoliated two-dimensional flakes are typically drawn by hand for each flake, placing contacts and gates to match its position and outline in optical micrographs. To our knowledge, no executable benchmark tests whether language-model agents can perform this flake-specific construction reliably. We introduce 2D-FET-Bench V2, a benchmark of 128 layout tasks built from microscopy-derived flake contours, including hole-containing flakes and multi-flake tasks. Each task supplies a textual device specification and contour coordinates. An agent generates typed polygon and path operations rendered to GDSII. A deterministic verifier checks geometric and structural requirements, and a separate integrity check verifies that the supplied contours remain unchanged. Scripted reference layouts pass all 128 tasks, showing that every task is solvable. We evaluate six models and seven workflow and scaffold variants of GPT5.6-Luna, with five attempts per task. The best-performing configuration in the six-model panel, GPT5.6-Luna with ReAct-3, passes 62.3% of attempts and solves 80.5% of tasks at least once (coverage) and 43.8% in all five attempts (consistency). ReAct-3 exceeds the one-pass Plan-and-Execute by 27.0 pass@1 points at 2.46 times the tokens. An expert audit of one sampled verifier-passing layout per covered task, across five ReAct-3 configurations, accepts 56.4% to 63.5% of them. The benchmark evaluates geometric and structural FET layout construction.