CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited ability of vision-language models to maintain component identity and externalize geometric relationships across heterogeneous architectural drawings—namely floor plans, sections, and elevations. To this end, we propose CrossProjection, an evaluation framework that establishes the first geometry-consistency benchmark tailored to architectural drawings. Leveraging anchor points, the framework systematically assesses model performance across three dimensions: matching, registration, and geometric localization. To ensure reproducibility and auditability, it incorporates drawing anchors, fixed-denominator scoring, and hash-locked artifacts. Experiments reveal that while GPT-5.5 achieves 82.4% accuracy on classification tasks, it exhibits fragility in unconstrained geometric localization. In contrast, human experts achieve accuracies of 87.3–93.3%, confirming the benchmark’s validity and highlighting a critical gap: current models lack robust candidate-free spatial reasoning capabilities.
📝 Abstract
Architectural drawings violate the usual assumption behind multi-view reasoning: plans and sections are cuts, while elevations are facade projections, so corresponding components change appearance in ways camera motion cannot explain. We introduce CrossProjection, an anchor-grounded diagnostic of whether vision-language models preserve component identity and externalize geometry across heterogeneous architectural views. It evaluates Matching, Registration, and Geometric Grounding through categorical judgments, candidate selection, and free point, line, and region localization. Across 23 real drawing sets and 1,954 categorical conditions per model, GPT-5.5 scores 82.4%, Qwen3-VL-32B-Instruct 62.2%, and GLM-4.5V 57.2%. A matched 200-target study crosses natural and vector-text-suppressed drawings with closed-candidate and free-geometry outputs. Candidate-supported performance is often higher, but free localization remains fragile: on natural drawings, point/region PCK@.05 is 54-76% for GPT, 8-10% for Qwen, and 14-36% for GLM; line endpoint PCK@.05 is 22%, 4%, and 0%. A coordinate grid recovers some GPT point/region precision but not lines. Three architecture-trained participants reach 87.3-93.3% categorical accuracy and 76-92% GT-region hit, supporting task feasibility rather than a population-level human ceiling. Because the categorical families do not form a same-item Matching-Registration contrast and interface controls alter multiple burdens, we avoid mechanistic claims. The supported conclusion is narrower: closed-choice or marked-element success does not entail reliable explicit geometric grounding. For drawing-guided CAD/BIM systems, categorical correctness should not be treated as evidence of candidate-free spatial reliability. Reusable on-sheet anchors, fixed-denominator scoring, and hash-locked artifacts establish an audit trail for this gap.
Problem

Research questions and friction points this paper is trying to address.

architectural drawings
geometric grounding
multi-view reasoning
vision-language models
component identity
Innovation

Methods, ideas, or system contributions that make the work stand out.

CrossProjection
Geometric Grounding
Architectural Drawings
Vision-Language Models
Multi-view Reasoning
🔎 Similar Papers
No similar papers found.