Enhancing Building Semantics Preservation in AI Model Training with Large Language Model Encodings

📅 2026-02-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional encoding schemes—such as one-hot encoding—in capturing fine-grained semantic relationships among architectural component subtypes, which hinders semantic understanding in artificial intelligence applications within the AECO (Architecture, Engineering, Construction, and Operations) domain. To overcome this, the authors propose a novel semantic encoding approach that integrates large language model (LLM) embeddings with Matryoshka-compressed representations. Specifically, they combine LLM-generated semantic vectors—derived from GPT and LLaMA—with GraphSAGE, a graph neural network, for semantic classification across 42 architectural component categories. Experimental results demonstrate that the proposed method substantially outperforms the one-hot baseline, with the compressed LLaMA-3 embeddings achieving a weighted F1 score of 0.8766—an improvement of approximately 2.9%—while effectively preserving fine-grained semantic relationships and enabling efficient modeling.

Technology Category

Machine Learning: Deep Generative Models & AutoencodersCognitive Modeling & Cognitive Systems: Agent ArchitecturesNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd work
📝 Abstract
Accurate representation of building semantics, encompassing both generic object types and specific subtypes, is essential for effective AI model training in the architecture, engineering, construction, and operation (AECO) industry. Conventional encoding methods (e.g., one-hot) often fail to convey the nuanced relationships among closely related subtypes, limiting AI's semantic comprehension. To address this limitation, this study proposes a novel training approach that employs large language model (LLM) embeddings (e.g., OpenAI GPT and Meta LLaMA) as encodings to preserve finer distinctions in building semantics. We evaluated the proposed method by training GraphSAGE models to classify 42 building object subtypes across five high-rise residential building information models (BIMs). Various embedding dimensions were tested, including original high-dimensional LLM embeddings (1,536, 3,072, or 4,096) and 1,024-dimensional compacted embeddings generated via the Matryoshka representation model. Experimental results demonstrated that LLM encodings outperformed the conventional one-hot baseline, with the llama-3 (compacted) embedding achieving a weighted average F1-score of 0.8766, compared to 0.8475 for one-hot encoding. The results underscore the promise of leveraging LLM-based encodings to enhance AI's ability to interpret complex, domain-specific building semantics. As the capabilities of LLMs and dimensionality reduction techniques continue to evolve, this approach holds considerable potential for broad application in semantic elaboration tasks throughout the AECO industry.
Problem

Research questions and friction points this paper is trying to address.

building semantics
AI model training
semantic representation
AECO industry
subtype relationships
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM embeddings
building semantics
GraphSAGE
semantic encoding
Matryoshka representation
💼 Related Jobs
No related jobs found.
S
Suhyung Jang
Building Informatics Group, Department of Architecture and Architectural Engineering, Yonsei University, Republic of Korea
G
Ghang Lee
Building Informatics Group, Department of Architecture and Architectural Engineering, Yonsei University, Republic of Korea; Institute for Advanced Studies, Technical University of Munich, Germany
J
Jaekun Lee
Building Informatics Group, Department of Architecture and Architectural Engineering, Yonsei University, Republic of Korea
H
Hyunjun Lee
Building Informatics Group, Department of Architecture and Architectural Engineering, Yonsei University, Republic of Korea