Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML

📅 2025-09-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study evaluates the applicability of large language models (LLMs) to generate **functional and maintainable code** in highly specialized, closed industrial software environments—exemplified by ASML’s ecosystem—facing two key challenges: stringent domain-specific constraints and cross-module code dependencies. We propose a **customized evaluation framework**, introducing the novel metric **build@k**, which quantifies compilation and integration success rates of generated code within real industrial repositories, and establish the first code-generation benchmark tailored to ASML’s proprietary ecosystem. Leveraging few-shot prompting and chain-of-thought strategies, we conduct systematic comparisons across general-purpose and code-specialized LLMs at multiple scales, using both matching-based and execution-based evaluation. Results show that prompt engineering substantially improves build success (few-shot + CoT yields optimal performance), and model specialization benefits exhibit family-level dependency. Our core contributions are an industrial-grade evaluation paradigm for deployable code generation and the first empirically grounded benchmark for this domain.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language models have shown impressive performance in various domains, including code generation across diverse open-source domains. However, their applicability in proprietary industrial settings, where domain-specific constraints and code interdependencies are prevalent, remains largely unexplored. We present a case study conducted in collaboration with the leveling department at ASML to investigate the performance of LLMs in generating functional, maintainable code within a closed, highly specialized software environment. We developed an evaluation framework tailored to ASML's proprietary codebase and introduced a new benchmark. Additionally, we proposed a new evaluation metric, build@k, to assess whether LLM-generated code successfully compiles and integrates within real industrial repositories. We investigate various prompting techniques, compare the performance of generic and code-specific LLMs, and examine the impact of model size on code generation capabilities, using both match-based and execution-based metrics. The findings reveal that prompting techniques and model size have a significant impact on output quality, with few-shot and chain-of-thought prompting yielding the highest build success rates. The difference in performance between the code-specific LLMs and generic LLMs was less pronounced and varied substantially across different model families.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLM-generated code functionality in industrial proprietary settings
Assessing code maintainability and integration within specialized software environments
Investigating impact of prompting techniques and model size on output quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tailored evaluation framework for proprietary codebase
Introduced build@k metric for compilation success
Investigated prompting techniques and model size impact
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yash Mundhra
Delft University of Technology, Delft, Netherlands
M
Max Valk
ASML, Veldhoven, Netherlands
Maliheh Izadi
Maliheh Izadi
Assistant Professor @ Delft University of Technology, The Netherlands
Software engineeringEvaluationAI4SELLM4CodeAgents