coding standards enforcement

Designs, implements, and evaluates the processes, rules, tools, and metrics that define and enforce coding standards so software artifacts meet production-quality, maintainability, and style requirements. Work includes authoring style guides and code-review policies, building or integrating linters, static/dynamic analysis, CI gating, quality dashboards, automated enforcement and measurement, and editor tooling such as code-completion and refactoring support to drive clean, compliant code.

codingstandardsenforcement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Checkstyle+: Reducing Technical Debt Through The Use of Linters with LLMs

Oct 27, 2025
ED
Ella Dodor
🏛️ University of California, Irvine

Traditional static analysis tools (e.g., Checkstyle) struggle to detect code style issues requiring deep semantic understanding. This paper proposes the first hybrid code style detection framework that integrates large language models (LLMs) into the Checkstyle pipeline, overcoming inherent limitations of rule-based engines in expressive power and contextual modeling. The method leverages Checkstyle for syntactic and structural validation while employing LLMs—guided by prompt engineering—to perform context-aware, semantic-level identification of style violations (e.g., naming intent consistency, logical block readability). Evaluated on 380 real-world Java source files, our framework achieves a 27.3% improvement in detection accuracy and a 41.6% increase in coverage over standard Checkstyle for complex semantic style rules. These gains significantly enhance code readability and maintainability, demonstrating the efficacy of combining lightweight static analysis with LLM-driven semantic reasoning in practical code quality assurance.

Augmenting linters with LLMs to detect nuanced style violationsIdentifying code style rules requiring semantic understanding beyond static analysisImproving detection of style violations in Java programs using hybrid approach

This study addresses the widespread violations of coding styles and best practices in open-source Java projects, which adversely affect code quality and maintainability. It presents the first large-scale longitudinal analysis of 1,036 popular GitHub Java repositories, systematically examining the evolution and compliance of their code style through static code analysis, temporal tracking, and adherence to established guidelines such as the Google Java Style Guide. The findings reveal that missing Javadoc documentation and inconsistent naming conventions are the most prevalent issues, with numerous projects violating style rules that fall outside the scope of conventional static analysis tools. Although some projects explicitly claim conformance to style guides, their overall compliance remains only marginally better. This work highlights critical blind spots in current automated detection approaches and provides empirical evidence to inform strategies for improving code quality in open-source ecosystems.

best practicescode stylecoding standards

MaintainCoder: Maintainable Code Generation Under Dynamic Requirements

Mar 31, 2025
ZW
Zhengren Wang
🏛️ Peking University | McGill University | Institute for Advanced Algorithms Research

This work addresses the critical limitation of existing code generation systems—neglect of maintainability and poor adaptability to dynamic requirement changes—by pioneering maintainability as the primary optimization objective. We propose a novel code generation framework designed for continuous evolution, emphasizing high cohesion, low coupling, and easy adaptability. Methodologically, we (1) introduce MaintainBench, the first dynamic maintainability evaluation benchmark; (2) integrate waterfall-style phased governance, design-pattern-driven architectural generation, and multi-agent collaborative reasoning; and (3) incorporate a quantitative dynamic maintenance cost assessment model. Experimental results demonstrate a 14–30% improvement in maintainability metrics on MaintainBench, while simultaneously achieving superior pass@k functional correctness over baseline methods. All code and the MaintainBench benchmark are publicly released.

Addressing maintainability gaps in existing code generation systemsEnhancing code maintainability under dynamic requirementsReducing rework through systematic cohesion and coupling improvements

This study addresses the lack of systematic evaluation of non-functional quality—specifically security, maintainability, and performance efficiency—in code generated by large language models (LLMs). Grounded in the ISO/IEC 25010 standard, it integrates a systematic literature review, dual-industry workshops, and multi-model empirical experiments (GPT-4, Claude, CodeLlama) to conduct multidimensional quality analysis on real-world software defect-fix patches. It introduces the first non-functional quality assessment framework reconciling academic rigor with industrial relevance, uncovering significant trade-offs among the three quality attributes and exposing gaps between LLM outputs and actual engineering requirements—including technical debt accumulation. Results demonstrate that functional correctness does not imply high non-functional quality, and that model architecture and optimization strategies yield markedly divergent outcomes across non-functional dimensions. The work provides both theoretical foundations and actionable guidelines for designing robust quality assurance mechanisms for LLM-generated code.

Addressing quality trade-offs in generated patches for security, maintainability and performanceEvaluating non-functional quality of LLM-generated code beyond functional correctnessInvestigating mismatches between academic focus and industry priorities on code quality

This study addresses the limitation of existing templates in specification-driven development, which fail to evaluate specification clarity and completeness. To overcome this, we propose EPIC, a framework grounded in the ISO/IEC/IEEE 29148 standard that conducts quantitative assessments of open-source repositories. By distilling an optimal specification taxonomy encompassing ten quality dimensions and forty practices, EPIC guides developers in clarifying expectations and bridging specification gaps. Empirical evaluations demonstrate that high-quality specifications reduce the proportion of bug-fixing commits to 11.8% and yield a fourfold increase in the median number of contributors. These findings indicate that adopting rigorous specification practices significantly enhances both collaborative efficiency and software quality in open-source projects.

Coding AgentsPrompt AmbiguitySoftware Engineering

Latest Papers

What's happening recently
View more

This study addresses the accountability deficit in agent development arising from the misalignment between platform controls and service provider terms. By analyzing four categories of tools and policy documents, we map workflow responsibilities and propose a novel grid model distinguishing verification mandates from executors. This framework reveals structural deficiencies in approval mechanisms, demonstrating that responsibility gaps have evolved from human oversight to inherent product attributes. Empirical findings indicate conflicting accountabilities across layers, contradictory attribution logic, and insufficient efficacy of approval artifacts. To support further research, we release a comprehensive dataset and validation scripts as open-source resources. Collectively, this work provides both theoretical grounding and empirical evidence necessary for reconstructing accountability frameworks in agent-based software systems, highlighting the urgent need to address systemic rather than incidental failures in current governance architectures.

AccountabilityAgentic Software DevelopmentCode Review

This study addresses the limitation that analyzing prompts alone is insufficient for comprehensively evaluating developer interactions with AI programming agents. To overcome this, we propose a novel multidimensional interaction analysis framework termed "Say-Do-Understand," which integrates prompt data, screen activity, and comprehension metrics through a systematic five-stage end-to-end workflow. Employing an observational methodology, the analysis utilizes a prompt codebook, a screen activity coding scheme, and dual scoring rubrics. An empirical study involving ten experienced developers validates the proposed approach. Furthermore, four developer personas synthesizing task performance and comprehension levels are introduced to elucidate behavioral variations. Notably, the findings reveal that excessive reliance on agent self-checking significantly reduces developers' autonomous testing time.

coding agentsdeveloper understandinghuman-AI interaction

This study addresses the limited understanding of how Agent Control Files (ACFs)—instruction documents guiding autonomous coding agents—evolve, are maintained, and relate to code quality. Through large-scale repository mining, the authors reconstruct the commit-level evolutionary history of ACFs and propose, for the first time, a taxonomy of ACF changes grounded in software maintenance theory. By integrating qualitative content analysis, statistical testing, and code quality metrics, they empirically demonstrate how different types of maintenance activities differentially impact code quality and reveal dynamic patterns in these effects across the software development lifecycle. The findings provide both theoretical grounding and practical guidance for the governance of autonomous coding agents.

Agent Context Filesautonomous coding agentscode quality

This study addresses the issue that coding agents, despite passing functional tests, frequently violate repository governance standards. To this end, we introduce SWE-CC, a benchmark for systematically evaluating both code and process compliance. Methodologically, we construct 823 machine-verifiable policies, design an auditing mechanism that integrates runtime behavior with final deliverables, and propose an evaluation framework combining semi-automated document conversion, deterministic checking, and LLM agent workflows. Experimental results reveal that modern agents exhibit a policy violation rate of 43.1%, with nearly half occurring during intermediate execution steps rather than in final outputs. These findings underscore the critical necessity of process-level compliance auditing to ensure that autonomous coding agents adhere not only to functional requirements but also to established software engineering practices and repository governance norms.

benchmarkingcoding agentscontribution governance

This study addresses maintainability concerns in AI-generated code by presenting the first systematic quantitative comparison of web application structural quality across three mainstream Vibe Coding tools, including Lovable. Leveraging SonarQube static analysis to evaluate code smells, complexity, and remediation costs, this research reveals distinct structural trade-offs inherent to each platform. Specifically, Lovable exhibits a "high-density, low-severity" pattern, whereas v0 and Replit demonstrate high-severity issues accompanied by significant code redundancy. By delineating these divergent quality profiles, this work provides empirical evidence to inform tool selection and code governance strategies in AI-assisted programming environments.

AI Code GenerationCode QualityMaintainability

Hot Scholars

MR

Michael R. Lyu

Professor of Computer Science & Engineering, The Chinese University of Hong Kong
software engineeringsoftware reliabilityfault tolerancemachine learning
SG

Sin G. Teo

Institute for Infocomm Research
Applied cryptographydata privacy and securityprivacy-preserving technologiesdeep learning
HS

Houari Sahraoui

Professor of Computer Science, Université de Montréal
Software EngineeringArtificial IntelligenceAutomated software engineeringMDE
OB

Oussama Ben Sghaier

PhD Student, University of Montreal
Software EngineeringMachine learningNatural Language ProcessingDeep Learning
DT

Davide Taibi

Full Professor, University of Oulu (M3S Cloud)
Software ArchitectureCloud ContinuumMicroservicesServerless