Score
Designs and performs mappings and assessments that relate system behaviors, artifacts, and records to legal, regulatory, safety, security, and licensing requirements; this includes building compliance mappings, gap analyses, mitigation plans, and audit evidence criteria. Translates legal norms into technical specifications and traceability requirements, and validates that runtime logs, records, and system traces satisfy evidentiary and regulatory obligations.
This work addresses the error-prone and labor-intensive process of manually translating regulatory texts such as the GDPR and the EU AI Act into actionable software requirements. The authors propose Reg2Req, the first end-to-end automated pipeline that leverages natural language processing to identify regulatory provisions, generate system-agnostic software requirements accompanied by plain-language explanations, and establish traceability links. The approach supports requirement classification, use case seed generation, and cross-reference analysis, achieving macro-averaged F1 scores of 0.82 on the GDPR and 0.78 on the EU AI Act. A user study demonstrates that the generated plain-language explanations significantly enhance users’ comprehension and confidence in taking compliance actions (p < 0.001), with all participants expressing willingness to adopt the output as a starting point for compliance efforts.
This paper addresses the high manual effort and poor generalizability in establishing traceability between software requirements and legal regulations (e.g., GDPR). We propose an automated legal traceability method and conduct the first systematic comparison of classification-based (Kashif, built on fine-tuned Sentence-BERT) and generation-based (Rice, a prompt-engineering-driven large language model) approaches on legal traceability tasks. Results demonstrate that the generation-based approach substantially outperforms the fine-tuned classifier: Kashif achieves 67% recall on the benchmark dataset (a 54-percentage-point improvement over baseline), whereas Rice attains 84% recall on real-world GDPR documents—69 percentage points higher than Kashif. Our core contribution lies in empirically validating that prompt-engineering-driven generative paradigms exhibit superior zero-shot generalization capability and greater practical deployability for legal traceability tasks.
This study addresses the challenges posed by the proliferation, complexity, and expanding scope of regulatory requirements in software engineering, which hinder their systematic integration into development processes. To tackle this issue, the paper proposes a viewpoint-centered, artifact-based approach to regulatory requirements engineering. The approach innovatively integrates viewpoint analysis with artifact modeling to develop the AM4RRE (Artifact Modeling for Regulatory Requirements Engineering) framework, which facilitates cross-functional collaboration and ensures consistency in compliance-driven design. Preliminary validation demonstrates that AM4RRE effectively bridges the gap between organizational regulatory processes and software development practices, enabling a shift from ad hoc compliance responses toward systematic integration. This foundational work paves the way for further empirical investigation into scalable and sustainable regulatory compliance in software engineering.
Normative requirements—encompassing Social, Legal, Ethical, Empathic, and Cultural (SLEEC) dimensions—are notoriously difficult to comprehend, debug, and verify in multi-stakeholder collaborative settings due to their inherent ambiguity and non-technical nature. Method: This paper introduces SLEEC-LLM, the first framework to leverage large language models (LLMs) for generating natural-language explanations of counterexamples revealing SLEEC requirement inconsistencies—thereby bridging the cognitive gap between formal verification outputs and non-technical stakeholders. It integrates a domain-specific language (DSL), model checking, and LLM-based explanation generation to produce human-readable, semantically precise interpretations. Results: Evaluated on two real-world case studies, SLEEC-LLM significantly improves non-technical stakeholders’ comprehension speed (62% reduction in time-to-understanding) and conflict identification accuracy (+38%). It markedly reduces cognitive load during requirement iteration and advances explainable, collaborative requirements engineering.
To address challenges in legal compliance checking—including high subjectivity in regulatory interpretation, dynamic evolution of legislation, and difficulties in cross-disciplinary collaboration—this paper introduces eFLINT, a domain-specific language for computable modeling and automated verification of legal rules, regulatory requirements, and contractual clauses. eFLINT integrates declarative and procedural paradigms, explicitly linking legal concepts to executable computational logic. It combines formal specification, context-aware reasoning, and scenario-based modeling to enable dynamic, end-to-end compliance verification across system design, runtime, and post-execution phases. Designed to balance expressiveness and executability, eFLINT reconciles conflicting requirements through principled language design. Drawing on multi-scenario industrial deployments, the paper distills actionable design principles and a methodology for automation-oriented compliance languages. It contributes both a reusable technical framework and theoretical foundations for computable regulation research in legal technology.
Current AI systems rely heavily on manual auditing and documentation, which hinders scalable governance for automated services. This work proposes Ontological Knowledge Blocks (OKBs), a novel framework that formalizes regulatory obligations as quintuples comprising ontologies, SHACL rules, evidence requirements, and provenance links. By leveraging RDF/OWL modeling, PROV-O for provenance tracking, and an intermediate representation–driven deterministic compiler, the approach enables dynamic switching of governance configurations without modifying service code. Evaluation in an AI-assisted HPC scheduling scenario demonstrates that compliance checks are configuration-sensitive, violations accumulate strictly additively, SHACL validation incurs only 12.6–100.3 milliseconds of latency, and the Combined configuration provides the most comprehensive coverage.
This study addresses the high complexity and labor-intensive challenges of accurately translating privacy regulations such as Brazil’s General Data Protection Law (LGPD) into actionable software requirements. It presents the first systematic exploration of leveraging large language models (LLMs) for generating LGPD-compliant requirements, proposing an automated approach that integrates legal text analysis with requirements engineering to directly map statutory provisions into user stories and acceptance test scenarios. Experimental results demonstrate that the method efficiently produces high-quality, executable compliance requirements, significantly supporting regulatory adherence during early-stage software development. This work thus offers an innovative and practical technical pathway for privacy regulation–driven requirements engineering.
This study addresses the compliance challenges faced by data practitioners in machine learning systems under regulations such as the GDPR and the AI Act, particularly concerning data quality. Through semi-structured interviews with practitioners in the European Union, combined with thematic analysis of regulatory texts and engineering workflows, the research systematically uncovers a structural disconnect between regulation-driven data quality requirements and ML engineering practices. It identifies five core challenges: misalignment between legal principles and engineering implementation, fragmented data pipelines, lack of purpose-built compliance tools, ambiguous accountability, and reactive responses to audits. Building on these findings, the work proposes directions for designing compliance-oriented tooling, establishing effective governance mechanisms, and fostering cultural transformation to bridge the gap between regulatory mandates and practical ML development.
This work addresses key challenges enterprises face when transitioning large language model (LLM) prototypes into production—namely, insufficient auditability, unpredictable behavior, and the absence of enforceable behavioral guarantees. The authors propose “harness engineering,” a methodology that restructures prompt-driven prototypes into auditable, traceable LLM agent architectures. For the first time, enterprise-grade behavioral contracts—including entity routing, source attribution, and output formatting—are formally encoded as executable code, establishing model-agnostic, verifiable safety boundaries. Through codified contracts, runtime validators, fault-injection testing, and model-swapping evaluations on data from 25 Korean publicly listed companies, the approach demonstrates 100% compliance with specified contracts across three hosted LLMs while preserving full functional utility (120/120), significantly outperforming pure prompting or external guardrail strategies.
This work addresses the lack of traceable and tamper-resistant transparency mechanisms in large language models (LLMs) deployed in high-stakes decision-making contexts, which undermines accountability. To bridge this gap, the paper introduces the first LLM lifecycle auditing framework that integrates technical provenance with governance records. It proposes a reference architecture enabling cross-organizational traceability and implements a lightweight, open-source Python-based auditing layer. By leveraging append-only logs, event emitters, structured metadata, and an auditor interface, the system seamlessly integrates into existing LLM workflows with minimal intrusiveness. This design ensures complete, tamper-evident traceability across critical stages—including training, deployment, and monitoring—thereby facilitating robust accountability and responsibility attribution throughout the model’s lifecycle.