Score
Designs and produces formal or structured specifications and interpretations of legal texts—including statutes, regulations, contracts, and case law—by identifying and articulating obligations, permissions, prohibitions, and exceptions. Analyzes and translates those legal requirements into explicit decision rules, compliance criteria, testable requirements, or machine-readable representations for use in policies, systems, or processes.
This survey addresses key challenges in legal NLP: strong long-range dependencies, domain-specific linguistic complexity, severe scarcity of annotated data, and inadequate model interpretability. Following the PRISMA framework, the authors rigorously select and analyze 127 studies to construct the first comprehensive taxonomy spanning legal summarization, named entity recognition, question answering, argument mining, classification, and judgment prediction. They identify 15 open challenges—centering on mitigating AI bias, enhancing robustness of legal reasoning, and advancing trustworthy modeling. A systematic evaluation is conducted across domain-specific models (e.g., Legal-BERT), adaptation techniques (fine-tuning, prompt engineering, domain-adaptive pretraining), and task-specific metrics, yielding a task-oriented benchmarking framework. The work establishes a foundational reference for developing compliant, interpretable, and production-ready legal AI systems, charting a clear evolutionary pathway for future research and deployment.
This paper addresses the fragmented state of research and inconsistent evaluation practices in legal artificial intelligence (Legal AI) powered by large language models (LLMs). To tackle these challenges, we propose the first comprehensive technical landscape: systematically cataloging 16 LLM families, 47 task-specific frameworks, 15 benchmark suites, and 29 domain-specific datasets; introducing a multidimensional evaluation framework covering legal understanding, reasoning, and generation; and open-sourcing an integrated resource platform. Our analysis identifies persistent limitations across current models—including domain expertise, interpretability, and regulatory compliance—while offering Legal AI researchers a structured entry point and reusable infrastructure. The work significantly enhances the systematicity, reproducibility, and cross-study comparability of Legal AI research.
This study addresses three critical challenges in applying large language models (LLMs) to legal text interpretation: hallucination, algorithmic homogeneity, and cross-jurisdictional regulatory compliance. Methodologically, we propose a systematic optimization framework tailored for the legal domain. We introduce, for the first time, a dual-benchmark evaluation system jointly assessing algorithmic robustness and regulatory compliance—covering the EU AI Act, emerging U.S. regulations, and China’s AI governance framework. The framework integrates semantic understanding, instruction fine-tuning, multi-source legal knowledge alignment, compliance-constrained decoding, and hallucination suppression mechanisms. Experimental results demonstrate state-of-the-art performance on key tasks—including contractual clause identification and case-based analogical reasoning—while significantly improving accuracy, interpretability, and cross-jurisdictional adaptability in legal summarization, contract negotiation support, and legal information retrieval.
Small organizations and startups often lack access to legal expertise, hindering their ability to interpret and comply with regulatory requirements. Method: This work proposes a lightweight, low-supervision legal text structuring approach that integrates textual entailment recognition, in-context learning, and domain-specific meta-model-driven Python class generation to automatically extract semantic legal metadata and explicitly encode logical relationships—yielding verifiable, executable compliance representations. Contribution/Results: The method avoids reliance on large-scale annotated datasets, thereby enhancing cross-regulation generalizability; crucially, it directly maps legal provisions to runnable code. Evaluated on data breach notification laws across 13 U.S. states, the generated representations pass 89.4% of test cases, achieving 82.2% precision and 88.7% recall.
Natural language often introduces execution ambiguity in computational legal applications, whereas formal languages risk undermining legal legitimacy and public accessibility. This tension raises fundamental trade-offs among normative clarity, public comprehensibility, and algorithmic executability. Method: Drawing on an EU road transport regulation case study, the paper conducts a comparative analysis of natural-language legal texts and formal computational models, integrating jurisprudential reasoning, normative linguistics, and computational logic evaluation. Contribution/Results: The study systematically identifies the inherent tensions between natural and formal languages in representing core legal principles—particularly interpretability, traceability, and intelligibility—and proposes design principles for hybrid normative frameworks that simultaneously preserve legal integrity and ensure machine operability. It is the first work to rigorously characterize the tripartite trade-off across normative precision, democratic accessibility, and computational enforceability, offering a foundational methodology for legally sound legal informatics.
Legal compliance of machine learning models cannot be directly encoded; instead, abstract legal obligations must be “indirectly operationalized” into verifiable model design choices. Existing approaches either focus narrowly on software-level compliance or overlook legal complexity, failing to address two core challenges: the multiplicity of legal interpretations and the unpredictability of performance–compliance trade-offs. Method: We propose a five-stage interdisciplinary framework introducing the first legal–ML co-modeling paradigm, embedding legal reasoning throughout the ML development lifecycle. It features a legally adaptable operationalization mechanism and a multi-objective trade-off evaluation system. Contribution/Results: Evaluated in an anti-money laundering use case, the framework identifies an optimal configuration achieving both high detection accuracy (12% F1-score improvement) and legal defensibility, demonstrating its systematic capacity to jointly optimize predictive performance and legal legitimacy.
This work addresses the error-prone and labor-intensive process of manually translating regulatory texts such as the GDPR and the EU AI Act into actionable software requirements. The authors propose Reg2Req, the first end-to-end automated pipeline that leverages natural language processing to identify regulatory provisions, generate system-agnostic software requirements accompanied by plain-language explanations, and establish traceability links. The approach supports requirement classification, use case seed generation, and cross-reference analysis, achieving macro-averaged F1 scores of 0.82 on the GDPR and 0.78 on the EU AI Act. A user study demonstrates that the generated plain-language explanations significantly enhance users’ comprehension and confidence in taking compliance actions (p < 0.001), with all participants expressing willingness to adopt the output as a starting point for compliance efforts.
To address challenges in legal compliance checking—including high subjectivity in regulatory interpretation, dynamic evolution of legislation, and difficulties in cross-disciplinary collaboration—this paper introduces eFLINT, a domain-specific language for computable modeling and automated verification of legal rules, regulatory requirements, and contractual clauses. eFLINT integrates declarative and procedural paradigms, explicitly linking legal concepts to executable computational logic. It combines formal specification, context-aware reasoning, and scenario-based modeling to enable dynamic, end-to-end compliance verification across system design, runtime, and post-execution phases. Designed to balance expressiveness and executability, eFLINT reconciles conflicting requirements through principled language design. Drawing on multi-scenario industrial deployments, the paper distills actionable design principles and a methodology for automation-oriented compliance languages. It contributes both a reusable technical framework and theoretical foundations for computable regulation research in legal technology.
The technological neutrality of legal texts impedes software engineers from efficiently and accurately translating them into executable compliance requirements; current manual translation processes are time-consuming, error-prone, and heavily reliant on domain experts. Method: This paper proposes an automated approach leveraging large language models (LLMs)—specifically Claude and Llama—to precisely map food safety regulations into Gherkin-formatted behavioral specifications, enabling Behavior-Driven Development (BDD)-driven compliance engineering and testing. Contribution/Results: Through a human-centered quasi-experiment—the first systematic evaluation of LLM-generated behavioral specifications—the study demonstrates high scores across relevance, clarity, and completeness. Empirical results show significant reduction in requirement translation time and broad developer acceptance of the method’s engineering utility. The core contribution is an end-to-end generative framework that bridges regulatory texts to testable behavioral specifications, empirically validated for effectiveness and feasibility.
Current large language models often rely on extratextual assumptions in legal reasoning, leading to logically unfaithful and unverifiable conclusions that fail to meet the legal profession’s stringent demands for rigor and accountability. This work proposes a neuro-symbolic framework that integrates the expressive power of large language models with formal logical verification to ensure that all inferences remain strictly grounded in the original legal text, thereby eliminating unwarranted assumptions. The resulting approach yields a traceable and verifiable legal reasoning mechanism that substantially reduces hypothetical errors, alleviates the burden of manual review, and enhances system trustworthiness while preserving logical soundness and accountability.
This work addresses the critical challenge of efficiently and accurately translating natural language legal provisions into machine-executable logic programs for intelligent regulatory compliance. The authors propose an end-to-end framework that, for the first time, leverages a single large language model (LLM) prompt to simultaneously generate human-readable if-then rules and their corresponding formal encodings in PROLEG. The approach incorporates a closed-loop refinement mechanism guided by legal expert feedback to enhance accuracy and fidelity. Demonstrated on Article 6 of the GDPR, the method successfully produces executable logic programs that support precise legal reasoning while also generating interpretable, human-readable explanations of regulatory decisions. This study validates the feasibility and practical utility of an automated pipeline for transforming unstructured legal texts into executable logical representations.
This work addresses the unpredictable interpretive choices often implicit in large language model (LLM) formalizations of legal provisions, which undermine the comparability and explainability of reasoning outcomes. The authors propose a systematic approach that integrates graph node matching with SAT solvers to enumerate divergent inferences arising from alternative formalizations when applied to identical legal cases. These divergences are then rendered into natural-language scenarios amenable to expert legal review. For the first time, this method maps formalization discrepancies onto intelligible edge cases, revealing their qualitative connection to real-world legal disputes. Experiments on ten EU legal provisions demonstrate that structural similarity among formalizations correlates poorly with behavioral agreement, whereas the generated divergence cases effectively capture actual conflicts in legal interpretation.