Score
Designs and conducts structured comparisons of statutes, case law, doctrines, policies, and regulatory regimes across jurisdictions or legal systems to identify substantive and procedural differences, overlapping remedies and limitations, and doctrinal coherence. Produces synthesized assessments and actionable recommendations for harmonization, reform, interpretation, or regulatory alignment.
This work addresses the unpredictable interpretive choices often implicit in large language model (LLM) formalizations of legal provisions, which undermine the comparability and explainability of reasoning outcomes. The authors propose a systematic approach that integrates graph node matching with SAT solvers to enumerate divergent inferences arising from alternative formalizations when applied to identical legal cases. These divergences are then rendered into natural-language scenarios amenable to expert legal review. For the first time, this method maps formalization discrepancies onto intelligible edge cases, revealing their qualitative connection to real-world legal disputes. Experiments on ten EU legal provisions demonstrate that structural similarity among formalizations correlates poorly with behavioral agreement, whereas the generated divergence cases effectively capture actual conflicts in legal interpretation.
Legal compliance of machine learning models cannot be directly encoded; instead, abstract legal obligations must be “indirectly operationalized” into verifiable model design choices. Existing approaches either focus narrowly on software-level compliance or overlook legal complexity, failing to address two core challenges: the multiplicity of legal interpretations and the unpredictability of performance–compliance trade-offs. Method: We propose a five-stage interdisciplinary framework introducing the first legal–ML co-modeling paradigm, embedding legal reasoning throughout the ML development lifecycle. It features a legally adaptable operationalization mechanism and a multi-objective trade-off evaluation system. Contribution/Results: Evaluated in an anti-money laundering use case, the framework identifies an optimal configuration achieving both high detection accuracy (12% F1-score improvement) and legal defensibility, demonstrating its systematic capacity to jointly optimize predictive performance and legal legitimacy.
In administrative law case processing, manual compliance review is highly susceptible to legal dynamism and case complexity, resulting in elevated error rates, decision latency, and diminished judicial accessibility. To address these challenges, this paper proposes a model-driven automated compliance framework that innovatively integrates the domain-specific language eFLINT with Model-Driven Engineering (MDE). The framework implements an interpretable, configurable normative reasoning engine and a prototype case management system. It enables real-time modeling, dynamic parsing, and automated execution of administrative regulations, thereby enforcing normative constraints and ensuring decision transparency. Experimental evaluation demonstrates significant reductions in human bias and processing latency, alongside improved compliance consistency and enhanced citizen access to justice. The framework provides a scalable, methodology-oriented foundation for intelligent administration under digital rule of law.
This work proposes a framework to achieve interoperability between Japanese domestic legal data and international legal standards such as Akoma Ntoso, enabling semantic comparison of legal provisions across jurisdictions. We develop the first structural transformation pipeline from Japan’s Legal Standard XML (JLS) to LegalDocML (Akoma Ntoso) and integrate multilingual embeddings, FAISS-based vector retrieval, and Cross-Encoder re-ranking to enable high-precision cross-lingual alignment of legal articles. The resulting prototype system automatically generates correspondences between provisions in Japanese and European legal systems and supports comparative legal research through an interactive visualization of legal networks. This effort marks the first integration of Japan’s legal framework into the global ecosystem of semantic legal interoperability.
To address the challenge of inefficient analysis of the vast corpus of judgments from the European Court of Human Rights (ECtHR), this paper proposes the first structured legal report generation method designed for multi-case synthesis. Methodologically, it introduces an end-to-end pipeline integrating semantic retrieval to locate relevant judgment excerpts, unsupervised clustering to identify thematic clusters across cases, and a domain-finetuned large language model guided by hierarchical structural prompts to generate coherent, standardized reports. Its key contribution lies in transcending the conventional single-case summarization paradigm by enabling automatic, cross-case extraction and organization of legal issues, governing principles, and rulings—structured at three hierarchical levels. Evaluated on real ECtHR judgments, the approach achieves 92% structural compliance and 86% content accuracy. Expert legal evaluation confirms substantial improvements in analytical efficiency and scalability, establishing a novel paradigm for case-law knowledge discovery.
This study addresses the challenge of ensuring AI systems consistently adhere to human values in novel environments by drawing an analogy to judicial reasoning in legal systems. It proposes a novel bidirectional framework that bridges jurisprudence and AI alignment, integrating Dworkin’s interpretivism and Sunstein’s analogical legal reasoning with constitutional AI and case-based reasoning methods to explore the synergistic role of rules and precedents in alignment fine-tuning. The work uncovers deep structural parallels between legal interpretation and AI alignment in terms of linguistic norms and value interpretation, offering a new theoretical pathway toward building robust, scalable AI systems aligned with human objectives. Furthermore, it demonstrates how advances in AI can reciprocally inform and refine legal practice.
This work addresses the longstanding challenge of automatically translating legal texts into executable decision logic, which has traditionally relied on manual encoding and evaluation. The authors propose enhancing large language models by introducing an intermediate structured representation and present the first systematic assessment—based on real-world data from the Dutch Environmental Planning Act—of how input/output constraints and semantic role labeling influence the structural and functional equivalence of generated logic. Experimental results demonstrate that incorporating I/O constraints improves structural similarity by 37–54%, achieves functional equivalence in 51–53% of test cases, and automatically eliminates 45–55% of redundant logic nodes, thereby revealing a notable inconsistency between structural similarity and functional equivalence.
Existing benchmarks struggle to evaluate large language models’ ability to discern jurisdictional differences in legal rules under identical factual scenarios and reach correct conclusions. This work introduces the first cross-jurisdictional legal reasoning benchmark grounded in authoritative statutes and case law, spanning China, California, and Germany, with 6,149 aligned instances across 55 legal issues. It proposes three complementary tasks to assess both intra- and cross-jurisdictional reasoning capabilities. The study innovatively adopts a “same facts, source-grounded” data construction paradigm and introduces the Grounded Joint metric, which jointly evaluates answer correctness and accuracy of cited legal authorities. Experimental results demonstrate that while current models can generate plausible answers, they remain notably deficient in providing accurate cross-jurisdictional legal citations.
Current evaluations of legal AI systems are largely confined to auxiliary tasks and fail to adequately assess doctrinal legal reasoning, leaving the European Union’s Artificial Intelligence Act’s requirement of “appropriate accuracy” without operationalizable standards. This study introduces the first benchmark specifically designed for evaluating doctrinal legal reasoning, translating regulatory compliance demands into technically measurable indicators. By integrating legal doctrinal analysis with large language model evaluation methodologies, the work formulates specialized tasks to assess capabilities in legal interpretation and reasoning. The proposed framework addresses a critical gap in evaluating high-risk judicial AI systems at the level of professional legal reasoning, thereby providing essential support for advancing legal AI beyond mere text generation toward normatively compliant, expert-level reasoning.