Score
Analyze statutes, regulations, judicial opinions, contracts, and rights- and privacy-related doctrines to map factual scenarios to applicable legal rules, assess obligations, liabilities, and individual or organizational rights, and identify compliance risks. Translate and formalize legal language and doctrines into structured representations or arguments for use in compliance, risk assessment, or legal decision-making.
Legal compliance of machine learning models cannot be directly encoded; instead, abstract legal obligations must be “indirectly operationalized” into verifiable model design choices. Existing approaches either focus narrowly on software-level compliance or overlook legal complexity, failing to address two core challenges: the multiplicity of legal interpretations and the unpredictability of performance–compliance trade-offs. Method: We propose a five-stage interdisciplinary framework introducing the first legal–ML co-modeling paradigm, embedding legal reasoning throughout the ML development lifecycle. It features a legally adaptable operationalization mechanism and a multi-objective trade-off evaluation system. Contribution/Results: Evaluated in an anti-money laundering use case, the framework identifies an optimal configuration achieving both high detection accuracy (12% F1-score improvement) and legal defensibility, demonstrating its systematic capacity to jointly optimize predictive performance and legal legitimacy.
This study addresses three critical challenges in applying large language models (LLMs) to legal text interpretation: hallucination, algorithmic homogeneity, and cross-jurisdictional regulatory compliance. Methodologically, we propose a systematic optimization framework tailored for the legal domain. We introduce, for the first time, a dual-benchmark evaluation system jointly assessing algorithmic robustness and regulatory compliance—covering the EU AI Act, emerging U.S. regulations, and China’s AI governance framework. The framework integrates semantic understanding, instruction fine-tuning, multi-source legal knowledge alignment, compliance-constrained decoding, and hallucination suppression mechanisms. Experimental results demonstrate state-of-the-art performance on key tasks—including contractual clause identification and case-based analogical reasoning—while significantly improving accuracy, interpretability, and cross-jurisdictional adaptability in legal summarization, contract negotiation support, and legal information retrieval.
Current large language models often rely on extratextual assumptions in legal reasoning, leading to logically unfaithful and unverifiable conclusions that fail to meet the legal profession’s stringent demands for rigor and accountability. This work proposes a neuro-symbolic framework that integrates the expressive power of large language models with formal logical verification to ensure that all inferences remain strictly grounded in the original legal text, thereby eliminating unwarranted assumptions. The resulting approach yields a traceable and verifiable legal reasoning mechanism that substantially reduces hypothetical errors, alleviates the burden of manual review, and enhances system trustworthiness while preserving logical soundness and accountability.
Large language models (LLMs) exhibit limited performance in Malaysian contract law IRAC (Issue, Rule, Application, Conclusion) reasoning due to insufficient domain-specific legal terminology mastery and shallow legal knowledge grounding. Method: We introduce LEGALSEMI—the first semi-structured benchmark explicitly designed for IRAC-based legal reasoning—comprising 54 expert-annotated scenarios and a complementary Structured Knowledge Graph (SKG). Our methodology systematically integrates the IRAC framework into dataset design, proposing a novel “semi-structured scenario annotation + SKG co-enhancement” paradigm. Contribution/Results: Integrating SKG into Llama-3, Qwen, GPT-4, and Claude-3 yields an average 23.7% improvement in F1 scores across all four IRAC subtasks. This demonstrates the efficacy of a data–knowledge dual-driven approach for modeling multi-step legal reasoning, establishing a foundational resource and methodology for domain-adapted legal LLM evaluation and enhancement.
This survey addresses key challenges in legal NLP: strong long-range dependencies, domain-specific linguistic complexity, severe scarcity of annotated data, and inadequate model interpretability. Following the PRISMA framework, the authors rigorously select and analyze 127 studies to construct the first comprehensive taxonomy spanning legal summarization, named entity recognition, question answering, argument mining, classification, and judgment prediction. They identify 15 open challenges—centering on mitigating AI bias, enhancing robustness of legal reasoning, and advancing trustworthy modeling. A systematic evaluation is conducted across domain-specific models (e.g., Legal-BERT), adaptation techniques (fine-tuning, prompt engineering, domain-adaptive pretraining), and task-specific metrics, yielding a task-oriented benchmarking framework. The work establishes a foundational reference for developing compliant, interpretable, and production-ready legal AI systems, charting a clear evolutionary pathway for future research and deployment.
This work addresses the critical challenge of efficiently and accurately translating natural language legal provisions into machine-executable logic programs for intelligent regulatory compliance. The authors propose an end-to-end framework that, for the first time, leverages a single large language model (LLM) prompt to simultaneously generate human-readable if-then rules and their corresponding formal encodings in PROLEG. The approach incorporates a closed-loop refinement mechanism guided by legal expert feedback to enhance accuracy and fidelity. Demonstrated on Article 6 of the GDPR, the method successfully produces executable logic programs that support precise legal reasoning while also generating interpretable, human-readable explanations of regulatory decisions. This study validates the feasibility and practical utility of an automated pipeline for transforming unstructured legal texts into executable logical representations.
AI systems rely on natural language rules to align with human intent, yet the inherent ambiguity in rule interpretation leads to inconsistent behavior and lacks institutionalized mechanisms—akin to legal systems—for constraining interpretive divergence. This paper pioneers the integration of legal interpretation theory into AI alignment, proposing a dual computational framework: “rule refinement” iteratively reduces syntactic and semantic ambiguity in rule formulations, while “interpretive constraint” employs prompt engineering and consistency-aware modeling to govern rule execution. Evaluated on 5,000 multi-scenario judgment tasks from the WildChat dataset, the approach significantly improves inter-annotator agreement among plausible interpreters (p < 0.01) and enhances model robustness in following complex linguistic instructions. This work systematically addresses interpretive ambiguity in natural language rules for AI—a longstanding challenge—and establishes a novel paradigm for building trustworthy, legally informed AI systems.
Existing legal reasoning research predominantly relies on generic frameworks (e.g., syllogism, IRAC), overlooking fine-grained reasoning processes in civil cases—particularly Chinese tort disputes—and lacks domain-specific evaluation benchmarks. To address this, we propose LawChain, the first structured legal reasoning framework tailored to Chinese civil tort cases, comprising three modules: attribution analysis, liability determination, and damage quantification. Concurrently, we introduce LawChain$_{eval}$, a dedicated benchmark for rigorous evaluation. Innovatively, LawChain integrates downstream tasks—including legal named entity recognition and compensation calculation—to assess generalization capability. Experimental results reveal that mainstream large language models exhibit significant weaknesses in critical reasoning steps. Incorporating LawChain yields substantial multi-task performance gains over baselines, establishing a new paradigm for fine-grained modeling and evaluation in civil legal AI.
This work addresses the longstanding challenge of automatically translating legal texts into executable decision logic, which has traditionally relied on manual encoding and evaluation. The authors propose enhancing large language models by introducing an intermediate structured representation and present the first systematic assessment—based on real-world data from the Dutch Environmental Planning Act—of how input/output constraints and semantic role labeling influence the structural and functional equivalence of generated logic. Experimental results demonstrate that incorporating I/O constraints improves structural similarity by 37–54%, achieves functional equivalence in 51–53% of test cases, and automatically eliminates 45–55% of redundant logic nodes, thereby revealing a notable inconsistency between structural similarity and functional equivalence.
Criminal judgment documents lack structured annotations of judicial reasoning elements, hindering large-scale empirical legal research. Method: This paper proposes the first structured annotation schema for three core judicial reasoning components in criminal judgments—holding, evidentiary consideration, and subsumption—and constructs the first bilingual judicial reasoning annotation dataset. Leveraging this dataset, we conduct the first investigation into few-shot automatic identification of these reasoning components using a large language model (ChatGLM2), followed by fine-tuning a multiclass classifier for end-to-end extraction. Contribution/Results: Experimental results show an 80% accuracy, demonstrating the feasibility of computationally modeling judicial reasoning logic. This work provides a reproducible technical pipeline and foundational resources for legal AI, advancing quantitative analysis and theoretical modeling of judicial decision-making logic.