CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks struggle to evaluate large language models’ ability to discern jurisdictional differences in legal rules under identical factual scenarios and reach correct conclusions. This work introduces the first cross-jurisdictional legal reasoning benchmark grounded in authoritative statutes and case law, spanning China, California, and Germany, with 6,149 aligned instances across 55 legal issues. It proposes three complementary tasks to assess both intra- and cross-jurisdictional reasoning capabilities. The study innovatively adopts a “same facts, source-grounded” data construction paradigm and introduces the Grounded Joint metric, which jointly evaluates answer correctness and accuracy of cited legal authorities. Experimental results demonstrate that while current models can generate plausible answers, they remain notably deficient in providing accurate cross-jurisdictional legal citations.
📝 Abstract
Legal reasoning is inherently jurisdiction-dependent: the same facts can call for different legal rules and yield different conclusions across legal systems. Yet existing benchmarks rarely evaluate whether large language models (LLMs) can recognize such jurisdiction-specific variation, especially when identical fact patterns lead to divergent legal outcomes.We introduce CrossLex, a same-fact, legal-source-grounded benchmark for evaluating cross-jurisdictional legal reasoning in LLMs across three jurisdictions: China, California, and Germany. Built from authoritative legal sources, CrossLex aligns 55 legal issues spanning contract, consumer, criminal, family, and labor law, and constructs jurisdiction-aligned questions paired with answers and supporting citations. In total, CrossLex contains 6,149 instances organized into 385 fact groups, with all legal issues, answers, and cited authorities reviewed by legal professionals.To disentangle basic legal knowledge from cross-jurisdictional reasoning, CrossLex defines three complementary tasks: single-jurisdiction reasoning (T1), joint cross-jurisdictional comparison (T2), and fine-grained cross-jurisdictional evaluation (T3). We further propose Grounded Joint, a metric that jointly assesses answer correctness and legal-source grounding, and provide a unified evaluation for streamlined benchmarking. Extensive experiments on representative LLMs show that, although current models can often answer legal questions correctly, they struggle to provide accurate cross-jurisdictional legal citations.We hope that CrossLex will facilitate future research on source-grounded cross-jurisdictional legal reasoning.
Problem

Research questions and friction points this paper is trying to address.

cross-jurisdictional legal reasoning
large language models
legal benchmark
jurisdiction-dependent reasoning
legal source grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-jurisdictional legal reasoning
source-grounded benchmark
legal citation grounding
large language models
jurisdiction-aware evaluation
Xiaocui Yang
Xiaocui Yang
Lecturer, Northeastern University (China)
Multimodal Sentiment AnalysisData MiningMultimodal Large Language Models
X
Xican Tan
Institute of Intelligent Computing, University of Electronic Science and Technology of China, Chengdu, China
S
Shoujie Chen
School of Computer Science and Engineering, Northeastern University, Shenyang, China
Shihan Xiao
Shihan Xiao
Tsinghua University
Networking
K
Keke Tong
Chongqing Branch of China Unicom, Chongqing, China
X
Xinyu Zhou
Institute of Intelligent Computing, University of Electronic Science and Technology of China, Chengdu, China