JusticeAxis: Benchmarking Legal Judgment between Rigid Rule Application and Ungrounded Discretion

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the imbalance between rigid statutory matching and unfounded judicial discretion in legal judgment by constructing a multimodal benchmark and proposing an agent-based framework. Methodologically, it formalizes reference-anchored facts to define the task and introduces Bayesian credible bounds to filter empirical skills, thereby enabling skill-driven adjudication. Furthermore, it reveals a failure-direction drift phenomenon induced by model scaling. Experiments demonstrate that this framework effectively characterizes the bias profiles of diverse models, while its plug-and-play architecture elevates open-source models to commercial-grade adjudicatory performance, offering a reliable technical pathway toward balanced legal decision-making.
📝 Abstract
A sound judgment applies the law to established facts and weighs the circumstances in which they arose. However, existing methods swing between rigid statute matching and ungrounded discretion, benchmarks score a label or a rubric, and the experience that would supply the balance stays unverified. We formalize legal judgment as a reference-anchored task, whose object is a single decision that stays tied to the statute and to the circumstances at once. We introduce JusticeAxis, 256 real-world criminal cases from 18 countries with audio, image, and text evidence, and three lawyer-written judgments for every case: the recorded one and one for each failure. We further propose JusticeAgent, a harness whose element agents establish the facts and whose judge agent applies the law under skills carrying experience of the circumstances. Skills are distilled from execution trajectories and admitted only under Bayesian credible bounds. Experiments show that failure turns direction with scale: open-weight backbones drift to unsupported grounds, frontier models to the statutory default. We further verify that JusticeAgent, as a simple yet effective plugin, carries a frozen open-weight backbone to commercial level. Project resources are available at https://github.com/beita6969/JusticeAxis.
Problem

Research questions and friction points this paper is trying to address.

Legal Judgment
Benchmarking
Rigid Rule Application
Ungrounded Discretion
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Legal Judgment Benchmarking
Multi-Agent Framework
Skill Distillation
Bayesian Credible Bounds
Multimodal Evidence
🔎 Similar Papers
No similar papers found.
Zhengkai Tu
Zhengkai Tu
The Chinese University of Hong Kong
Mingda Zhang
Mingda Zhang
Google DeepMind
Computer VisionMulti-modal UnderstandingVideo Generation
Z
Zijia Wang
University of Oxford
X
Xiaoying Tang
The Chinese University of Hong Kong, Shenzhen
J
Jimmy Huang
York University