Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过对比Jev与九种语言模型在ContractNLI上的表现,评估了合同推理的成本、响应时间及准确性,揭示了配置变化对个别判断正确性的影响。
📝 Abstract
Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating inference cost, response time, average correctness, and correctness across repeated request conditions. Controlled comparisons vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev has the lowest cost and median response time among the evaluated configurations, while hosted language models achieve higher baseline accuracy. Rankings by baseline accuracy differ from rankings by correctness across every condition and repeat, although small differences in the latter do not establish a general stability advantage. Development diagnostics further reveal compensating corrections and regressions, as well as persistent errors. These findings motivate evaluating cost and response time alongside whether individual judgments remain correct as the request configuration changes. Code: https://github.com/ZF-Utokyo/Jev-Benchmark
Problem

Research questions and friction points this paper is trying to address.

Contract Inference
Individual Decisions
Repetition Consistency
Request Configuration
Innovation

Methods, ideas, or system contributions that make the work stand out.

ContractNLI
JEV
response time
cost evaluation
stability
💼 Related Jobs
No related jobs found.
F
Fan Zhang
The University of Tokyo; MBZUAI
Yankai Chen
Yankai Chen
Postdoctoral Associate, Cornell University
Information RetrievalKnowledge MiningLarge Language ModelsAgentic AI
Zhuohan Xie
Zhuohan Xie
MBZUAI
Financial AIReasoningNatural Language ProcessingComputational LinguisticsDeep Learning
Y
Yixi Zhou
Hong Kong Baptist University
S
Sijia Peng
Fudan University
L
Lei Fan
University of Illinois Urbana-Champaign
X
Xinhua Ji
UCloud
C
Cunyuan Zheng
Columbia University
H
Huangyong Shan
The University of Hong Kong; Quantell Capital
Philip S. Yu
Philip S. Yu
Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
X
Xue Liu
MBZUAI; McGill University
Y
Yu Chen
The University of Tokyo
Preslav Nakov
Preslav Nakov
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Computational LinguisticsLarge Language ModelsFact-checkingFake News
S
Songwei He
The University of Hong Kong; Quantell Capital