WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement

πŸ“… 2026-07-20
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of current large language models (LLMs) in environmental enforcement, where they struggle to produce traceable and reliable decisions and lack systematic evaluation benchmarks. To bridge this gap, the authors construct the first fine-grained, multi-task evaluation framework tailored to environmental enforcement, grounded in real-world cases, regulatory standards, and expert review. The framework encompasses 14 enforcement tasks across 12 pollution media and introduces two novel metricsβ€”the Absolute Enforcement Score (AES) and the Intelligent Enforcement Index (IEI)β€”to holistically assess model performance in capability, response quality, and resource efficiency. Experimental results reveal that while LLMs perform adequately on rule-constrained tasks, they remain unreliable in critical aspects such as evidence chain construction, inconsistency detection, multi-source integration, and procedural judgment. Notably, medium-scale models already approach the performance of leading models on structured tasks, and model scaling exhibits diminishing marginal returns.
πŸ“ Abstract
Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.
Problem

Research questions and friction points this paper is trying to address.

environmental law enforcement
large language models
traceable decision-making
evidence reasoning
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

environmental law enforcement
large language models
benchmark
evidence-based reasoning
evaluation metrics
πŸ”Ž Similar Papers
Z
Ziliang Yang
School of Environment, Tsinghua University; College of Economics and Management, Beijing University of Technology
Yi Zhang
Yi Zhang
Tsinghua University
computational imagingscatteringlight-field imaging
K
Kaijun Lin
State Key Laboratory of Iron and Steel Industry Environmental Protection, School of Environment, Tsinghua University
J
Jiachao Ke
State Key Laboratory of Iron and Steel Industry Environmental Protection, School of Environment, Tsinghua University
H
Haihong Xu
Appraisal Center for Environmental Engineering, Ministry of Ecology and Environment
Z
Zongguo Wen
State Key Laboratory of Iron and Steel Industry Environmental Protection, School of Environment, Tsinghua University