π€ AI Summary
This study addresses the limitations of current large language models (LLMs) in environmental enforcement, where they struggle to produce traceable and reliable decisions and lack systematic evaluation benchmarks. To bridge this gap, the authors construct the first fine-grained, multi-task evaluation framework tailored to environmental enforcement, grounded in real-world cases, regulatory standards, and expert review. The framework encompasses 14 enforcement tasks across 12 pollution media and introduces two novel metricsβthe Absolute Enforcement Score (AES) and the Intelligent Enforcement Index (IEI)βto holistically assess model performance in capability, response quality, and resource efficiency. Experimental results reveal that while LLMs perform adequately on rule-constrained tasks, they remain unreliable in critical aspects such as evidence chain construction, inconsistency detection, multi-source integration, and procedural judgment. Notably, medium-scale models already approach the performance of leading models on structured tasks, and model scaling exhibits diminishing marginal returns.
π Abstract
Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.