LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credibility of legal citations in lengthy Chinese legal research reports by proposing the first multidimensional evaluation framework tailored to Chinese legal texts. The authors introduce the LegalCiteTrust benchmark, which operationalizes citation assessment along three dimensions—coverage, support, and trustworthiness—through an Existence/Faithfulness/Applicability (E/F/A) framework. Combining dense human annotations with multi-system comparative experiments, the research reveals that while reliance on retrieval tools enhances evidential support, it does not necessarily improve citation trustworthiness. In contrast, revision strategies grounded in the E/F/A criteria significantly outperform approaches that merely filter out non-existent citations, leading to marked improvements in overall report quality and underscoring the critical role of post-citation governance.
📝 Abstract
Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, but also on whether cited legal authorities are trustworthy. A citation can be risky even when it points to a real source: the report may omit limiting conditions, misdescribe the authority, or use it to support a stronger claim than the source allows. We introduce LegalCiteTrust, a benchmark for evaluating citation trustworthiness in Chinese long-form legal research reports. It contains 72 densely annotated report-level tasks and evaluates reports along three dimensions: Coverage, Support, and Citation Trustworthiness. Citation Trustworthiness is operationalized through citation-level Existence, Fidelity, and Applicability (E/F/A). Experiments on general-purpose LLMs, deep-research systems, and legal-specific systems show that task completion, evidence richness, citation density, and citation reliability expose different system behaviors. Retrieval tools can improve evidence support without reliably improving the Trust score, while E/F/A-based revision improves Trust and Final score more clearly than existence-only filtering. These results suggest that trustworthy legal research generation requires citation-aware evidence governance after retrieval: systems must not only retrieve legal authorities, but also select, describe, and apply them reliably.
Problem

Research questions and friction points this paper is trying to address.

citation trustworthiness
legal research reports
LLMs
evidence governance
Chinese legal domain
Innovation

Methods, ideas, or system contributions that make the work stand out.

Citation Trustworthiness
Legal Research Benchmark
Evidence Governance
E/F/A Framework
Long-form Legal Generation