🤖 AI Summary
This study addresses the lack of real-world consultation scenarios in existing Vietnamese legal benchmarks, which hinders the effective evaluation of trustworthy legal AI. To this end, we construct a large-scale legal benchmark derived from authentic citizen-lawyer interactions, comprising 172,000 questions across 34 legal domains. The dataset provides expert-verified legal evidence and professional answers to support retrieval as well as extractive and generative question-answering tasks. We evaluate hybrid retrieval techniques alongside pretrained language models through comprehensive experiments. Results demonstrate that hybrid retrieval strategies achieve optimal performance while revealing fundamental challenges in mapping natural language queries to statutory provisions. This work bridges the evaluation gap for Vietnamese legal AI and establishes a rigorous, high-difficulty benchmark for future research.
📝 Abstract
Trustworthy Legal AI requires systems that can answer legal questions while grounding their responses in authoritative sources. However, existing Vietnamese legal benchmarks provide limited coverage of real-world legal consultations. We introduce \textbf{ViLegalExpert}, a large-scale benchmark constructed from authentic citizen--lawyer consultations, containing over \textbf{172K} questions across \textbf{34 legal domains}, together with professional answers and expert-verified legal evidence. ViLegalExpert supports legal information retrieval, extractive QA, and abstractive QA. Experiments with representative retrieval methods and language models reveal substantial challenges in evidence retrieval and grounded answer generation. While pretrained models perform strongly on QA, hybrid retrieval achieves the best retrieval performance. These results demonstrate the difficulty of mapping naturally expressed legal questions to authoritative provisions and establish ViLegalExpert as a challenging benchmark for reliable Vietnamese Legal AI.