🤖 AI Summary
This study addresses the challenge of legal understanding and reasoning for driving regulations in low-resource languages, specifically Romanian, by introducing Romanian-DrivingQA—the first multimodal question-answering benchmark tailored to the Romanian driving license examination. It comprises text- and image-based questions, annotated with both legal provisions and human-written explanatory rationales to support interpretable reasoning. Methodologically, the framework integrates retrieval-augmented generation (RAG), dense retrieval, chain-of-thought prompting, and domain-specific reasoning models, enabling systematic evaluation across text QA, visual QA, and cross-modal reasoning. Experiments demonstrate that domain-adaptive fine-tuning substantially improves retrieval quality; optimized reasoning strategies achieve text QA accuracy exceeding the passing threshold (90%), whereas visual reasoning remains a critical bottleneck. Key contributions include: (1) the first multilingual, multimodal driving examination benchmark; (2) a dual-annotation schema linking questions to legal sources and pedagogical explanations; and (3) empirical validation of RAG and CoT efficacy in legal education contexts.
📝 Abstract
The intersection of AI and legal systems presents a growing need for tools that support legal education, particularly in under-resourced languages such as Romanian. In this work, we aim to evaluate the capabilities of Large Language Models (LLMs) and Vision-Language Models (VLMs) in understanding and reasoning about Romanian driving law through textual and visual question-answering tasks. To facilitate this, we introduce RoD-TAL, a novel multimodal dataset comprising Romanian driving test questions, text-based and image-based, alongside annotated legal references and human explanations. We implement and assess retrieval-augmented generation (RAG) pipelines, dense retrievers, and reasoning-optimized models across tasks including Information Retrieval (IR), Question Answering (QA), Visual IR, and Visual QA. Our experiments demonstrate that domain-specific fine-tuning significantly enhances retrieval performance. At the same time, chain-of-thought prompting and specialized reasoning models improve QA accuracy, surpassing the minimum grades required to pass driving exams. However, visual reasoning remains challenging, highlighting the potential and the limitations of applying LLMs and VLMs to legal education.