🤖 AI Summary
This study presents the first systematic Turing test of off-the-shelf large language models (LLMs) on authentic legal professional examinations—specifically those for attorneys, judges, and notaries—without any task-specific fine-tuning, to delineate the boundaries of their legal reasoning and practical competence. Employing a blind evaluation protocol, expert graders assessed complete model-generated responses against official scoring rubrics under anonymized conditions. Results reveal that certain LLMs match or even surpass top human candidates in adversarial argumentation and doctrinal analysis, yet consistently fail in the notary examination, which imposes stringent formal and substantive constraints. These findings demonstrate that LLMs’ legal capabilities are highly task-dependent and provide an empirical foundation and methodological framework for evaluating the applicability of general-purpose LLMs in high-stakes professional domains.
📝 Abstract
The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and recurring legal failure patterns. Although limited to out-of-the-box systems, the findings provide qualitative evidence on the current scope and boundaries of the legal competence of LLMs across distinct professional tasks.