From Process to Evidence: How Computing Can Ground Appropriate Reliance on Legal AI

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current legal practice exacerbates judicial inequity by overemphasizing procedural compliance while neglecting actual performance, due to a lack of empirical evidence on the error rates, failure modes, and detectability of AI tools. This study addresses this gap by systematically integrating legal duties with the human–AI interaction concept of “appropriate reliance.” Leveraging official records from New York courts, the research develops a comprehensive framework comprising legal text analysis, a task taxonomy, error metrics, benchmark maintenance protocols, and a private-data evaluation infrastructure. By distilling core requirements of judicial contexts, the work advances an agenda for evaluative infrastructure tailored to legal AI, advocating for shared error benchmarks and sustainable assessment tools that help narrow—rather than widen—the justice gap.
📝 Abstract
Lawyers and self-represented litigants are already using artificial intelligence (AI) to draft legal documents, and courts are responding with rules. After more than 1,500 cases involving AI hallucinations, lawyers have been instructed to perform careful, independent review of AI-assisted filings. Discharging these duties requires what the human-computer interaction (HCI) literature calls ``appropriate reliance,'' which cannot be calibrated without evidence on how often, how badly, and how detectably these tools fail at legal work. Existing research barely describes any of the three. We analyze the official record of the New York court system. The documents repeatedly call for evidence that does not exist (e.g., error rates, do-not-use lists). In its place they invoke procedure, including training mandates, checklists, and uncalibrated human review. The burden falls hardest on those least equipped to bear it: legal aid programs are told to track their own error rates, and judges are left to improvise their own tests. The paper makes four contributions: (1) a mapping from the legal duties to concepts in HCI; (2) a set of requirements elicited from the official record; (3) an analysis of how the legal system substitutes process for evidence; and (4) a research agenda for computing, including task taxonomies, shared error metrics, maintained benchmarks, and test harnesses for evaluations on private data. The computing community must supply what the justice system lacks; in doing so, it can help close, rather than widen, the justice gap.
Problem

Research questions and friction points this paper is trying to address.

appropriate reliance
legal AI
AI hallucinations
error rates
justice gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

appropriate reliance
legal AI evaluation
error metrics
task taxonomies
test harnesses
🔎 Similar Papers
No similar papers found.