Law And Order: Tax Law Autoformalization

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of accurately translating tax law texts into executable symbolic representations by proposing a neuro-symbolic framework that automatically formalizes tax forms into programs. The method establishes a dual correspondence between legal semantics and logical structures, integrating large language model-based code generation with a cell-level iterative verification and repair mechanism to ensure precise conversion. Evaluated on 51 independent test tax forms, the proposed framework achieves 100% accuracy at both the cell and form levels, significantly outperforming purely LLM-driven approaches. These results demonstrate the framework’s effectiveness in bridging natural language legal provisions and computational execution, offering a new paradigm for the reliable automated implementation of legal texts.
πŸ“ Abstract
Legal systems are increasingly implemented through software, yet scalable methods for translating legal texts into accurate symbolic representations remain underdeveloped. We study this problem through tax law, where forms and filing instructions define large computational structures involving arithmetic, branching, recursion, and tabular reasoning. We propose Law&Order, a neuro-symbolic framework for automatically formalizing tax forms and instructions into executable symbolic programs. Our approach establishes two forms of correspondence between law and logic: structural correspondence, which aligns legal and symbolic components such as cells and schedules, and denotational correspondence, which requires symbolic components to implement the computations specified by their legal counterparts. We combine large language model synthesis with cell-level verification and iterative localized error repair using human-written OpenTaxSolver tax returns. We then evaluate the resulting formalizations on independently authored, held-out TaxCalcBench returns, that are never exposed during generation or repair. Although the most advanced LLM achieves only 66% accuracy, Law&Order achieves 100% cell-level and form-level accuracy on 51 held-out returns, demonstrating the effectiveness of combining LLM-based synthesis with symbolic verification for scalable and verifiable large-scale legal autoformalization compared with using an LLM alone.
Problem

Research questions and friction points this paper is trying to address.

legal autoformalization
tax law
symbolic representation
neuro-symbolic framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neuro-symbolic framework
Autoformalization
Tax law
Large language models
Symbolic verification
S
Sophia Simeng Han
Stanford University, Stanford Law School
Y
Yoshiki Takashima
Pramaana Labs
Anjiang Wei
Anjiang Wei
Stanford University
Computer Science
Z
Zhaoyu Li
University of Toronto
M
Michael Genesereth
Stanford University, Stanford Law School