Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between statutory citation and judicial decision-making in large language models, alongside the lack of faithfulness in their legal explanations. To investigate this, we propose a counterfactual auditing framework based on hidden-state decoding, integrating LoRA fine-tuning with red-teaming evaluations to systematically assess the causal faithfulness of legal chain-of-thought reasoning. Experiments across seven models and four benchmarks reveal that although citation accuracy ranges from 66.7% to 100%, judgment sensitivity upon substituting cited statutes remains merely 0%–50%, with models highly susceptible to adversarial instruction manipulation. Our findings demonstrate that neither scaling nor domain-specific fine-tuning resolves this decoupling problem, exposing significant risks in treating generative legal explanations as reliable evidence of compliance.
📝 Abstract
Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model (a best-effort LoRA reproduction; Section 6) closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts (73.3%-96.4%) exceeds verdict-swap sensitivity by a wide margin, holding without exception across model rankings. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, while the same verdict remains separately vulnerable to adversarial manipulation. Both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on the use of generated legal explanations as compliance or audit artefacts.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought faithfulness
legal reasoning
large language models
adversarial robustness
counterfactual audit
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Audit
Chain-of-Thought Faithfulness
Legal Reasoning
Hidden State Decoding
Adversarial Red-teaming
🔎 Similar Papers
No similar papers found.