🤖 AI Summary
This study addresses the lack of interpretability in identifying ambiguous clauses within legal contracts by proposing a post-training framework tailored for small models. Methodologically, it transfers the reasoning capabilities of large language models to lightweight architectures via knowledge distillation, employing a joint optimization objective for classification and rationale generation. Furthermore, this work introduces a novel IRAC-Unlearning prompting technique that effectively balances interpretability with predictive performance. Experimental results demonstrate that a model with merely 250 million parameters achieves state-of-the-art interpretability across multiple benchmarks while matching optimal predictive accuracy. Ultimately, this approach facilitates efficient and lightweight deployment for legal NLP tasks without compromising explanatory depth or performance.
📝 Abstract
Legal contracts contain ambiguities that expose enterprises to financial and legal risks. Some ambiguities allow flexible interpretation without triggering disputes, while others lead to significant legal conflicts. This makes identification alone insufficient, and interpretable rationale analysis essential. We propose LAURA, a post-training framework for interpretable ambiguous clause identification. LAURA leverages knowledge distillation with an IRAC-Unlearning prompting technique to transfer knowledge from a teacher LLM to an open-weight student model (<=1B parameters), which is then trained using a joint objective combining classification and rationale generation losses. The framework supports both legal and non-legal stakeholders in making informed decisions about which ambiguities require further attention. Extensive experiments across 7 baselines and 7 open-weight models demonstrate that LAURA with Flan-T5 (250M) delivers state-of-the-art interpretability over all interpretable baselines while matching the identification performance of the best-performing opaque baseline.