🤖 AI Summary
This study addresses the limitations of existing reward signals for legal large language models, which are typically coarse-grained and lack domain specificity and interpretability. To overcome these issues, we propose a three-dimensional classification framework for legal response quality that integrates style, essential elements, and chain-of-thought reasoning. Building upon this taxonomy, we develop a fine-grained reward model and jointly optimize it using Direct Preference Optimization (DPO), reinforcement learning, and rubric-based training strategies. Experimental results demonstrate that our approach significantly enhances legal model performance across all three dimensions. Furthermore, the proposed reward model delivers interpretable, domain-specific feedback without requiring reference answers, effectively guiding and empowering downstream model alignment and optimization.
📝 Abstract
Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing reward methods, however, often rely on coarse-grained holistic judgments, providing limited domain specificity and interpretability. We introduce LexReward, a taxonomy-driven framework for legal reward modeling. LexReward characterizes legal response quality along three complementary dimensions: Style, covering lexical and syntactic quality; Element, assessing legal subjects, facts, statutes, and decisions; and Chain, evaluating the order, completeness, correctness, and non-redundancy of legal reasoning. For each dimension, we develop rubrics that specify evaluation criteria and quality levels. The resulting rewards are used to construct pairwise preference data for Direct Preference Optimization (DPO) and reward-model training. Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions. The learned reward models, LexRM, also support effective downstream optimization: each dimension-specific reward model improves policy performance in its corresponding dimension through reinforcement learning, without requiring reference answers at reward time. Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.