JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew

πŸ“… 2026-04-20
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the critical challenge of personalizing large language models to accurately emulate individual judges’ judicial reasoning styles under low-resource conditions. The authors propose a novel synthetic-organic supervision pipeline that automatically transforms raw case law into high-quality instruction-tuning data. By integrating causal language modeling, parameter-efficient fine-tuning, and structured processing of legal texts, the approach enables judge-level personalization in a Hebrew-language setting. Experimental results demonstrate that the method significantly outperforms existing techniques across three evaluation tasks, generating judicial reasoning that is semantically, stylistically, and lexically indistinguishable from that of human judges.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
πŸ“ Abstract
Despite significant advances in large language models, personalizing them for individual decision-makers remains an open problem. Here, we introduce a synthetic-organic supervision pipeline that transforms raw judicial decisions into instruction-tuning data, enabling parameter-efficient fine-tuning of personalized models for individual judges in low-resource settings. We compare our approach to state-of-the-art personalization techniques across three different tasks and settings. The results show that Causal Language Modeling followed by synthetically generated instruction-tuning significantly outperforms all other baselines, providing significant improvements across lexical, stylistic, and semantic similarity. Notably, our model-generated outputs are indistinguishable from the reasoning of human judges, highlighting the viability of efficient personalization, even in low-resource settings.
Problem

Research questions and friction points this paper is trying to address.

personalization
judicial reasoning
large language models
low-resource settings
instruction-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

personalized LLMs
synthetic instruction tuning
judicial reasoning
low-resource fine-tuning
parameter-efficient adaptation
πŸ”Ž Similar Papers