🤖 AI Summary
This work addresses the misalignment between large language models’ (LLMs) reasoning logic and human cognition in legal domains. We propose the first fine-grained evaluation framework explicitly designed for cognitive alignment—moving beyond mere output correctness to interrogate internal reasoning processes. Grounded in interactive attribution, the framework explicitly models LLMs’ raw decision logic as quantifiable objects and establishes multi-dimensional logical consistency metrics grounded in mathematical fidelity theory. Empirical evaluation on legal reasoning tasks reveals that over 60% of correctly answered instances exhibit internal reasoning that significantly violates domain-specific human knowledge and established legal inference patterns. This study is the first to systematically expose the pervasive “superficially correct but logically flawed” behavior of LLMs, empirically validating the necessity of logic-level assessment. Our framework introduces a novel paradigm for enhancing model trustworthiness and enabling effective human-AI collaboration in high-stakes legal applications.
📝 Abstract
This paper presents a method to evaluate the alignment between the decision-making logic of Large Language Models (LLMs) and human cognition in a case study on legal LLMs. Unlike traditional evaluations on language generation results, we propose to evaluate the correctness of the detailed decision-making logic of an LLM behind its seemingly correct outputs, which represents the core challenge for an LLM to earn human trust. To this end, we quantify the interactions encoded by the LLM as primitive decision-making logic, because recent theoretical achievements have proven several mathematical guarantees of the faithfulness of the interaction-based explanation. We design a set of metrics to evaluate the detailed decision-making logic of LLMs. Experiments show that even when the language generation results appear correct, a significant portion of the internal inference logic contains notable issues.