Alignment Between the Decision-Making Logic of LLMs and Human Cognition: A Case Study on Legal LLMs

📅 2024-10-06
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the misalignment between large language models’ (LLMs) reasoning logic and human cognition in legal domains. We propose the first fine-grained evaluation framework explicitly designed for cognitive alignment—moving beyond mere output correctness to interrogate internal reasoning processes. Grounded in interactive attribution, the framework explicitly models LLMs’ raw decision logic as quantifiable objects and establishes multi-dimensional logical consistency metrics grounded in mathematical fidelity theory. Empirical evaluation on legal reasoning tasks reveals that over 60% of correctly answered instances exhibit internal reasoning that significantly violates domain-specific human knowledge and established legal inference patterns. This study is the first to systematically expose the pervasive “superficially correct but logically flawed” behavior of LLMs, empirically validating the necessity of logic-level assessment. Our framework introduces a novel paradigm for enhancing model trustworthiness and enabling effective human-AI collaboration in high-stakes legal applications.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningNatural Language Processing: (Large) Language Models

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
This paper presents a method to evaluate the alignment between the decision-making logic of Large Language Models (LLMs) and human cognition in a case study on legal LLMs. Unlike traditional evaluations on language generation results, we propose to evaluate the correctness of the detailed decision-making logic of an LLM behind its seemingly correct outputs, which represents the core challenge for an LLM to earn human trust. To this end, we quantify the interactions encoded by the LLM as primitive decision-making logic, because recent theoretical achievements have proven several mathematical guarantees of the faithfulness of the interaction-based explanation. We design a set of metrics to evaluate the detailed decision-making logic of LLMs. Experiments show that even when the language generation results appear correct, a significant portion of the internal inference logic contains notable issues.
Problem

Research questions and friction points this paper is trying to address.

Analyzing LLM inference patterns for legal judgments
Identifying incorrect LLM representations via human knowledge
Evaluating faithfulness of interaction-based inference patterns
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyze LLM inference patterns for judgment correctness
Quantify input phrase interactions as primitive patterns
Design metrics to evaluate detailed inference logic
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shanghai Jiao Tong University | Beijing Institute for General Artificial Intelligence
L
Lu Chen
Shanghai Jiao Tong University, Shanghai, China
Y
Yuxuan Huang
Shanghai Jiao Tong University, Shanghai, China
Y
Yixing Li
Shanghai Jiao Tong University, Shanghai, China
Yaohui Jin
Yaohui Jin
Shanghai Jiao Tong University
S
Shuai Zhao
Shanghai Jiao Tong University, Shanghai, China
Z
Zilong Zheng
Beijing Institute for General Artificial Intelligence, China
Quanshi Zhang
Quanshi Zhang
Shanghai Jiao Tong University
Interpretable Machine Learning