The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "canonical order problem" in large language models (LLMs) during multi-value relation generation, wherein models tend to organize multi-value entities according to canonical sequences such as alphabetical or chronological order, leading to significant degradation in generation completeness and reliability when deviating from such sequences. To investigate this phenomenon, this work formally introduces the concept of the canonical order problem and employs mechanistic interpretability analysis to examine the internal probability distributions of LLMs, revealing that set generation comprises three distinct stages: retrieval, sorting, and selection. The findings confirm that internal canonical ordering constitutes a critical factor constraining the completeness of multi-value relations, thereby providing a theoretical foundation and novel perspective for constructing highly reliable LLM-based knowledge repositories.
📝 Abstract
Large language models (LLMs) are increasingly used as knowledge bases (KBs) due to the vast amount of knowledge they acquire during pre-training. While many works focus on extracting single relational triples, most real-world relations are multi-valued and require generating sets of entities. In this paper, we investigate how LLMs represent and generate multi-valued relations. We identify the canonical order problem: The probabilistic distributions inside LLMs organize many multi-valued relations according to a canonical ordering (e.g., alphabetical or chronological). Through mechanistic analysis, we show that set generation in LLMs can be thought of in terms of three phases: (1) retrieval of candidate entities, (2) internal sorting, and (3) selection of the next element. As a result, prompts aiming to construct KBs that deviate from this internal canonical ordering lead to a markedly reduced reliability of LLMs when aiming to generate complete sets for multi-valued relations.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Multi-Valued Relations
Canonical Order Problem
Knowledge Bases
Set Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Canonical Order Problem
Multi-Valued Relations
Mechanistic Analysis
Knowledge Bases
🔎 Similar Papers
No similar papers found.