🤖 AI Summary
This study addresses the limitations of traditional single-mechanism perspectives in accurately assessing the trustworthiness of large language model (LLM) outputs, which leads to the failure of understanding attribution. We reveal that LLM generation exhibits a "polyphonic" nature arising from the parallel collaboration of multiple mechanisms, thereby transcending the conventional monophonic paradigm. Accordingly, this work proposes a conceptual framework for understanding tailored to polyphonic AI. Methodologically, we employ mechanistic evidence analysis and parallel mechanism reliability modeling to transform abstract trust assessment into actionable, structured analysis. By rendering understanding attribution computationally tractable, this research provides effective guidance for trust-based decision-making regarding AI-generated outputs.
📝 Abstract
When a doctor, a judge, or an engineer must decide whether to trust an AI model's output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs without reintroducing the question of AI understanding in all but name. Yet the question is ill-framed as it stands, because the inherited concept operates within a monophonic paradigm: the idea that a cognitive system's understanding of something must be localised to a single mechanism underpinning all the capacities conferred by such understanding. Drawing on a wide range of mechanistic evidence, we show that LLMs are pervasively polyphonic: outputs emerge from coalitions of parallel mechanisms of uneven reliability, which variously complement, duplicate, or drown out one another, with several coalitions sufficing for a task without any one being indispensable. Polyphony not only complicates attributions of understanding, but renders monophonic inference patterns hazardous. In response, we develop a conception of understanding fit for polyphonic AI. It centres on sound circuitry that is reliably and correctly recruited and in control of outputs. Attributions of understanding thereby become tractable claims about internal organisation, and can do the work of guiding trust in AI.