Prospective Interpretation Risk: Principled Communication Control Between LLMs

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of LLM-based multi-agent systems to communication failures arising from potential receiver misinterpretation of messages. To mitigate this, it introduces a prospective interpretation risk metric alongside a theory of explanatory information value, enabling pre-transmission prediction and optimization of query strategies. Furthermore, black-box probing is employed to model receiver types, integrating Bayesian posterior inference with supervised learning via heterogeneous frozen models to guide dynamic message revision and selection. Empirical evaluations demonstrate that the proposed approach reduces calibration error by 68% and interpretation failure rate by 44%, significantly outperforming existing baseline methods at lower computational cost.
📝 Abstract
Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in heterogeneous systems, where capable receivers can reconstruct different tasks from the same message. We model this as a sender-receiver problem with a latent receiver type and define prospective interpretation risk (PIR): the probability that a receiver reconstructs a task other than intended. Rather than model an LLM's full input-output behaviour, we use black-box probes relating messages, intended tasks, and receiver-specific reconstructions, yielding scalable supervision while separating interpretation from downstream capability failure. Offline, heterogeneous frozen receivers provide supervision for receiver-conditioned risk and the effects of predefined mutable message features. At deployment, history induces a posterior over receiver types, guiding message revision and selection. We introduce value of interpretation information (VoII), querying for receiver information only when its expected communication benefit exceeds its cost. Our theory characterises when receiver information has decision value and bounds such queries. Empirically, interpretation-failure rates vary by 4-13x across receivers. Receiver information reduces PIR calibration error by 68% relative to a receiver-agnostic predictor, largely by correcting receiver-specific risk levels. PIR-guided revision reduces interpretation failure by 44% relative to the original message and 40% relative to a generic rewrite, mostly through a repair that helps every receiver. VoII outperforms information-gain and random querying at matched cost on the interpretation objective it optimises, lowering interpretation failure from 3.84% to 3.79% while querying 18.2% of episodes.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Prospective Interpretation Risk
Multi-agent Communication
Sender-Receiver Problem
Heterogeneous Systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prospective Interpretation Risk
Value of Interpretation Information
Sender-Receiver Communication
Heterogeneous LLM Agents
Black-box Probes
🔎 Similar Papers
No similar papers found.