Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of inferring latent preferences and the accumulation of belief update errors in LLM-based multi-agent orchestration by proposing the HARP framework. Its core innovation lies in shifting preference beliefs from prompt text to a numerical posterior space, enabling closed-loop updates via Bayes' rule. Furthermore, this work decouples estimation from reasoning and introduces a discriminative exploration reward (HARP+), with theoretical analysis proving it achieves an optimal Bayesian regret bound. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines across multiple scenarios, establishing itself as the strongest non-oracle method and effectively enhancing agent coordination efficiency.
📝 Abstract
LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pursues but does not reveal. Inferring such hidden preferences from behavior has been a subject of long-standing research in game theory and multi-agent systems. The core challenge lies in maintaining a belief over every agent's preference and updating it from the agents' observed actions. Existing LLM orchestrators carry that belief as prompt text with no explicit update rule. This lets early errors persist and propagate rather than be corrected. We therefore propose \textbf{HARP} (Heterogeneous-preference Agent oRchestration via Preference inference), a novel framework that moves the belief out of the prompt. Specifically, HARP maintains one numeric posterior per agent over a finite set of candidate preferences and updates it in closed form by Bayes' rule. The language model supplies only actions and per-candidate likelihoods, so estimation is decoupled from its reasoning. We prove that HARP attains the same $\tilde O(\sqrt K)$ Bayesian regret as explicit joint inference when the factorization is exact. Furthermore, HARP\textsuperscript{+} augments planning with a bonus for actions that distinguish the candidates, so inference continues even when the optimal action is uninformative. Empirical results on three substrates, ranging from payoffs the preferences fully determine, through payoffs that depend on more than them, to scales where explicit joint inference is infeasible, demonstrate that HARP\textsuperscript{+} is the strongest non-oracle method across the class our theory identifies.
Problem

Research questions and friction points this paper is trying to address.

LLM orchestration
heterogeneous preferences
preference inference
multi-agent systems
belief updating
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Orchestration
Preference Inference
Bayesian Updating
Multi-Agent Systems
Exploration Bonus