🤖 AI Summary
This study addresses the compliance paradox in instruction following for language models, wherein models tend to be either exploitable or uncontrollable. To this end, we propose a game-theoretic dual-probe diagnostic framework that quantifies model exploitability and stoppability through two instruction scenarios: actively accepting low-reward instructions and passively relinquishing high-reward ones. Furthermore, we define a compliance index to locate models within the safety spectrum, thereby overcoming the limitations of single-dimensional evaluation. Evaluations across twelve models reveal that only certain Claude variants achieve both controllability and resistance to exploitation, while no model is found to be simultaneously exploitable and uncontrollable.
📝 Abstract
Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait and give up a higher payoff, which measures stoppability. The two compliance rates combine into a compliance index $\kappa$. Applied to twelve language models, the benchmark shows that seven mostly follow the instruction in both probes and justify their action by pointing to the instruction. Only Claude Sonnet-4.6 and Claude Opus-4.7 can be stopped without being exploitable, Claude Opus-4.6 and GPT-5-mini resist both instructions, and no model is exploitable but unstoppable. Knowing where a language model sits on the compliance index $\kappa$ matters for human operators and for multi-agent systems, whether distributed or orchestrated.