🤖 AI Summary
This study addresses the vulnerability of language models to blind reliance on unreliable external workflows, highlighting their lack of selective discernment. To tackle this, we introduce the "thinking outside the box" paradigm and the Box²-Bench benchmark to evaluate models' capacity for selective reliance. We optimize model behavior through counterfactual supervised fine-tuning (SFT) and outcome-based reinforcement learning (RL), enhancing their ability to circumvent misleading information while leveraging valid guidance. This work reveals a novel dimension of agent reliability beyond task performance, demonstrating that such capabilities transfer to other external information scenarios. Experiments show that our approach significantly improves model robustness, enabling effective adoption of reliable guidance and resistance to unreliable interference, thereby enhancing peer error correction and memory resilience against perturbations.
📝 Abstract
Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.