🤖 AI Summary
This study addresses the challenge that language models face in comprehending external API functionalities through black-box interactions when source code is inaccessible. To this end, this work proposes PAU, a benchmark requiring models to explore and understand APIs solely via input-output queries, thereby simulating unsupervised tool cognition scenarios. It introduces the first paradigm for understanding APIs through black-box interaction and incorporates an asymmetric actor-critic (AAC) architecture from robot learning to optimize exploration strategies. The findings reveal that frontier models exhibit insufficient exploration due to overconfidence. Notably, Qwen3-8B fine-tuned with AAC post-training achieves performance comparable to GPT-5-mini, demonstrating the substantial potential of smaller models for black-box API understanding tasks.
📝 Abstract
Empowered by advances in Language Model agents, systems have made substantial strides in code generation and understanding. However, these approaches often rely on read access to the relevant code, an assumption which does not hold when dealing with external APIs. In this work, we introduce the PAU (Python API Understanding) benchmark, where we provide models with black-box, API-level access to code snippets. Models must query the API with exploratory inputs and draw insights from the resulting outputs, with the goal of describing the snippet's true functionality. By treating the code snippets as external tools that must be understood via interaction alone, PAU studies the more general problem of unsupervised tool understanding, specifically for tools implemented as Python methods. Despite recent progress in coding agents, even frontier models struggle to achieve high performance on PAU, with the best model (Claude-4-Opus) failing to understand over 45% of the PAU test set. An investigation into the common error modes reveals that models are overconfident; they often overrate the quality of their current hypothesis, leading to insufficient exploration and premature termination. Finally, we take inspiration from the Asymmetric Actor Critic (AAC) paradigm, frequently used in robot learning, to post-train models for interactive code understanding. Models trained with AAC conduct more active exploration of the APIs, with an AAC-tuned Qwen3-8B model matching the performance of GPT-5-mini.