🤖 AI Summary
Existing large language models (LLMs) suffer from task interference and rigid format constraints in JSON-based tool calling, leading to degraded performance and poor robustness. To address this, we propose Natural Language Tools (NLT), a novel framework that introduces a pure natural-language paradigm for tool invocation—decoupling tool selection from response generation and eliminating reliance on structured output formats. NLT is fully compatible with any open-weight LLM, requiring no native tool-calling support or architectural modifications. In cross-domain evaluations spanning customer service and mental health assistance, NLT achieves an average accuracy improvement of 18.4 percentage points across 10 models and 6,400 test instances, reduces output variance by 70%, and demonstrates strong robustness against prompt perturbations. Notably, NLT enables open-weight models to outperform state-of-the-art closed-source flagship models in tool-calling performance for the first time.
📝 Abstract
We present Natural Language Tools (NLT), a framework that replaces programmatic JSON tool calling in large language models (LLMs) with natural language outputs. By decoupling tool selection from response generation, NLT eliminates task interference and format constraints that degrade tool call performance. When evaluated across 10 models and 6,400 trials spanning customer service and mental health domains, NLT improves tool calling accuracy by 18.4 percentage points while reducing output variance by 70%. Open-weight models see the largest gains, surpassing flagship closed-weight alternatives, with implications for model training in both reinforcement learning and supervised fine-tuning stages. These improvements persist under prompt perturbations and extend tool-calling capabilities to models lacking native support.