Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek
This study addresses the insufficient robustness of tool-calling retrieval to user input variations, such as diacritic removal and Latin transliteration, within small bilingual knowledge bases. Using the KyGround benchmark, we systematically compare tool calling with vector-based Retrieval-Augmented Generation (RAG), employing Claude Haiku as both the routing and generation model alongside custom tool agent interfaces. Results indicate that literal matching deficiencies cause elevated refusal rates in tool agents, whereas vector RAG achieves significantly higher accuracy on standardized Greek question answering. The primary contribution of this work lies in revealing how diverse input forms substantially affect retrieval performance and demonstrating that incorporating stemming and diacritic-removal preprocessing effectively enhances the robustness of tool calling, thereby considerably narrowing its performance gap with vector RAG.