๐ค AI Summary
This study addresses the insufficient robustness of tool-calling retrieval to user input variations, such as diacritic removal and Latin transliteration, within small bilingual knowledge bases. Using the KyGround benchmark, we systematically compare tool calling with vector-based Retrieval-Augmented Generation (RAG), employing Claude Haiku as both the routing and generation model alongside custom tool agent interfaces. Results indicate that literal matching deficiencies cause elevated refusal rates in tool agents, whereas vector RAG achieves significantly higher accuracy on standardized Greek question answering. The primary contribution of this work lies in revealing how diverse input forms substantially affect retrieval performance and demonstrating that incorporating stemming and diacritic-removal preprocessing effectively enhances the robustness of tool calling, thereby considerably narrowing its performance gap with vector RAG.
๐ Abstract
Assistants grounded in a small, frequently edited knowledge base can retrieve through tool calls to a live data interface or through vector retrieval-augmented generation (RAG). We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes. With Claude Haiku 4.5 as router and answer model, a reconstruction of the platform's tool agent answered 71.6\% of canonical Greek questions correctly and vector RAG 95.3\% (difference $-23.6$ percentage points, 95\% CI $-33.1$ to $-15.1$). Letting the router write the vector query changed nothing, and placing the whole knowledge base of about 26,000 tokens in the prompt reached 99.3\%. The tool agent's losses arose in retrieval. Its literal searches returned nothing when the router's arguments did not occur verbatim in a record, for example when it transliterated Greek into Latin script or combined words that occur in a record but not as one phrase, and the agent then abstained. Unaccented and capitalised questions cost the tool agent about 20 points and vector RAG at most 2; accent-insensitive search removed this loss, and matching stemmed tokens raised the tool agent to 83.8\% on canonical Greek. Greeklish cost both designs about 21 to 32 points. Tool interfaces for community knowledge bases need search that tolerates how users type.