🤖 AI Summary
Existing IR benchmarks inadequately capture the complex, domain-specific information needs inherent in banking operations, while constructing realistic, legally compliant evaluation datasets remains hindered by regulatory constraints and high annotation costs. To address this, we propose a systematic, large language model–based query generation framework that integrates single-document and multi-document joint sampling, augmented by a novel reasoning-enhanced answerability assessment mechanism—significantly improving alignment between generated queries and authentic business requirements. Our method explicitly supports modeling of complex multi-document retrieval scenarios. Leveraging it, we construct KoBankIR—the first Korean-language, banking-domain IR benchmark—comprising 815 high-quality, expert-validated queries. Empirical evaluation reveals substantial performance degradation of state-of-the-art retrieval models on multi-document tasks under KoBankIR, confirming its rigor and practical relevance. KoBankIR provides a reproducible, legally compliant, and high-fidelity evaluation foundation for financial-domain IR research.
📝 Abstract
As financial applications of large language models (LLMs) gain attention, accurate Information Retrieval (IR) remains crucial for reliable AI services. However, existing benchmarks fail to capture the complex and domain-specific information needs of real-world banking scenarios. Building domain-specific IR benchmarks is costly and constrained by legal restrictions on using real customer data. To address these challenges, we propose a systematic methodology for constructing domain-specific IR benchmarks through LLM-based query generation. As a concrete implementation of this methodology, our pipeline combines single and multi-document query generation with an enhanced and reasoning-augmented answerability assessment method, achieving stronger alignment with human judgments than prior approaches. Using this methodology, we construct KoBankIR, comprising 815 queries derived from 204 official banking documents. Our experiments show that existing retrieval models struggle with the complex multi-document queries in KoBankIR, demonstrating the value of our systematic approach for domain-specific benchmark construction and underscoring the need for improved retrieval techniques in financial domains.