Do Static Embeddings Add Value to Hybrid Dutch Retrieval?

πŸ“… 2026-08-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates whether static embeddings retain complementary value in hybrid retrieval systems that combine lexical and Transformer-based approaches. On five Dutch-language tasks from the MTEB-NL benchmark, the authors integrate BM25, Qwen3-Embedding-0.6B, and two multilingual static embeddings using reciprocal rank fusion (RRF), evaluating performance rigorously via ten-fold query-level cross-validation, bootstrap confidence intervals, and sign randomization tests. Results show that fusing BM25 with Qwen significantly improves MRR on four tasks (by up to +0.061), whereas adding static embeddings yields no significant gains and occasionally degrades performance. The optimal fusion weights consistently lie on the BM25–Qwen boundary, supporting a dual-retriever architecture as a robust default and challenging conventional practices that rely on standalone embedding evaluations.
πŸ“ Abstract
Embedding benchmarks measure standalone model quality, but they do not establish whether a low-cost retriever contributes complementary ranking information once lexical and transformer-based retrieval are already combined. We present a controlled evaluation of this question across Dutch retrieval tasks from the Massive Text Embedding Benchmark for Dutch (MTEB-NL). Weighted reciprocal rank fusion (RRF) combines Best Matching 25 (BM25), Qwen/Qwen3-Embedding-0.6B (Qwen), and two multilingual static embedding models. Five datasets comprising 14,500 queries and 786,573 documents are scored exhaustively, and fusion weights are searched on a simplex in increments of 0.1. Ten-fold query-level cross-validation selects weights on nine folds and evaluates them on the held-out fold; paired bootstrap confidence intervals and sign-randomisation tests quantify the resulting differences. Fusion improves over the training-selected individual retriever by 0.061 mean reciprocal rank (MRR) on Dutch News, 0.029 on VABB, 0.004 on WebFAQ NL, and 0.025 on Wikipedia NL, while matching BM25 on Open Tender. All four positive differences remain distinguishable from zero after Holm correction. No unrestricted fold assigns positive weight to either static retriever: all 50 selections lie on the BM25-Qwen edge, and forcing a static contribution reduces effectiveness. Leave-one-dataset-out selection chooses equal BM25-Qwen weighting in every iteration and outperforms the cross-domain-selected individual retriever on every held-out task. The results support a two-retriever lexical-transformer architecture as a robust tested default across the evaluated Dutch tasks and show that standalone benchmark performance is insufficient to establish marginal value in hybrid retrieval.
Problem

Research questions and friction points this paper is trying to address.

static embeddings
hybrid retrieval
Dutch retrieval
lexical retrieval
transformer-based retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

hybrid retrieval
static embeddings
reciprocal rank fusion
controlled evaluation
marginal value
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
A
AntΓ³nio Pereira Barata