🤖 AI Summary
This study systematically evaluates how text embedding models and retrieval pipeline configurations impact AI search performance, using high-quality evaluation data derived from U.S. city council meeting transcripts. Methodologically, it compares sentence-transformers models (All-MPNet, BGE, GTE) with a generative embedding model (Qwen3-Embedding-8B), incorporating fine-grained text chunking (512 characters), high-dimensional embeddings (4096 dimensions), Milvus indexing (HNSW/IVF), and neural re-ranking. It further introduces a local LLM-driven synthetic data generation framework and a CI/CD-automated evaluation pipeline. Key contributions include: (1) empirical validation that high-dimensional generative embeddings substantially improve long-tail query recall (Top-3 accuracy = 0.571); (2) identification of synergistic gains from fine-grained chunking and neural re-ranking; and (3) proposal of a reproducible, end-to-end optimized evaluation paradigm for AI search systems.
📝 Abstract
We evaluate the performance of various text embedding models and pipeline configurations for AI-driven search systems. We compare sentence-transformer and generative embedding models (e.g., All-MPNet, BGE, GTE, and Qwen) at different dimensions, indexing methods (Milvus HNSW/IVF), and chunking strategies. A custom evaluation dataset of 11,975 query-chunk pairs was synthesized from US City Council meeting transcripts using a local large language model (LLM). The data pipeline includes preprocessing, automated question generation per chunk, manual validation, and continuous integration/continuous deployment (CI/CD) integration. We measure retrieval accuracy using reference-based metrics: Top-K Accuracy and Normalized Discounted Cumulative Gain (NDCG). Our results demonstrate that higher-dimensional embeddings significantly boost search quality (e.g., Qwen3-Embedding-8B/4096 achieves Top-3 accuracy about 0.571 versus 0.412 for GTE-large/1024), and that neural re-rankers (e.g., a BGE cross-encoder) further improve ranking accuracy (Top-3 up to 0.527). Finer-grained chunking (512 characters versus 2000 characters) also improves accuracy. We discuss the impact of these factors and outline future directions for pipeline automation and evaluation.