Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

📅 2026-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of text embedding and retrieval for low-resource, morphologically complex agricultural documents in Khmer by systematically evaluating four chunking strategies—recursive, Khmer-aware, sentence-based, and large language model–based—within a retrieval-augmented generation (RAG) framework. Using the BGE-M3 multilingual embedding model and FAISS for dense retrieval, the authors conduct a multidimensional assessment incorporating L2 distance, Khmer IoU, answer relevance, and 5-fold cross-validation. The work reveals, for the first time, the critical impact of chunk granularity and structural preservation on retrieval performance in low-resource languages. The recursive chunking strategy with 300-character segments achieves the best results (L2 distance: 0.4295; answer relevance: 0.8663), significantly outperforming sentence-based chunking (p = 0.0121).
📝 Abstract
In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Augmented Generation (RAG) framework applied to Khmer agricultural documents. The document chunks are encoded using the BGE-M3 multilingual embedding model and retrieved using the FAISS library. Performance is evaluated using four metrics: Average Retrieval Score (L2 distance), Answer Relevance, Khmer Coverage, and Khmer Intersection over Union, all measured against ground-truth question-answer pairs. For evaluation, we perform 5-fold cross-validation over 18 question-answer pairs. We observe the best performance for the character-based Recursive chunking method with a chunk size of 300 characters, achieving the lowest L2 distance (0.4295 +- 0.0461), highest Answer Relevance (0.8663 +- 0.0199), and highest Khmer IoU (0.6441 +- 0.0347). A paired t-test shows a statistically significant improvement over the Sentence-Based chunking method in L2 distance (p = 0.0121). These results highlight the importance of segmentation granularity and structural preservation for optimizing dense retrieval in morphologically complex, low-resource languages such as Khmer.
Problem

Research questions and friction points this paper is trying to address.

chunking
low-resource language
text embedding
agricultural documents
Khmer
Innovation

Methods, ideas, or system contributions that make the work stand out.

chunking strategy
low-resource language
text embedding
RAG
Khmer NLP
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sovandara Chhoun
Department of Big Data, Chungbuk National University, Cheongju-si, South Korea
P
Pichdara Po
Department of Computer Science, Chungbuk National University, Cheongju-si, South Korea
S
Sereiwathna Ros
Department of Computer Science, Chungbuk National University, Cheongju-si, South Korea
W
Wan-Sup Cho
BigDataLabs Co., Ltd. Department of Management Information Systems, Chungbuk National University, South Korea
S
Saksonita Khoeurn
BigDataLabs Co., Ltd. Department of Management Information Systems, Chungbuk National University, South Korea