Listwise Explanation of Embedding-Based Rankings via Semantic Chunk Grouping

📅 2026-06-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the granularity mismatch between existing listwise explanation methods—which rely on isolated terms—and the semantic chunk representations leveraged by dense retrievers. To bridge this gap, the authors propose ChunkGroupSHAP, the first approach to incorporate semantic chunk clustering into Shapley value computation. By aligning attribution with the dense retriever’s representation granularity through cross-document semantic groupings, ChunkGroupSHAP enhances attribution consistency while preserving the listwise explanation framework. Experiments reveal that the optimal explanation unit varies with both the retriever and corpus: BM25 performs best with word-level units, dense models like E5 benefit from corpus-level groupings, and heterogeneous retrieval settings gain from query-local groupings. The method demonstrates consistent effectiveness across MS MARCO, FinanceBench, AILACaseDocs, and FinQA benchmarks.
📝 Abstract
Dense embedding rankers score documents through contextual sentence- and passage-level representations. Yet many listwise explanation methods still attribute rankings to isolated words. This feature-unit mismatch leaves word-level features too fragmented for dense semantic ranking. We introduce ChunkGroupSHAP, a listwise Shapley method that clusters semantically related chunks into shared cross-document features. Masking a group perturbs all documents with related evidence, attributing rankings at a granularity closer to dense representations while preserving the listwise setup. Our findings across MS MARCO, FinanceBench, AILACaseDocs, and FinQA with E5 rankers and BM25 show that the best explanation unit is setting-dependent: word features for lexical BM25, corpus-level groups for dense rankers, and query-local grouping for heterogeneous web retrieval. Feature units should thus follow both the ranker's representational granularity and the structure of the retrieved corpus.
Problem

Research questions and friction points this paper is trying to address.

listwise explanation
embedding-based ranking
feature-unit mismatch
semantic chunk grouping
dense rankers
Innovation

Methods, ideas, or system contributions that make the work stand out.

ChunkGroupSHAP
listwise explanation
semantic chunk grouping
dense embedding rankers
feature granularity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hyunkyu Kim
Financial Tech Lab, KakaoBank Corp.
Y
Yeeun Yoo
Financial Tech Lab, KakaoBank Corp.
Y
Youngjun Kwak
Financial Tech Lab, KakaoBank Corp.