🤖 AI Summary
This work addresses the limited interpretability and inefficient context utilization in traditional rerankers, which output only scalar relevance scores without refined evidential support. To overcome this, the authors propose Prism-Reranker, a family of multi-scale models built upon Qwen3.5 that jointly models contribution statement generation and evidence passage distillation, yielding both relevance judgments and structured explanations. The approach leverages LLM-as-Judge data curation, hybrid training on real and synthetic queries, keyword-based query rewriting, and a combination of pointwise distillation with supervised fine-tuning to substantially enhance generalization. Evaluated on BEIR-QA subsets, the enhanced Qwen3-Reranker-4B achieves an average NDCG@10 improvement of 1.54, with both contribution clarity and evidence quality validated by LLM-based assessment. The models and training pipeline are publicly released.
📝 Abstract
Modern retrieval pipelines increasingly serve downstream consumers like retrieval-augmented generation (RAG) and autonomous agents that need more than a scalar relevance score. A reranker that only tells the caller "how relevant" forces the agent to dump entire documents into the language-model context, wasting tokens on tangential passages and boilerplate. We introduce Prism-Reranker, a family of reranker models built on Qwen3.5 at four sizes (0.8B, 2B, 4B, 9B) that goes beyond scalar scoring. In addition to the standard yes/no relevance judgement, whenever the verdict is yes the model emits (i) a contribution statement summarizing how the document helps the query, and (ii) an evidence passage: a self-contained rewrite that preserves every query-relevant signal while discarding noise. Prism-Reranker is trained with a hybrid objective combining point-wise distillation from a strong commercial reranker API with supervised fine-tuning on contribution and evidence targets. We curate training data from KaLM-Embedding's open-source aggregation, augmented with real web documents retrieved via commercial search APIs for open-domain queries and LLM-synthesized variants, and rewrite a portion of queries into keyword-style reformulations to adapt the model to agent-issued traffic. To reconcile inconsistent labels across open corpora and obtain crisp binary supervision, we relabel data with an LLM-as-Judge ensemble aggregating votes from five frontier LLMs. On a QA subset of BEIR and on an LLM-judged evaluation of contribution and evidence quality, Prism-Reranker attains solid results across all four sizes. We further show that the same recipe extends existing LLM-based rerankers, augmenting Qwen3-Reranker-4B with contribution and evidence capabilities while improving its average BEIR-QA NDCG@10 by +1.54 over the base model. Model weights, training recipe, and evaluation suite are released.