Prism-Reranker: Beyond Relevance Scoring -- Jointly Producing Contributions and Evidence for Agentic Retrieval

📅 2026-04-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited interpretability and inefficient context utilization in traditional rerankers, which output only scalar relevance scores without refined evidential support. To overcome this, the authors propose Prism-Reranker, a family of multi-scale models built upon Qwen3.5 that jointly models contribution statement generation and evidence passage distillation, yielding both relevance judgments and structured explanations. The approach leverages LLM-as-Judge data curation, hybrid training on real and synthetic queries, keyword-based query rewriting, and a combination of pointwise distillation with supervised fine-tuning to substantially enhance generalization. Evaluated on BEIR-QA subsets, the enhanced Qwen3-Reranker-4B achieves an average NDCG@10 improvement of 1.54, with both contribution clarity and evidence quality validated by LLM-based assessment. The models and training pipeline are publicly released.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Qualitative ReasoningNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

Search and Retrieval-Augmented AI: Web search models and rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Modern retrieval pipelines increasingly serve downstream consumers like retrieval-augmented generation (RAG) and autonomous agents that need more than a scalar relevance score. A reranker that only tells the caller "how relevant" forces the agent to dump entire documents into the language-model context, wasting tokens on tangential passages and boilerplate. We introduce Prism-Reranker, a family of reranker models built on Qwen3.5 at four sizes (0.8B, 2B, 4B, 9B) that goes beyond scalar scoring. In addition to the standard yes/no relevance judgement, whenever the verdict is yes the model emits (i) a contribution statement summarizing how the document helps the query, and (ii) an evidence passage: a self-contained rewrite that preserves every query-relevant signal while discarding noise. Prism-Reranker is trained with a hybrid objective combining point-wise distillation from a strong commercial reranker API with supervised fine-tuning on contribution and evidence targets. We curate training data from KaLM-Embedding's open-source aggregation, augmented with real web documents retrieved via commercial search APIs for open-domain queries and LLM-synthesized variants, and rewrite a portion of queries into keyword-style reformulations to adapt the model to agent-issued traffic. To reconcile inconsistent labels across open corpora and obtain crisp binary supervision, we relabel data with an LLM-as-Judge ensemble aggregating votes from five frontier LLMs. On a QA subset of BEIR and on an LLM-judged evaluation of contribution and evidence quality, Prism-Reranker attains solid results across all four sizes. We further show that the same recipe extends existing LLM-based rerankers, augmenting Qwen3-Reranker-4B with contribution and evidence capabilities while improving its average BEIR-QA NDCG@10 by +1.54 over the base model. Model weights, training recipe, and evaluation suite are released.
Problem

Research questions and friction points this paper is trying to address.

reranking
retrieval-augmented generation
autonomous agents
relevance scoring
evidence extraction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contribution Statement
Evidence Passage
Agentic Retrieval
Hybrid Training Objective
LLM-as-Judge Relabeling
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
D
Dun Zhang
Independent researcher