Listwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure Reranking

📅 2026-08-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of medical procedure reranking caused by the lexical gap between patient queries and clinical terminology. For the first time in a real-world health insurance setting, it systematically compares the performance and efficiency of lightweight cross-encoders against large language model (LLM)-based instruction rerankers. The authors fine-tune domain-specific models such as MedCPT and MiniLM-L12 using listwise ranking losses like ListNet and introduce a GPT-4–driven agent-based prompt optimization pipeline to enhance the Qwen3-Reranker-4B. Experimental results demonstrate that a compact cross-encoder with only 109 million parameters outperforms the 4-billion-parameter LLM by 2.6 points in NDCG@3 and by 13.3 points in Spearman correlation, while using merely 1/37th of the parameters—highlighting the efficacy and scalability of small models for specialized domain tasks.
📝 Abstract
Reranking medical procedures against patient queries is a critical component of health insurance information retrieval, complicated by a substantial lexical gap between patient language and clinical nomenclature. We present a systematic comparison of two reranking paradigms for this production task: (1) small cross-encoders (MedCPT, MiniLM-L12) fine-tuned with listwise learning-to-rank objectives across layer freezing configurations, and (2) Qwen3-Reranker-4B, a 4B-parameter instruction reranker whose prompt is iteratively refined via an agentic optimization loop driven by GPT-4.1. On a purpose-built dataset of 2,647 queries across 708 insurance services, we find that a 109M-parameter cross-encoder fine-tuned with ListNet outperforms the 4B-parameter model by 2.6 percentage points on NDCG@3 and 13.3 points on Spearman correlation - at 37x fewer parameters. We report practical findings, a scalable LLM based dataset construction pipeline, and deployment trade-offs relevant to production reranking systems. We release our code and a sample dataset to support reproducibility and adaptation to other domains.
Problem

Research questions and friction points this paper is trying to address.

medical procedure reranking
lexical gap
health insurance information retrieval
patient queries
clinical nomenclature
Innovation

Methods, ideas, or system contributions that make the work stand out.

listwise fine-tuning
agentic instruction tuning
medical procedure reranking
cross-encoder
LLM-based dataset construction