RAIM: Robust Aggregation of Inexpensive Models for Hallucination Detection

📅 2026-09-30
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
This study addresses the high cost and limited throughput of employing frontier large language models as judges for hallucination detection by proposing RAIM, a framework that aggregates inexpensive open-source small models to enable efficient automated faithfulness evaluation. Methodologically, RAIM introduces a robust aggregation mechanism resilient to correlated errors alongside an output-based acceptability test, integrating these with cross-fitted stacked logistic regression to construct a multi-model ensemble system. Experimental results demonstrate that RAIM retains 93% decision agreement with frontier models at merely one-sixty-fourth of the computational cost, achieving performance comparable to specialized detectors across multiple benchmarks. Consequently, this work provides a scalable, low-cost, and high-precision solution for monitoring hallucinations in large language models.
📝 Abstract
Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary frontier models, costly and ill-suited to high-throughput monitoring. We investigate whether a panel of cheap open-weight judges (4--9B) can be aggregated to stand in for a frontier one, what the substitution sacrifices, and when it is worth making. We propose RAIM, an aggregation scheme robust to the members' correlated errors, coupling a cross-fitted stacked logistic regression with an admissibility test that, read from the members' own outputs, identifies when aggregating them improves on their best member and stays within reach of the frontier judge. We instantiate RAIM with ten judges from disjoint families across eight faithfulness benchmarks. Against Claude Sonnet, the panel retains a median 93% of its Cohen's $Îș$ and gives up only 2.9 points of balanced accuracy on average; read as paired differences, it clearly improves on one benchmark and clearly worsens on three (only two by a non-negligible margin), leaving four unresolved. At a sixty-fourth of the frontier's inference price, the operative expense is a one-time in-domain calibration on 50--100 labelled records. The panel is also competitive with purpose-trained detectors on their home benchmarks (within 1.3 accuracy points of GPT-4o and 1.9 of the LLM-AggreFact leader), and beats the strongest one we reran by 6 points on our grounded sets. Whether aggregation pays depends on the members themselves: where several capable members err on different items, the panel improves on its best judge and approaches the frontier; where one dominates, the stacker recovers the leader, and only there does the frontier remain materially ahead. Both conditions are read off the calibration set at no further cost, so a cheap panel can stand in for a frontier one wherever this audit admits it.
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
faithfulness evaluation
LLM-as-a-judge
model aggregation
cost-efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hallucination Detection
Model Aggregation
Stacked Logistic Regression
LLM-as-a-Judge
Faithfulness Evaluation
🔎 Similar Papers
No similar papers found.