Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation in calibration and the lack of cross-model score comparability when transferring pre-generation success probes to multilingual scenarios. By constructing probes from the hidden-layer activations of large language models, this work reveals the failure of English-trained probes to maintain calibration across languages. To overcome this limitation, it proposes a multilingual pooled supervision method coupled with a routing strategy. Experimental results demonstrate that, compared to the English-training baseline, the proposed approach preserves both score comparability and reliability while improving the test success rate of the pooled router by 0.7% and reducing modeling costs by 13.0%. These findings indicate that the method achieves efficient and robust multilingual reasoning routing.
📝 Abstract
Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across languages along three dimensions: (1) whether they preserve the ranking of likely successes and failures (DISCRIMINATION); (2) whether they retain probabilities that match observed success frequencies (CALIBRATION); and (3) whether they produce scores comparable enough across candidate models for cost-aware multilingual routing (UTILITY). Using 3,000 MATH problems in 10 languages and 8 open-weight model configurations, we compare cross-lingual transfer from English-trained probes and equal-budget pooled multilingual probes. English-trained probes retain useful cross-lingual discrimination but become less well calibrated after transfer. Pooled multilingual supervision improves both properties and yields more reliable estimates of success. In routing experiments, the pooled router achieves a 0.7% higher test success rate while reducing modeled cost by 13.0% relative to always selecting the model with the highest average success. These results show that multilingual routing requires success estimates that remain well calibrated and comparable across languages and models.
Problem

Research questions and friction points this paper is trying to address.

cross-lingual calibration
pre-generation success probes
multilingual LLM routing
cost-aware routing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-lingual Calibration
Pre-generation Success Probes
Multilingual LLM Routing
Cost-aware Routing
Pooled Multilingual Supervision
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Andrea Paganelli
The University of Queensland, Polytechnic University of Milan
S
Stefano Civelli
The University of Queensland
P
Pietro Bernardelle
The University of Queensland
Gianluca Demartini
Gianluca Demartini
Professor at the University of Queensland
Information RetrievalSemantic WebHuman ComputationCrowdsourcing