When Your LLM Reaches End-of-Life: A Framework for Confident Model Migration in Production Systems
This work addresses the lack of efficient and reliable methods for transferring large language models (LLMs) into production systems by proposing a Bayesian statistical framework for transfer evaluation. The approach leverages Bayesian calibration to align automated evaluation metrics with limited human annotations, significantly improving assessment efficiency while maintaining high quality. It enables reproducible transfer decisions across multiple regions and deployment scenarios by innovatively integrating calibrated automated metrics with human judgment. The framework effectively evaluates substitute models along critical dimensions—including factual correctness, refusal behavior, and stylistic consistency—and has been successfully deployed in a commercial question-answering system handling 5.3 million monthly interactions, where it accurately identified suitable model replacements.