🤖 AI Summary
Large language models (LLMs) suffer from uncontrollable output safety and quality. Method: We propose a rejection-based safety assurance framework grounded in stochastic voting. It employs an ensemble of multiple independent checkers, integrates randomized re-generation, and applies an adaptive threshold optimization algorithm to jointly assess output safety after each generation. Crucially, it jointly models and optimizes both failure rate and computational cost, yielding theoretically provable asymptotic safety. Contributions/Results: (1) The first output assurance framework enabling reliable failure-rate estimation under few-shot settings; (2) An optimal exponential trade-off between failure rate and computational cost; (3) A lightweight, plug-and-play deployment scheme requiring no additional training. Experiments demonstrate accurate system behavior prediction even with limited labeled data, significantly enhancing LLM output safety, controllability, and practical utility.
📝 Abstract
This paper proposes a new method for preventing unsafe or otherwise low quality large language model (LLM) outputs, by leveraging the stochasticity of LLMs. We propose a system whereby LLM checkers vote on the acceptability of a generated output, regenerating it if a threshold of disapproval is reached, until sufficient checkers approve. We further propose estimators for cost and failure rate, and based on those estimators and experimental data tailored to the application, we propose an algorithm that achieves a desired failure rate at the least possible cost. We demonstrate that, under these models, failure rate decreases exponentially as a function of cost when voter count and threshold are chosen according to the algorithm, and that the models reasonably estimate the actual performance of such a system in action, even with limited data.