A Voter-Based Stochastic Rejection-Method Framework for Asymptotically Safe Language Model Outputs

📅 2024-07-24
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) suffer from uncontrollable output safety and quality. Method: We propose a rejection-based safety assurance framework grounded in stochastic voting. It employs an ensemble of multiple independent checkers, integrates randomized re-generation, and applies an adaptive threshold optimization algorithm to jointly assess output safety after each generation. Crucially, it jointly models and optimizes both failure rate and computational cost, yielding theoretically provable asymptotic safety. Contributions/Results: (1) The first output assurance framework enabling reliable failure-rate estimation under few-shot settings; (2) An optimal exponential trade-off between failure rate and computational cost; (3) A lightweight, plug-and-play deployment scheme requiring no additional training. Experiments demonstrate accurate system behavior prediction even with limited labeled data, significantly enhancing LLM output safety, controllability, and practical utility.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)Reasoning under Uncertainty: Stochastic Optimization

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
This paper proposes a new method for preventing unsafe or otherwise low quality large language model (LLM) outputs, by leveraging the stochasticity of LLMs. We propose a system whereby LLM checkers vote on the acceptability of a generated output, regenerating it if a threshold of disapproval is reached, until sufficient checkers approve. We further propose estimators for cost and failure rate, and based on those estimators and experimental data tailored to the application, we propose an algorithm that achieves a desired failure rate at the least possible cost. We demonstrate that, under these models, failure rate decreases exponentially as a function of cost when voter count and threshold are chosen according to the algorithm, and that the models reasonably estimate the actual performance of such a system in action, even with limited data.
Problem

Research questions and friction points this paper is trying to address.

Preventing unsafe language model outputs through stochastic rejection methods
Achieving exponentially decreasing failure rates with Pareto-optimal costs
Enabling small language models to constrain complex models' outputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Voter-based stochastic rejection method for safety
Repeated checking with regeneration for outputs
Exponential failure rate reduction with cost scaling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Georgia Institute of Technology
J
Jake R. Watts
Georgia Institute of Technology
J
Joel Sokol
Georgia Institute of Technology