Automated Consistency Analysis of LLMs

📅 2024-10-28
🏛️ International Conference on Trust, Privacy and Security in Intelligent Systems and Applications
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) exhibit inconsistent and low-reliability responses in high-stakes domains such as cybersecurity, undermining their trustworthiness for critical applications. Method: This paper formally defines LLM response consistency for the first time and introduces a dual-path verification framework—comprising intra-model self-verification and inter-model cross-verification—alongside a security-oriented consistency evaluation framework and benchmark. The methodology includes formal consistency modeling, semantic consistency measurement, multi-model comparative analysis, and domain-specific question-answering benchmark design. Contribution/Results: Empirical evaluation across mainstream models—including GPT-4o Mini, GPT-3.5, Gemini, Cohere, and Llama3—reveals significant response inconsistency in cybersecurity QA tasks, highlighting urgent reliability gaps. The proposed framework provides a reproducible, principled assessment paradigm and empirical foundation for deploying LLMs safely in high-risk operational environments.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessComputer Vision: Large Vision Models

Application Category

Security and Privacy: Large-scale security measurementsSearch and Retrieval-Augmented AI: Large language models for searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Generative AI (Gen AI) with large language models (LLMs) are being widely adopted across the industry, academia and government. Cybersecurity is one of the key sectors where LLMs can be and/or are already being used. There are a number of problems that inhibit the adoption of trustworthy Gen AI and LLMs in cybersecurity and such other critical areas. One of the key challenge to the trustworthiness and reliability of LLMs is: how consistent an LLM is in its responses?In this paper, we have analyzed and developed a formal definition of consistency of responses of LLMs. We have formally defined what is consistency of responses and then develop a framework for consistency evaluation. The paper proposes two approaches to validate consistency: self-validation, and validation across multiple LLMs. We have carried out extensive experiments for several LLMs such as GPT4oMini, GPT3.5, Gemini, Cohere, and Llama3, on a security benchmark consisting of several cybersecurity questions: informational and situational. Our experiments corroborate the fact that even though these LLMs are being considered and/or already being used for several cybersecurity tasks today, they are often inconsistent in their responses, and thus are untrustworthy and unreliable for cybersecurity.
Problem

Research questions and friction points this paper is trying to address.

Defining LLM response consistency
Framework for consistency evaluation
Validating LLM consistency in cybersecurity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Formal consistency definition
Self-validation approach
Multi-LLM validation framework
💼 Related Jobs
No related jobs found.