🤖 AI Summary
This study addresses the absence of capability evaluations for large language models (LLMs) in cryptographic protocol verification by proposing the first systematic assessment framework. Methodologically, it integrates LLMs with symbolic protocol verifiers through theoretical analysis and empirical validation, introducing CRoST, a novel proof-tree-based metric designed to quantify proof coverage. Experimental results demonstrate that models achieve an average coverage rate of 38.82% and effectively generate auxiliary lemmas; however, they also exhibit significant failure modes and computational overhead in complex protocol scenarios. By revealing both the potential and limitations of LLM-assisted protocol verification, this work establishes a critical benchmark and methodological foundation for future research in this direction.
📝 Abstract
Large language models (LLMs) have shown strong potential for assisting software and security analysis tasks, yet their effectiveness in cryptographic symbolic protocol verification remains insufficiently understood. In this paper, we conduct the first systematic evaluation of the capability of state-of-the-art LLMs in cryptographic symbolic protocol verification. To quantify this capability, we propose \textsc{CRoST} (Coverage Rate of Solve Tree), a proof-based metric derived from the verifier's proof skeleton that measures the similarity between generated lemmas and reference lemmas. We then establish the rationale of \textsc{CRoST} through both theoretical analysis and empirical validation. The evaluation results show that state-of-the-art models achieve 38.82\% coverage on average, with 14.4\% of generated lemmas exceeding 80\% coverage, indicating that LLMs can already generate useful lemmas to a certain extent. However, they still exhibit non-trivial failure modes on complex multi-phase protocols, show diminishing returns under naive scaling, and incur substantial verification overhead. These findings clarify the practical potential and limitations of LLMs for protocol verification and motivate future work on complex real-world protocols.