CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建CESBench,一个包含380个问题的基准测试集,来评估大型语言模型在物联网设备密码工程安全领域的表现,涵盖六个子领域和四种任务类型。
📝 Abstract
For Internet of Things (IoT) devices, a secure algorithm alone is not enough: an attacker with physical access can attack the implementation directly, and its flaws are hard to fix once deployed. Large language models (LLMs) are now used to build and analyze such implementations. LLM benchmarks exist for cryptography and general cybersecurity, but none covers cryptographic engineering. In this paper, we present CESBench, 380 expert-written items across six sub-domains of cryptographic engineering security for IoT devices: side-channel, fault injection, implementation, countermeasures, evaluation, and integration. Four task types target different competences: 209 multiple-choice items test recall, 67 judgment items require a security verdict and its justification, 63 scenario items require an engineering diagnosis, and 41 code tasks are graded by 572 test cases. To validate the benchmark, 11 open-weight and proprietary LLMs answer every item. Multiple-choice and code responses are scored automatically, and judgment and scenario responses by an LLM judge, whose scores are checked against a second judge from another model family and human re-scoring. Composite scores range from 54.4% to 83.6%. The top score on each task type is 98.6% for multiple choice, 95.1% for code, and 88.4% for scenario diagnosis, but only 58.8% for judgment. Across models, 88.5% of verdicts are correct, yet their justifications earn only 53.4% of the rubric marks. Multiple choice is near its ceiling for the strongest models and most code tasks are solved, whereas justifying a security verdict remains the weakest competence. The benchmark, prompts, and per-item results are public.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Cryptographic Engineering Security
IoT Devices
Benchmarking
Physical Access Attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

CESBench
cryptographic engineering security
IoT devices
large language models
benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wenquan Zhou
School of Cyberspace Science and Technology, Beijing Institute of Technology, Beijing 100081, China; and also with the State Key Laboratory of Cryptography and Digital Economy Security, Shandong University, Qingdao 266237, China
An Wang
An Wang
Tokyo Institute of Technology
Nature Languge ProcessingDeep LearningLarge Language Model
J
Jing Liang
School of Cyberspace Science and Technology, Beijing Institute of Technology, Beijing 100081, China
P
Peien Feng
Xuteli School, Beijing Institute of Technology, Beijing 100081, China
J
Jingqi Zhang
School of Cyberspace Science and Technology, Beijing Institute of Technology, Beijing 100081, China
Y
Yaoling Ding
School of Cyberspace Science and Technology, Beijing Institute of Technology, Beijing 100081, China
L
Liehuang Zhu
School of Cyberspace Science and Technology, Beijing Institute of Technology, Beijing 100081, China