Demystifying Agent Skills for Smart Contract Auditing: Design, Effectiveness, Behavioral Impact

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of clarity in skill design and the absence of quantitative evaluation regarding the effectiveness of LLM-based agents for smart contract auditing. We curate a corpus of 83 auditing skills and conduct comparative experiments across multiple models and agent configurations using the EVMBench benchmark. To our knowledge, this is the first work to systematically quantify the differential gains that specific skills confer upon various models and frameworks while identifying execution bottlenecks. Experimental results demonstrate that incorporating these skills yields a 22.8% improvement in detection scores and a 43.2% increase in capture rewards for Codex/GPT-5.5. Furthermore, this paper elucidates the behavioral mechanisms through which skills influence agent performance and releases the first open-source skill corpus dedicated to smart contract auditing.
📝 Abstract
LLM agents, notably Claude Code and OpenAI Codex, are emerging as versatile tools beyond coding agents only. These agents can be enhanced with skills---reusable artifacts that package domain knowledge, workflows, and tool-use instructions. To date, however, little is known about how such skills are designed or how they affect agent effectiveness and behavior in practice. In this paper, we investigate these questions in smart contract security auditing, a domain in which agents have shown substantial promise. We systematically collect 83 smart contract audit skills from the wild and evaluate them on EVMBench across seven agent--model configurations. Our study examines three dimensions: (i) the design characteristics of audit skills, including their structure, knowledge representations, workflows, and tool dependencies; (ii) their effectiveness in improving vulnerability detection; and (iii) their influence on agent execution trajectories. We find that audit skills are mostly lightweight but heterogeneous in design, covering a broad yet imbalanced range of vulnerability types. Their effectiveness is determined primarily by the model rather than the agent harness: Codex/GPT-5.5 achieves the largest gains, improving detection score by 22.8% and captured award by 43.2%. We further find that skill triggering is a key bottleneck. When triggered, skills preserve a shared six-stage audit workflow while exhibiting distinct loading patterns and differential effects on agent behavior across configurations. We release our skill corpus and artifacts to support future research.
Problem

Research questions and friction points this paper is trying to address.

Smart Contract Auditing
LLM Agents
Agent Skills
Vulnerability Detection
Skill Effectiveness
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agents
Smart Contract Auditing
Agent Skills
Vulnerability Detection
Behavioral Impact
🔎 Similar Papers
No similar papers found.