News · DiggingBeagle record

Audit skills improved EVM vulnerability detection — when the agent actually loaded them

A September 24 study of 83 smart-contract audit skills found a large detection gain for Codex/GPT-5.5, but also showed that skill activation can be the bottleneck: several weaker configurations rarely or never loaded the available guidance.

Published
Sep 24, 2026

The report

The new result is not that an AI auditor can read one more checklist. It is that reusable audit skills measurably changed vulnerability detection in a strong agent, while the same skill corpus failed to help several weaker configurations because the agents did not reliably activate it.

What changed

Gao et al. collected 83 Solidity/EVM audit skills from public skill marketplaces, GitHub collections and an academic dataset, then evaluated them on EVMBench: 40 real-world audit tasks containing 120 annotated vulnerabilities. The skills package checklists, workflow instructions, vulnerability heuristics, examples, reporting templates and tool guidance into reusable SKILL.md-style artifacts.

The strongest row was Codex with GPT-5.5. With the skills available it detected 97 of 120 annotated vulnerabilities, versus 79 without them. That is a 22.8% relative improvement in Detect Score. The paper also reports a 43.2% increase in captured award value, from USD 108,241 to USD 154,983.

The bottleneck is activation

The corpus is not automatically useful just because it exists. The paper reports that Codex/GPT-5.5 activated skills in all 40 audits, while Codex running DeepSeek-V4-Pro activated none. In a same-backend ablation using DeepSeek-V4-Pro, Claude Code, Codex and OpenCode all showed negative Detect Score changes with skills available.

That makes skill discovery and activation part of the practical security-performance boundary. A well-written audit playbook can sit beside the agent and still contribute nothing if the model never loads it.

How this fits the existing EVM agent map

DiggingBeagle already tracks several different layers of AI-assisted EVM security:

Layer Existing example What it demonstrates
Guided discovery EchoFuzz LLM-guided transaction-sequence generation and fuzzing feedback
Exploit validation KASS Agentic exploit planning, Foundry execution and repair
Offensive strategy generation EVAge Multi-agent MEV strategy generation on historical forked state
Reusable audit guidance This study Domain skills can improve detection, but only when the agent activates and applies them

The new paper therefore does not show autonomous exploitation or a live attack. It measures a defensive auditing layer that sits before exploit synthesis or transaction execution.

Important limits

Most configurations were not repeated across many independent runs, and the paper's strongest claim that backend model capability matters more than harness design relies heavily on a same-backend ablation using one weaker backend. The evidence supports a strong model-dependence and activation-bottleneck finding; it does not yet establish that harness design is generally unimportant for stronger backends.

Research behind this

Cite this record

DiggingBeagle. “Audit skills improved EVM vulnerability detection — when the agent actually loaded them.” Published Sep 24, 2026. https://diggingbeagle.com/news/agent-skills-improve-evm-auditing-when-the-agent-actually-loads-them/

Citation guidance