The report
The new result is not that an AI auditor can read one more checklist. It is that reusable audit skills measurably changed vulnerability detection in a strong agent, while the same skill corpus failed to help several weaker configurations because the agents did not reliably activate it.
What changed
Gao et al. collected 83 Solidity/EVM audit skills from public skill marketplaces, GitHub collections and an academic dataset, then evaluated them on EVMBench: 40 real-world audit tasks containing 120 annotated vulnerabilities. The skills package checklists, workflow instructions, vulnerability heuristics, examples, reporting templates and tool guidance into reusable SKILL.md-style artifacts.
The strongest row was Codex with GPT-5.5. With the skills available it detected 97 of 120 annotated vulnerabilities, versus 79 without them. That is a 22.8% relative improvement in Detect Score. The paper also reports a 43.2% increase in captured award value, from USD 108,241 to USD 154,983.
The bottleneck is activation
The corpus is not automatically useful just because it exists. The paper reports that Codex/GPT-5.5 activated skills in all 40 audits, while Codex running DeepSeek-V4-Pro activated none. In a same-backend ablation using DeepSeek-V4-Pro, Claude Code, Codex and OpenCode all showed negative Detect Score changes with skills available.
That makes skill discovery and activation part of the practical security-performance boundary. A well-written audit playbook can sit beside the agent and still contribute nothing if the model never loads it.
How this fits the existing EVM agent map
DiggingBeagle already tracks several different layers of AI-assisted EVM security:
| Layer | Existing example | What it demonstrates |
|---|---|---|
| Guided discovery | EchoFuzz | LLM-guided transaction-sequence generation and fuzzing feedback |
| Exploit validation | KASS | Agentic exploit planning, Foundry execution and repair |
| Offensive strategy generation | EVAge | Multi-agent MEV strategy generation on historical forked state |
| Reusable audit guidance | This study | Domain skills can improve detection, but only when the agent activates and applies them |
The new paper therefore does not show autonomous exploitation or a live attack. It measures a defensive auditing layer that sits before exploit synthesis or transaction execution.
Important limits
Most configurations were not repeated across many independent runs, and the paper's strongest claim that backend model capability matters more than harness design relies heavily on a same-backend ablation using one weaker backend. The evidence supports a strong model-dependence and activation-bottleneck finding; it does not yet establish that harness design is generally unimportant for stronger backends.