Analysis · DiggingBeagle record
The new bottleneck in AI smart-contract auditing is whether the agent loads the skill
Smart-contract audit skills can materially improve a strong agent, but the September 2026 evidence makes activation and verification part of the real audit stack. That puts the new result beside EchoFuzz, KASS and EVAge without confusing defensive auditing with exploit automation.
- Published
- Oct 3, 2026
Overview
A smart-contract audit skill can package years of domain practice into a folder, but the new evidence says the useful question is no longer only what knowledge is available? It is also will the agent notice the skill, load it at the right moment, and turn it into evidence rather than ceremony?
That distinction matters because the EVM agent stack is no longer one capability. DiggingBeagle now has evidence for four different stages: guided vulnerability discovery, executable exploit synthesis, autonomous strategy generation on forked state, and reusable audit guidance. They should not be collapsed into a single claim that "AI can audit smart contracts."
The newest layer: portable audit guidance
The September 24 study by Gao et al. curated 83 Solidity/EVM audit skills from public marketplaces, GitHub collections and SkillsBench. The typical artifact is lightweight: a SKILL.md plus optional references, scripts or templates. Across the corpus, the common content is procedural and heuristic — checklists, vulnerability rules, validation gates, reporting structures and tool recipes — rather than a new model or analyzer.
The researchers then exposed seven agent/model configurations to the same skill corpus on EVMBench. The benchmark contains 40 audit tasks and 120 annotated vulnerabilities.
For Codex/GPT-5.5, skill availability changed the result materially: 97 vulnerabilities detected with skills versus 79 without. Captured award value increased from USD 108,241 to USD 154,983. The paper reports no computational-overhead penalty for that flagship row and describes activation as the decisive bottleneck.
Availability is not use
The most important finding is behavioral. Skills follow a staged lifecycle: discover metadata, activate the full instructions, then execute them. Only the middle step actually puts the detailed audit guidance into context.
Codex/GPT-5.5 loaded skills in all 40 audits. Under Codex/DeepSeek-V4-Pro, the paper reports zero activations in 40 audits. In the controlled weaker-backend comparison, the three tested harnesses all performed worse with the skill corpus available.
This supports a narrower and more useful conclusion than "skills work": skills can be a large multiplier when the backend model reliably discovers and applies them; a skill library that the agent does not activate is inert.
Four different EVM security layers
| Layer | DiggingBeagle record | Input → authority path | Evidence boundary |
|---|---|---|---|
| LLM-guided fuzzing | EchoFuzz | contract logic + static/runtime feedback → generated call sequences | Researcher-reported findings; public confirmation is incomplete |
| Exploit synthesis | KASS | vulnerability hypothesis → exploit plan → Foundry PoC → retry | Controlled exploit validation, not a live theft |
| MEV strategy generation | EVAge | protocol state → generated/modified MEV strategy → historical-fork validation | No production broadcast; revenue is replay-derived |
| Audit skills | September 24 study | skill metadata → activation → audit workflow → reported findings | Defensive benchmark on annotated tasks |
The rows share an EVM target but solve different problems. EchoFuzz changes where a fuzzer looks. KASS changes whether a suspected bug becomes an executable attack test. EVAge changes whether an agent can construct and adapt trading strategies. Audit skills change the instructions and reusable domain knowledge available to a general coding agent.
What the study does not prove
The paper's "model matters more than harness" conclusion deserves a qualification. The cleanest harness comparison holds DeepSeek-V4-Pro constant across three harnesses, and that weaker backend never benefits from the skills. That isolates an important failure mode, but it does not directly test whether two different harnesses can materially change skill activation under the same strong backend.
Most configurations were also run once rather than across a large set of repeated stochastic trials. The reported 22.8% and 43.2% gains are therefore concrete benchmark measurements, not an estimate of what every Solidity audit will gain in production.
Practical consequence
A team evaluating an AI auditor now needs at least three separate checks:
- Skill quality: does the artifact contain useful, scoped audit knowledge?
- Activation quality: does the agent actually load the relevant skill?
- Verification quality: does the resulting workflow produce evidence-backed findings rather than confident text?
That is a different engineering problem from simply collecting more SKILL.md files. The next useful benchmark is not another larger skill directory; it is controlled testing of activation, harness behavior and verification under the same strong model.
Research behind this
Cite this record
DiggingBeagle. “The new bottleneck in AI smart-contract auditing is whether the agent loads the skill.” Published Oct 3, 2026. https://diggingbeagle.com/articles/the-new-bottleneck-in-ai-smart-contract-auditing-is-skill-activation/