Analysis · DiggingBeagle record

The new bottleneck in AI smart-contract auditing is whether the agent loads the skill

Smart-contract audit skills can materially improve a strong agent, but the September 2026 evidence makes activation and verification part of the real audit stack. That puts the new result beside EchoFuzz, KASS and EVAge without confusing defensive auditing with exploit automation.

Published
Oct 3, 2026

Overview

A smart-contract audit skill can package years of domain practice into a folder, but the new evidence says the useful question is no longer only what knowledge is available? It is also will the agent notice the skill, load it at the right moment, and turn it into evidence rather than ceremony?

That distinction matters because the EVM agent stack is no longer one capability. DiggingBeagle now has evidence for four different stages: guided vulnerability discovery, executable exploit synthesis, autonomous strategy generation on forked state, and reusable audit guidance. They should not be collapsed into a single claim that "AI can audit smart contracts."

The newest layer: portable audit guidance

The September 24 study by Gao et al. curated 83 Solidity/EVM audit skills from public marketplaces, GitHub collections and SkillsBench. The typical artifact is lightweight: a SKILL.md plus optional references, scripts or templates. Across the corpus, the common content is procedural and heuristic — checklists, vulnerability rules, validation gates, reporting structures and tool recipes — rather than a new model or analyzer.

The researchers then exposed seven agent/model configurations to the same skill corpus on EVMBench. The benchmark contains 40 audit tasks and 120 annotated vulnerabilities.

For Codex/GPT-5.5, skill availability changed the result materially: 97 vulnerabilities detected with skills versus 79 without. Captured award value increased from USD 108,241 to USD 154,983. The paper reports no computational-overhead penalty for that flagship row and describes activation as the decisive bottleneck.

Availability is not use

The most important finding is behavioral. Skills follow a staged lifecycle: discover metadata, activate the full instructions, then execute them. Only the middle step actually puts the detailed audit guidance into context.

Codex/GPT-5.5 loaded skills in all 40 audits. Under Codex/DeepSeek-V4-Pro, the paper reports zero activations in 40 audits. In the controlled weaker-backend comparison, the three tested harnesses all performed worse with the skill corpus available.

This supports a narrower and more useful conclusion than "skills work": skills can be a large multiplier when the backend model reliably discovers and applies them; a skill library that the agent does not activate is inert.

Four different EVM security layers

Layer DiggingBeagle record Input → authority path Evidence boundary
LLM-guided fuzzing EchoFuzz contract logic + static/runtime feedback → generated call sequences Researcher-reported findings; public confirmation is incomplete
Exploit synthesis KASS vulnerability hypothesis → exploit plan → Foundry PoC → retry Controlled exploit validation, not a live theft
MEV strategy generation EVAge protocol state → generated/modified MEV strategy → historical-fork validation No production broadcast; revenue is replay-derived
Audit skills September 24 study skill metadata → activation → audit workflow → reported findings Defensive benchmark on annotated tasks

The rows share an EVM target but solve different problems. EchoFuzz changes where a fuzzer looks. KASS changes whether a suspected bug becomes an executable attack test. EVAge changes whether an agent can construct and adapt trading strategies. Audit skills change the instructions and reusable domain knowledge available to a general coding agent.

What the study does not prove

The paper's "model matters more than harness" conclusion deserves a qualification. The cleanest harness comparison holds DeepSeek-V4-Pro constant across three harnesses, and that weaker backend never benefits from the skills. That isolates an important failure mode, but it does not directly test whether two different harnesses can materially change skill activation under the same strong backend.

Most configurations were also run once rather than across a large set of repeated stochastic trials. The reported 22.8% and 43.2% gains are therefore concrete benchmark measurements, not an estimate of what every Solidity audit will gain in production.

Practical consequence

A team evaluating an AI auditor now needs at least three separate checks:

  1. Skill quality: does the artifact contain useful, scoped audit knowledge?
  2. Activation quality: does the agent actually load the relevant skill?
  3. Verification quality: does the resulting workflow produce evidence-backed findings rather than confident text?

That is a different engineering problem from simply collecting more SKILL.md files. The next useful benchmark is not another larger skill directory; it is controlled testing of activation, harness behavior and verification under the same strong model.

Research behind this

Cite this record

DiggingBeagle. “The new bottleneck in AI smart-contract auditing is whether the agent loads the skill.” Published Oct 3, 2026. https://diggingbeagle.com/articles/the-new-bottleneck-in-ai-smart-contract-auditing-is-skill-activation/

Citation guidance