News · DiggingBeagle record

CyberChainBench grades exploit agents against historical chain state and economic impact

CyberChainBench evaluates agents against 541 historical EVM incidents across nine chains. Its best setup achieved 43.7% exploitation on the evaluated set, reproducing historical profit in forked state rather than stealing live funds.

A dated report connected to the underlying research where available.

The report

CyberChainBench moves smart-contract exploit evaluation closer to real economic conditions by replaying historical blockchain state. The benchmark was built from 541 real exploit incidents across nine EVM chains and grades agents on detection, exploitation and patching.

The best reported setup reached 43.7% exploitation on the evaluated exploitation set and reproduced $57.4 million in historical exploit profit on forked state. That is materially stronger than generating plausible Solidity snippets, because the exploit has to execute against historical conditions.

It is still a benchmark. The targets are historical and the transactions run in controlled forks, so the result is evidence of exploit-generation capability, not live autonomous theft.

Research behind this

Cite this record

DiggingBeagle. “CyberChainBench grades exploit agents against historical chain state and economic impact.” https://diggingbeagle.com/news/cyberchainbench-economic-exploit-evaluation-june-2026/

Citation guidance

Why this archive exists

The source matters after the headline fades.

DiggingBeagle is a non profit research project documenting AI security incidents, agent failures, vulnerabilities and AI-assisted operations. A case keeps its claims beside the sources that support, contest or limit them. Later updates stay visible, so a reader can see when the account changed.

We publish case reconstructions, dated reporting and analysis across records. Each has a different evidentiary role. About the project and our methodology explain how the work is reviewed.