The report
CyberChainBench moves smart-contract exploit evaluation closer to real economic conditions by replaying historical blockchain state. The benchmark was built from 541 real exploit incidents across nine EVM chains and grades agents on detection, exploitation and patching.
The best reported setup reached 43.7% exploitation on the evaluated exploitation set and reproduced $57.4 million in historical exploit profit on forked state. That is materially stronger than generating plausible Solidity snippets, because the exploit has to execute against historical conditions.
It is still a benchmark. The targets are historical and the transactions run in controlled forks, so the result is evidence of exploit-generation capability, not live autonomous theft.