Evidence · DiggingBeagle record
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
The final ICML 2026 SUSVIBES paper evaluates 186 real-world feature-request tasks. It reports 57% functional correctness but only 11.8% secure solutions for SWE-Agent with Claude 4 Sonnet, and says preliminary security strategies did not eliminate the security gap.
- Published
- 2026-07 (month precision)
- Source role
- primary disclosure
Evidence record
The final ICML 2026 SUSVIBES paper evaluates 186 real-world feature-request tasks. It reports 57% functional correctness but only 11.8% secure solutions for SWE-Agent with Claude 4 Sonnet, and says preliminary security strategies did not eliminate the security gap.
Claim-level citations (2)
- supportsFive coding agents produced 69 findings across 15 matched application builds: The evidence does not support the blanket claim that generated security failures occur only when the requirement was absent from the prompt: Tenzai reports authorization failures despite detailed guidance, and SUSVIBES reports that preliminary security prompting did not eliminate the secure-coding gap.
Abstract, functional correctness versus secure solutions and preliminary security strategies
- supportsFive coding agents produced 69 findings across 15 matched application builds: Independent benchmarks support the distinction between functional completion and secure completion but do not reproduce Tenzai's 15-app result. BaxBench reports that end-to-end exploits succeeded against around half of functionally correct generated backends on average, while the final ICML 2026 SUSVIBES paper reports 57% functional correctness but only 11.8% secure solutions for SWE-Agent with Claude 4 Sonnet.
PMLR abstract: 186 real-world feature-request tasks across 12 coding-agent settings; SWE-Agent with Claude 4 Sonnet reached 57% functional correctness but 11.8% secure solutions, and preliminary security strategies did not eliminate the gap.