Evidence · DiggingBeagle record

Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks

The final ICML 2026 SUSVIBES paper evaluates 186 real-world feature-request tasks. It reports 57% functional correctness but only 11.8% secure solutions for SWE-Agent with Claude 4 Sonnet, and says preliminary security strategies did not eliminate the security gap.

Published
2026-07 (month precision)
Source role
primary disclosure

Evidence record

The final ICML 2026 SUSVIBES paper evaluates 186 real-world feature-request tasks. It reports 57% functional correctness but only 11.8% secure solutions for SWE-Agent with Claude 4 Sonnet, and says preliminary security strategies did not eliminate the security gap.

Read the original source ↗

Claim-level citations (2)