Evidence · DiggingBeagle record
Bad Vibes: Comparing the Secure Coding Capabilities of Popular Coding Agents
Tenzai's controlled comparison of Cursor, Claude Code, OpenAI Codex, Replit and Devin on three matched application specifications. It reports 69 vulnerabilities across 15 generated applications and documents recurring authorization, SSRF, business-logic and missing-control failures.
- Source role
- primary disclosure
Evidence record
Tenzai's controlled comparison of Cursor, Claude Code, OpenAI Codex, Replit and Devin on three matched application specifications. It reports 69 vulnerabilities across 15 generated applications and documents recurring authorization, SSRF, business-logic and missing-control failures.
Claim-level citations (5)
- supportsFive coding agents produced 69 findings across 15 matched application builds: Tenzai tested five coding agents on three matched application specifications each and reported 69 vulnerabilities across 15 generated applications.
Article lines 19-29, experiment setup and total of 69 vulnerabilities across 15 applications
- supportsFive coding agents produced 69 findings across 15 matched application builds: The reported failures were concentrated in contextual security decisions: no exploitable SQL injection or XSS was found, all five agents introduced SSRF in the link-preview task, and the study documented recurring authorization, business-logic and missing-control failures.
Sections 'The Good', 'Authorization', 'Business logic vulnerabilities', 'Unsolved vulnerability classes' and 'The Ugly'
- supportsFive coding agents produced 69 findings across 15 matched application builds: The evidence does not support the blanket claim that generated security failures occur only when the requirement was absent from the prompt: Tenzai reports authorization failures despite detailed guidance, and SUSVIBES reports that preliminary security prompting did not eliminate the secure-coding gap.
Authorization section stating agents struggled despite clear and detailed guidance in the prompts
- supportsFive coding agents produced 69 findings across 15 matched application builds: Tenzai reports a distinct missing-controls failure mode across the 15 generated applications: none included proper CSRF protection, none added the listed security headers, and 14 of 15 login flows lacked rate limiting or account lockout. The one rate-limiting implementation was reported bypassable through X-Forwarded-For.
Section 'The Ugly': CSRF Protection, Security Headers and Login Rate Limiting; none of 15 applications had proper CSRF protection, no tested application added the listed security headers, all but one login flow lacked rate limiting or lockout, and the lone rate-limiting attempt was bypassed via X-Forwarded-For.
- contextHolding Base44 constant, Tenzai found materially different security outcomes by model: The 254-finding Base44 result cannot be used as evidence that agent-generated software became less secure than in Tenzai's earlier 69-finding study because the experiments changed the surrounding platform, model cohort and evaluation context.
Earlier experiment setup using five different coding agents with matched prompts and technology stacks