Tenzai's experiment produced a more specific pattern than the Telegram summary suggests. Across the 15 generated applications, the researchers reported no exploitable SQL injection or XSS. The failures appeared when security depended on application context rather than a framework's standard safe primitive. All five agents introduced SSRF in a link-preview…
Inspect the ClaimsCase · DiggingBeagle record
Five coding agents produced 69 findings across 15 matched application builds
Tenzai gave five coding agents the same three application specifications and reported 69 vulnerabilities across 15 builds. Familiar framework protections held up better than authorization, business invariants, SSRF and deployment controls.
Controlled research using Cursor, Claude Code, OpenAI Codex, Replit and Devin with their default models during December 2025. Each agent generated the same three application types from matched prompts and technology stacks. Tenzai then tested the outputs with its own security agent, so the result is a benchmark finding rather than a production compromise rate.
At a glance
The benchmark does not show that coding agents always produce insecure software, and its exact agent ranking will age quickly. It does show that a plausible, functional implementation can still fail where security depends on contextual policy. Prompting remains useful for expressing intent, but the enforceable boundary has to exist in tests, runtime policy…
Read the implicationsTenzai supplied both the application-generation benchmark and the security-testing agent; the complete 69-finding set has not been independently reproduced in the reviewed evidence. The public article does not expose a complete machine-readable ledger of all 69 findings. The tested agent and default-model versions were those available in December 2025 and…
Limits and uncertaintyFull account
Tenzai's experiment produced a more specific pattern than the Telegram summary suggests. Across the 15 generated applications, the researchers reported no exploitable SQL injection or XSS. The failures appeared when security depended on application context rather than a framework's standard safe primitive. All five agents introduced SSRF in a link-preview feature that accepted user-controlled URLs, four of five allowed negative order quantities, three of five allowed negative product prices, and the study documents authorization mistakes that survived detailed prompt guidance.
The missing-control layer was broader. None of the 15 applications included proper CSRF protection, no application added the security headers Tenzai checked, and all but one login flow lacked rate limiting or lockout. The one rate-limiting implementation was reported bypassable through X-Forwarded-For. This means the useful conclusion is not that agents only omit whatever the user forgot to request. Some business invariants were absent from the specification, but authorization failures also occurred despite explicit requirements. BaxBench and SUSVIBES independently reinforce the larger point that functional completion and secure completion remain different objectives.
Mechanism and trust boundary
- 01
Generate an apparently functional application
A coding agent converts the requested feature set into a working web application that can pass ordinary functional checks.
Boundary: specification / generated implementation
- 02
Reach a context-dependent security decision
The generated code must decide object ownership, role permissions, numeric business invariants, outbound URL destinations or other policy that is not guaranteed by a framework primitive.
Boundary: generated implementation / application policy
- 03
Exercise the missing invariant
A crafted but syntactically valid request reaches the generated endpoint and crosses the policy boundary that the application failed to enforce.
Boundary: untrusted request / protected data or action
Timeline
- 2025-12 (month precision)Event type unspecified
Tenzai ran the matched coding-agent experiment
Five coding agents generated three matched applications each and Tenzai tested the 15 outputs.
- 2025-12 (month precision)occurrence
Occurrence began
Claims & evidence
CLM-TENZAI-69Tenzai tested five coding agents on three matched application specifications each and reported 69 vulnerabilities across 15 generated applications.supported
Basis: reported finding
- supportsBad Vibes: Comparing the Secure Coding Capabilities of Popular Coding Agentsprimary disclosure
Article lines 19-29, experiment setup and total of 69 vulnerabilities across 15 applications
CLM-TENZAI-PATTERNThe reported failures were concentrated in contextual security decisions: no exploitable SQL injection or XSS was found, all five agents introduced SSRF in the link-preview task, and the study documented recurring authorization, business-logic and missing-control failures.supported
Basis: reported finding
- supportsBad Vibes: Comparing the Secure Coding Capabilities of Popular Coding Agentsprimary disclosure
Sections 'The Good', 'Authorization', 'Business logic vulnerabilities', 'Unsolved vulnerability classes' and 'The Ugly'
CLM-TENZAI-PROMPT-LIMITThe evidence does not support the blanket claim that generated security failures occur only when the requirement was absent from the prompt: Tenzai reports authorization failures despite detailed guidance, and SUSVIBES reports that preliminary security prompting did not eliminate the secure-coding gap.supported
Basis: inference
- supportsBad Vibes: Comparing the Secure Coding Capabilities of Popular Coding Agentsprimary disclosure
Authorization section stating agents struggled despite clear and detailed guidance in the prompts
- supportsIs Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasksprimary disclosure
Abstract, functional correctness versus secure solutions and preliminary security strategies
- contextBaxBench: Can LLMs Generate Secure and Correct Backends?primary disclosure
Key Takeaways and prompt-setting descriptions
CLM-TENZAI-MISSING-CONTROLSTenzai reports a distinct missing-controls failure mode across the 15 generated applications: none included proper CSRF protection, none added the listed security headers, and 14 of 15 login flows lacked rate limiting or account lockout. The one rate-limiting implementation was reported bypassable through X-Forwarded-For.supported
Basis: reported finding
- supportsBad Vibes: Comparing the Secure Coding Capabilities of Popular Coding Agentsprimary disclosure
Section 'The Ugly': CSRF Protection, Security Headers and Login Rate Limiting; none of 15 applications had proper CSRF protection, no tested application added the listed security headers, all but one login flow lacked rate limiting or lockout, and the lone rate-limiting attempt was bypassed via X-Forwarded-For.
CLM-TENZAI-EXTERNAL-BENCHMARKSIndependent benchmarks support the distinction between functional completion and secure completion but do not reproduce Tenzai's 15-app result. BaxBench reports that end-to-end exploits succeeded against around half of functionally correct generated backends on average, while the final ICML 2026 SUSVIBES paper reports 57% functional correctness but only 11.8% secure solutions for SWE-Agent with Claude 4 Sonnet.supported
Basis: reported finding
- supportsBaxBench: Can LLMs Generate Secure and Correct Backends?primary disclosure
PMLR abstract and key findings: 392 backend-generation tasks evaluated with functional tests and end-to-end exploits; security exploits succeeded against around half of the functionally correct programs generated by each LLM on average.
- supportsIs Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasksprimary disclosure
PMLR abstract: 186 real-world feature-request tasks across 12 coding-agent settings; SWE-Agent with Claude 4 Sonnet reached 57% functional correctness but 11.8% secure solutions, and preliminary security strategies did not eliminate the gap.
Implications
The benchmark does not show that coding agents always produce insecure software, and its exact agent ranking will age quickly. It does show that a plausible, functional implementation can still fail where security depends on contextual policy. Prompting remains useful for expressing intent, but the enforceable boundary has to exist in tests, runtime policy and review outside the model.
Controls and mitigations
- Encode authorization, ownership and business invariants as executable tests that fail independently of the agent's explanation or confidence.
- Treat server-side URL fetching as an explicit network boundary and test private-address, redirect and resolver cases rather than relying on generic URL validation.
- Make CSRF protection, authentication throttling and required response headers deployment gates, then run independent dynamic security testing against the generated runtime.
Unknowns and contradictions
- Tenzai supplied both the application-generation benchmark and the security-testing agent; the complete 69-finding set has not been independently reproduced in the reviewed evidence.
- The public article does not expose a complete machine-readable ledger of all 69 findings.
- The tested agent and default-model versions were those available in December 2025 and may not represent later versions.
- Tenzai's blog index lists January 13, 2026 for the article, while the current article page displays June 25, 2026.
Sources and citation
Material revision history
- Oct 7, 2026 · Canonical change recorded · new in release · revision 92