Case · DiggingBeagle record

Five coding agents produced 69 findings across 15 matched application builds

Tenzai gave five coding agents the same three application specifications and reported 69 vulnerabilities across 15 builds. Familiar framework protections held up better than authorization, business invariants, SSRF and deployment controls.

Scope

Controlled research using Cursor, Claude Code, OpenAI Codex, Replit and Devin with their default models during December 2025. Each agent generated the same three application types from matched prompts and technology stacks. Tenzai then tested the outputs with its own security agent, so the result is a benchmark finding rather than a production compromise rate.

UnratedAI role: BY AIAssessment method

At a glance

Mechanism and trust boundary

  1. 01

    Generate an apparently functional application

    A coding agent converts the requested feature set into a working web application that can pass ordinary functional checks.

    Boundary: specification / generated implementation

  2. 02

    Reach a context-dependent security decision

    The generated code must decide object ownership, role permissions, numeric business invariants, outbound URL destinations or other policy that is not guaranteed by a framework primitive.

    Boundary: generated implementation / application policy

  3. 03

    Exercise the missing invariant

    A crafted but syntactically valid request reaches the generated endpoint and crosses the policy boundary that the application failed to enforce.

    Boundary: untrusted request / protected data or action

Timeline

  1. 2025-12 (month precision)
    Event type unspecified

    Tenzai ran the matched coding-agent experiment

    Five coding agents generated three matched applications each and Tenzai tested the 15 outputs.

  2. 2025-12 (month precision)
    occurrence

    Occurrence began

Claims & evidence

CLM-TENZAI-69Tenzai tested five coding agents on three matched application specifications each and reported 69 vulnerabilities across 15 generated applications.supported

Basis: reported finding

Link to claim
CLM-TENZAI-PATTERNThe reported failures were concentrated in contextual security decisions: no exploitable SQL injection or XSS was found, all five agents introduced SSRF in the link-preview task, and the study documented recurring authorization, business-logic and missing-control failures.supported

Basis: reported finding

Link to claim
CLM-TENZAI-PROMPT-LIMITThe evidence does not support the blanket claim that generated security failures occur only when the requirement was absent from the prompt: Tenzai reports authorization failures despite detailed guidance, and SUSVIBES reports that preliminary security prompting did not eliminate the secure-coding gap.supported

Basis: inference

Link to claim
CLM-TENZAI-MISSING-CONTROLSTenzai reports a distinct missing-controls failure mode across the 15 generated applications: none included proper CSRF protection, none added the listed security headers, and 14 of 15 login flows lacked rate limiting or account lockout. The one rate-limiting implementation was reported bypassable through X-Forwarded-For.supported

Basis: reported finding

  • supports
    Bad Vibes: Comparing the Secure Coding Capabilities of Popular Coding Agentsprimary disclosure

    Section 'The Ugly': CSRF Protection, Security Headers and Login Rate Limiting; none of 15 applications had proper CSRF protection, no tested application added the listed security headers, all but one login flow lacked rate limiting or lockout, and the lone rate-limiting attempt was bypassed via X-Forwarded-For.

Link to claim
CLM-TENZAI-EXTERNAL-BENCHMARKSIndependent benchmarks support the distinction between functional completion and secure completion but do not reproduce Tenzai's 15-app result. BaxBench reports that end-to-end exploits succeeded against around half of functionally correct generated backends on average, while the final ICML 2026 SUSVIBES paper reports 57% functional correctness but only 11.8% secure solutions for SWE-Agent with Claude 4 Sonnet.supported

Basis: reported finding

Link to claim

Implications

The benchmark does not show that coding agents always produce insecure software, and its exact agent ranking will age quickly. It does show that a plausible, functional implementation can still fail where security depends on contextual policy. Prompting remains useful for expressing intent, but the enforceable boundary has to exist in tests, runtime policy and review outside the model.

Controls and mitigations

  • Encode authorization, ownership and business invariants as executable tests that fail independently of the agent's explanation or confidence.
  • Treat server-side URL fetching as an explicit network boundary and test private-address, redirect and resolver cases rather than relying on generic URL validation.
  • Make CSRF protection, authentication throttling and required response headers deployment gates, then run independent dynamic security testing against the generated runtime.

Unknowns and contradictions

  • Tenzai supplied both the application-generation benchmark and the security-testing agent; the complete 69-finding set has not been independently reproduced in the reviewed evidence.
  • The public article does not expose a complete machine-readable ledger of all 69 findings.
  • The tested agent and default-model versions were those available in December 2025 and may not represent later versions.
  • Tenzai's blog index lists January 13, 2026 for the article, while the current article page displays June 25, 2026.

Sources and citation

Material revision history

  1. Oct 7, 2026 · Canonical change recorded · new in release · revision 92