Case · DiggingBeagle record

Holding Base44 constant, Tenzai found materially different security outcomes by model

Tenzai generated the same three Base44 application specifications with five model backends and reported 254 findings across 15 builds. Severity and vulnerability classes varied materially by model, so the surrounding builder did not make model choice security-neutral.

Scope

September 2026 controlled comparison on Base44 using Base1, Opus 5, DeepSeek v4 Flash, GLM 5.3 Flash and GPT-5.6 Sol. The builder and prompts were held constant while the model backend changed. Tenzai evaluated all 15 generated applications with its own autonomous security agent.

UnratedAI role: BY AIAssessment method

At a glance

Mechanism and trust boundary

  1. 01

    Hold the app-builder workflow constant

    The same platform and application specifications are used across multiple builds.

    Boundary: builder harness / model backend

  2. 02

    Swap the underlying model

    Different model backends produce different implementations while the surrounding builder remains fixed.

    Boundary: model generation / application code

  3. 03

    Runtime testing reaches different weaknesses

    Security probing identifies different severity and vulnerability-class profiles in the deployed outputs.

    Boundary: generated application / attacker-controlled input

Timeline

  1. 2026-09 (month precision)
    occurrence

    Occurrence began

  2. Sep 30, 2026
    disclosure

    Tenzai published the Base44 model comparison

    The report compared five model backends across 15 builds generated from three matched application specifications.

Claims & evidence

CLM-BASE44-254Tenzai reports 254 vulnerabilities across 15 Base44 applications generated from the same three specifications using five model backends.supported

Basis: reported finding

Link to claim
CLM-BASE44-SEVERITYBase1 was the only tested model with zero high or critical findings, while each of the other four models generated at least one high or critical flaw allowing unauthenticated access to sensitive data.supported

Basis: reported finding

Link to claim
CLM-BASE44-NOT-TRENDThe 254-finding Base44 result cannot be used as evidence that agent-generated software became less secure than in Tenzai's earlier 69-finding study because the experiments changed the surrounding platform, model cohort and evaluation context.supported

Basis: inference

Link to claim
CLM-BASE44-CLASS-DIVERGENCEWithin Tenzai's Base44 experiment, model changes altered vulnerability classes as well as total severity: Opus 5 and GPT-5.6 Sol had no high or critical injection findings, while DeepSeek v4 Flash and GLM 5.3 Flash each introduced high-severity NoSQL injection findings; Base1 had lower-severity injection findings. This is within-study evidence of model-dependent attack-surface differences, not a general ranking outside the tested builder, prompts and model versions.supported

Basis: reported finding

  • supports
    Five Models Walk Into an App Builder. Security Gets Interesting.primary disclosure

    Section 'Different models, different failure modes': Opus 5 and GPT-5.6 Sol registered zero high/critical injection findings; DeepSeek v4 Flash and GLM 5.3 Flash introduced high-severity NoSQL injection vulnerabilities; Base1 had lower-severity injection issues.

Link to claim

Implications

The surrounding app-builder harness can standardize generation and deployment without making the security properties of generated code uniform. Model routing and model upgrades therefore belong in the application's change surface, but the security decision still has to be made on the generated runtime rather than the model name alone.

Controls and mitigations

  • Treat model selection and model upgrades as security-relevant changes that trigger the same regression gates as source-code changes.
  • Evaluate generated applications with stable authorization, injection and business-invariant tests rather than relying on general model rankings.
  • Verify the deployed runtime because a stronger model can reduce findings in one benchmark without guaranteeing a secure application.

Unknowns and contradictions

  • No independent reproduction of Tenzai's 254 findings was located in the reviewed evidence.
  • The public article does not expose a complete machine-readable corpus of all findings.
  • The model and builder versions are a September 2026 snapshot and may change quickly.
  • The experiment measures generated-app findings, not real-world exploitation frequency.

Sources and citation

Material revision history

  1. Oct 7, 2026 · Canonical change recorded · new in release · revision 92