Tenzai's Base44 follow-up changes one variable that the earlier cross-agent experiment could not isolate. The builder and the three application specifications remained fixed while five model backends generated 15 builds. Tenzai reports 254 vulnerabilities across those outputs, with Base1 the only tested model to avoid high or critical findings and high or…
Inspect the ClaimsCase · DiggingBeagle record
Holding Base44 constant, Tenzai found materially different security outcomes by model
Tenzai generated the same three Base44 application specifications with five model backends and reported 254 findings across 15 builds. Severity and vulnerability classes varied materially by model, so the surrounding builder did not make model choice security-neutral.
September 2026 controlled comparison on Base44 using Base1, Opus 5, DeepSeek v4 Flash, GLM 5.3 Flash and GPT-5.6 Sol. The builder and prompts were held constant while the model backend changed. Tenzai evaluated all 15 generated applications with its own autonomous security agent.
At a glance
The surrounding app-builder harness can standardize generation and deployment without making the security properties of generated code uniform. Model routing and model upgrades therefore belong in the application's change surface, but the security decision still has to be made on the generated runtime rather than the model name alone.
Read the implicationsNo independent reproduction of Tenzai's 254 findings was located in the reviewed evidence. The public article does not expose a complete machine-readable corpus of all findings. The model and builder versions are a September 2026 snapshot and may change quickly. The experiment measures generated-app findings, not real-world exploitation frequency.
Limits and uncertaintyFull account
Tenzai's Base44 follow-up changes one variable that the earlier cross-agent experiment could not isolate. The builder and the three application specifications remained fixed while five model backends generated 15 builds. Tenzai reports 254 vulnerabilities across those outputs, with Base1 the only tested model to avoid high or critical findings and high or critical access-control failures. The other four models each produced at least one high or critical path to unauthenticated sensitive-data access, while injection profiles also differed.
This is evidence that the model inside an app-builder harness can materially change the generated attack surface. It is not evidence that security deteriorated from the earlier 69-finding Tenzai study. The two experiments used different platforms, model cohorts and contexts, and their raw counts are not a longitudinal metric. The defensible comparison is within the Base44 experiment itself: holding the surrounding builder constant did not eliminate meaningful model-dependent security variation.
Mechanism and trust boundary
- 01
Hold the app-builder workflow constant
The same platform and application specifications are used across multiple builds.
Boundary: builder harness / model backend
- 02
Swap the underlying model
Different model backends produce different implementations while the surrounding builder remains fixed.
Boundary: model generation / application code
- 03
Runtime testing reaches different weaknesses
Security probing identifies different severity and vulnerability-class profiles in the deployed outputs.
Boundary: generated application / attacker-controlled input
Timeline
- 2026-09 (month precision)occurrence
Occurrence began
- Sep 30, 2026disclosure
Tenzai published the Base44 model comparison
The report compared five model backends across 15 builds generated from three matched application specifications.
Claims & evidence
CLM-BASE44-254Tenzai reports 254 vulnerabilities across 15 Base44 applications generated from the same three specifications using five model backends.supported
Basis: reported finding
- supportsFive Models Walk Into an App Builder. Security Gets Interesting.primary disclosure
Experiment setup and 'Substantial variation in security outcomes' section
CLM-BASE44-SEVERITYBase1 was the only tested model with zero high or critical findings, while each of the other four models generated at least one high or critical flaw allowing unauthenticated access to sensitive data.supported
Basis: reported finding
- supportsFive Models Walk Into an App Builder. Security Gets Interesting.primary disclosure
Severity comparison and 'Divergence in access controls' section
CLM-BASE44-NOT-TRENDThe 254-finding Base44 result cannot be used as evidence that agent-generated software became less secure than in Tenzai's earlier 69-finding study because the experiments changed the surrounding platform, model cohort and evaluation context.supported
Basis: inference
- supportsFive Models Walk Into an App Builder. Security Gets Interesting.primary disclosure
What we tested, identical builder and prompts with different model backends
- contextBad Vibes: Comparing the Secure Coding Capabilities of Popular Coding Agentsprimary disclosure
Earlier experiment setup using five different coding agents with matched prompts and technology stacks
CLM-BASE44-CLASS-DIVERGENCEWithin Tenzai's Base44 experiment, model changes altered vulnerability classes as well as total severity: Opus 5 and GPT-5.6 Sol had no high or critical injection findings, while DeepSeek v4 Flash and GLM 5.3 Flash each introduced high-severity NoSQL injection findings; Base1 had lower-severity injection findings. This is within-study evidence of model-dependent attack-surface differences, not a general ranking outside the tested builder, prompts and model versions.supported
Basis: reported finding
- supportsFive Models Walk Into an App Builder. Security Gets Interesting.primary disclosure
Section 'Different models, different failure modes': Opus 5 and GPT-5.6 Sol registered zero high/critical injection findings; DeepSeek v4 Flash and GLM 5.3 Flash introduced high-severity NoSQL injection vulnerabilities; Base1 had lower-severity injection issues.
Implications
The surrounding app-builder harness can standardize generation and deployment without making the security properties of generated code uniform. Model routing and model upgrades therefore belong in the application's change surface, but the security decision still has to be made on the generated runtime rather than the model name alone.
Controls and mitigations
- Treat model selection and model upgrades as security-relevant changes that trigger the same regression gates as source-code changes.
- Evaluate generated applications with stable authorization, injection and business-invariant tests rather than relying on general model rankings.
- Verify the deployed runtime because a stronger model can reduce findings in one benchmark without guaranteeing a secure application.
Unknowns and contradictions
- No independent reproduction of Tenzai's 254 findings was located in the reviewed evidence.
- The public article does not expose a complete machine-readable corpus of all findings.
- The model and builder versions are a September 2026 snapshot and may change quickly.
- The experiment measures generated-app findings, not real-world exploitation frequency.
Sources and citation
Material revision history
- Oct 7, 2026 · Canonical change recorded · new in release · revision 92