Evidence · DiggingBeagle record
Five Models Walk Into an App Builder. Security Gets Interesting.
Tenzai's September 2026 follow-up holds the Base44 builder and application prompts constant while comparing five model backends. It reports 254 vulnerabilities across 15 builds and materially different severity and vulnerability profiles by model.
- Published
- Sep 30, 2026
- Source role
- primary disclosure
Evidence record
Tenzai's September 2026 follow-up holds the Base44 builder and application prompts constant while comparing five model backends. It reports 254 vulnerabilities across 15 builds and materially different severity and vulnerability profiles by model.
Claim-level citations (4)
- supportsHolding Base44 constant, Tenzai found materially different security outcomes by model: Tenzai reports 254 vulnerabilities across 15 Base44 applications generated from the same three specifications using five model backends.
Experiment setup and 'Substantial variation in security outcomes' section
- supportsHolding Base44 constant, Tenzai found materially different security outcomes by model: Base1 was the only tested model with zero high or critical findings, while each of the other four models generated at least one high or critical flaw allowing unauthenticated access to sensitive data.
Severity comparison and 'Divergence in access controls' section
- supportsHolding Base44 constant, Tenzai found materially different security outcomes by model: The 254-finding Base44 result cannot be used as evidence that agent-generated software became less secure than in Tenzai's earlier 69-finding study because the experiments changed the surrounding platform, model cohort and evaluation context.
What we tested, identical builder and prompts with different model backends
- supportsHolding Base44 constant, Tenzai found materially different security outcomes by model: Within Tenzai's Base44 experiment, model changes altered vulnerability classes as well as total severity: Opus 5 and GPT-5.6 Sol had no high or critical injection findings, while DeepSeek v4 Flash and GLM 5.3 Flash each introduced high-severity NoSQL injection findings; Base1 had lower-severity injection findings. This is within-study evidence of model-dependent attack-surface differences, not a general ranking outside the tested builder, prompts and model versions.
Section 'Different models, different failure modes': Opus 5 and GPT-5.6 Sol registered zero high/critical injection findings; DeepSeek v4 Flash and GLM 5.3 Flash introduced high-severity NoSQL injection vulnerabilities; Base1 had lower-severity injection issues.