Analysis · DiggingBeagle record

A working vibe-coded app tells you almost nothing about its security boundary

Across a Lovable RLS disclosure, a five-agent benchmark and a fixed-builder model comparison, the same separation keeps appearing: functional software can still fail at backend authorization, business rules, network policy and deployment controls. The studies use different methods and their raw counts should not be merged, but together they show why security has to be tested at the runtime boundary rather than inferred from a polished interface or a model name.

Published
Oct 7, 2026

Overview

Vibe-coded software can look complete, pass ordinary product checks and still lack a reliable security boundary. The evidence in this cluster comes from three different kinds of observation: a real-world scan of Lovable projects, controlled builds across five coding agents, and a fixed-builder comparison across five model backends. The measurements are not interchangeable, but they converge on one practical point: functional completion does not demonstrate that authorization, business rules, network access and deployment controls are correct.

The failure is also more specific than "AI writes insecure code." In the studies reviewed here, some familiar vulnerability classes were handled relatively well, while failures clustered around decisions that depend on application context. A framework can parameterize a SQL query automatically, but it cannot infer who should be allowed to delete an order, whether a negative quantity makes sense, which URL destinations a preview service may fetch, or whether a public browser client should be able to read a row from the backend.

Four measurements, four different questions

Evidence Environment Reported result What it actually supports
Lovable RLS disclosure 1,645 publicly discoverable Lovable projects, scanned in March 2025 303 inadequately protected endpoints across 170 projects Deployed projects in the sampled public set had backend authorization failures; this is not a census of all Lovable apps and not a current prevalence estimate
Escape vibe-app study More than 5,600 public apps across several platforms More than 2,000 vulnerabilities, 400+ exposed secrets and 175 PII exposures The broader public ecosystem contained repeated exposed-backend and secret-management failures; sampling, timing and platform imbalance limit prevalence comparisons
Tenzai cross-agent benchmark Five coding agents, three matched applications each 69 vulnerabilities across 15 generated apps Different agents repeatedly missed authorization, business invariants, SSRF defenses and deployment controls under a controlled setup
Tenzai Base44 model comparison One builder and the same three specifications, with five model backends 254 vulnerabilities across 15 builds Swapping only the underlying model changed severity and vulnerability classes inside that experiment

The temptation is to read 69 vulnerabilities in one study and 254 in another as a worsening trend. The evidence does not support that conclusion. The surrounding platform, model cohort and evaluation context changed, so the defensible comparison is within each experiment, not across the raw totals.

The first boundary can sit behind a perfectly normal interface

The Lovable disclosure shows why the generated frontend is the wrong place to stop a security review.

Lovable applications in the examined setup were primarily client-driven and used Supabase for backend storage and authentication. A Supabase anon credential appearing in browser code was not, by itself, the vulnerability. The real authority boundary was Row Level Security in the database. If that policy allowed a request that the interface never intended to expose, changing the browser-visible request could reach data outside the intended UI path.

Matt Palmer reported that a March 21, 2025 scan found 303 inadequately protected endpoints across 170 of 1,645 analyzed projects. The scan only inspected homepages and modified observed requests, so it was not an exhaustive assessment of every surface. In a later Linkable follow-up, Palmer reported a more concrete integrity failure: after changing the request context, an unauthenticated request could insert a record with payment_status set to paid.

That is the useful mechanism. The interface can be polished, the normal workflow can behave correctly, and the backend can still accept a request that the interface would never present.

The CVE record also matters because responsibility is disputed. The public record describes insufficient RLS that could permit unauthenticated reads or writes, while recording Lovable's position that customers are responsible for protecting their own application data. That dispute changes attribution, but it does not remove the engineering question: which server-side rule rejects a request once the frontend is bypassed?

The coding-agent failures were concentrated in contextual policy

Tenzai's cross-agent experiment gives a second view of the same gap. Cursor, Claude Code, OpenAI Codex, Replit and Devin each built the same three application types with matched prompts and technology stacks. Tenzai then tested the 15 generated applications and reported 69 vulnerabilities.

The interesting result was not that every familiar bug class collapsed. Tenzai reported no exploitable SQL injection or XSS across the tested applications. The agents generally used framework protections that already encoded a clear safe pattern.

The failures appeared when the application had to understand context.

Authorization was one example. Tenzai documented APIs that checked ownership only for some roles or only when a user was already authenticated, leaving other request paths improperly authorized. Business invariants were another: four of five agents allowed negative order quantities in one task, and three of five allowed negative product prices. All five agents introduced SSRF in the link-preview task, where the application had to decide which user-controlled URLs were safe for the server to fetch.

This also weakens the simple explanation that insecure output appears only when the user forgets to ask for security. Tenzai says authorization mistakes survived clear and detailed prompt guidance. Prompt detail matters, but it is not a substitute for an independent enforcement mechanism.

Missing controls are a different failure from buggy controls

A second layer in the same benchmark concerned protections that were mostly absent rather than incorrectly implemented.

None of the 15 generated applications included proper CSRF protection. None added the security headers Tenzai checked. Fourteen of the 15 login flows lacked rate limiting or account lockout, and the single rate-limiting implementation was reported bypassable through X-Forwarded-For.

That distinction matters operationally. A code review focused on the lines the agent wrote can miss controls that never appeared in the diff at all. Security therefore needs an expected-control inventory as well as vulnerability scanning: authentication throttling, CSRF protection, response headers, authorization tests, network egress rules and similar gates must be checked even when no generated line looks suspicious.

Independent benchmarks show the same correctness gap

The broader benchmark literature supports the separation between "works" and "safe" without reproducing Tenzai's exact experiment.

BaxBench evaluates 392 security-critical backend coding tasks with functional tests and expert-designed exploits. Its project reports that 62% of solutions from the best tested model were either incorrect or vulnerable, and that around half of functionally correct solutions were insecure on average.

SUSVIBES uses 186 feature-request tasks derived from real open-source changes that introduced vulnerabilities. In the published ICML 2026 results, SWE-Agent with Claude 4 Sonnet reached 57% functional correctness, while only 11.8% of solutions were secure. The paper also reports that preliminary security strategies, including adding vulnerability hints to the task, did not eliminate the gap.

These studies use different tasks, agents and evaluation methods, so their percentages should not be combined into a single industry-wide rate. Their shared value is structural: correctness metrics alone can reward software that satisfies the requested feature while leaving an exploitable security boundary behind it.

Model choice changes the attack surface, but it does not close it

Tenzai's September 2026 Base44 study isolates a different variable. The builder and the three application specifications stayed constant while the underlying model changed across Base1, Opus 5, DeepSeek v4 Flash, GLM 5.3 Flash and GPT-5.6 Sol.

Tenzai reported 254 vulnerabilities across the 15 builds. Base1 was the only tested model with no high or critical findings, while each of the other four models produced at least one high or critical flaw allowing unauthenticated access to sensitive data.

The vulnerability classes also moved. Opus 5 and GPT-5.6 Sol had no high or critical injection findings in the experiment, while DeepSeek v4 Flash and GLM 5.3 Flash produced high-severity NoSQL injection findings. Base1 had lower-severity injection issues.

This is useful evidence that a model swap is a security-relevant change to an app-builder pipeline. It is not a durable league table. The result comes from one builder, three specifications, one testing system and the model versions available at that time. Tenzai also tested its own generated applications with its own security agent, and the reviewed evidence does not include an independent reproduction of the complete finding set.

The practical consequence is narrower and stronger: changing the model can change the runtime attack surface, so a model upgrade or routing change should trigger the same regression gates as a source-code change.

The security boundary is spread across four layers

The studies become more useful when the controls are mapped to the layer that actually holds authority.

Boundary Typical failure in the evidence What should be tested independently
Backend authorization Missing or overly broad RLS, broken object ownership, role checks skipped on some paths Anonymous, authenticated and cross-user requests against the backend or API, not only through the UI
Business invariants Negative quantities, negative prices, destructive actions allowed in invalid states Executable tests for ownership, state transitions, numeric bounds and role-specific actions
Network authority SSRF through server-side URL fetching Private-address blocking, redirects, resolver behavior and explicit destination policy
Deployment controls Missing CSRF protection, headers, login throttling or lockout Release gates that verify the expected controls exist and work at runtime

This is where vibe coding changes the review problem. Generation compresses implementation time, but it does not compress the number of trust boundaries. A person can move from idea to deployed application before anyone has explicitly asked which component owns authorization, which values are legal, which network destinations are reachable, or which controls the deployment platform is expected to add.

Field evidence suggests the problem is not confined to benchmarks

Escape's broader study gives the controlled findings a real-world counterpart. Its researchers analyzed more than 5,600 publicly available vibe-coded applications across several platforms and reported more than 2,000 vulnerabilities, more than 400 exposed secrets and 175 PII exposures. The dataset was heavily weighted toward Lovable, with roughly 4,000+ Lovable applications in the initial platform coverage.

The methodology also documents limits that matter. Collection was a one-time snapshot, public and easily discoverable applications could be overrepresented, and platform coverage was uneven. Escape therefore provides evidence that exposed backends and secrets recur in deployed public apps, but it does not provide a clean platform-to-platform prevalence ranking.

The strongest overlap with the Lovable case is architectural. Escape describes finding public frontend artifacts that exposed routes, tokens or backend structure, then testing whether the server-side authorization actually held. Again, the browser-visible credential or route is not necessarily the bug. The decisive question is whether the backend accepts an unauthorized operation.

What this evidence does not support

The 170 of 1,645 Lovable figure should not be presented as the current probability that a Lovable application is vulnerable. It was a March 2025 sample of discoverable public projects, not a census, and the platform and user behavior have changed since then.

The 69 findings from the cross-agent experiment and the 254 findings from the Base44 experiment should not be plotted as a time series. They came from different test designs.

The model comparison does not establish that one named model is generally secure or insecure. It shows that, inside one controlled builder experiment, model choice changed the resulting security profile.

None of the benchmark counts is a production compromise rate. They describe findings in generated applications under controlled testing. The Lovable case is closer to deployed exposure, but even there the reviewed evidence does not establish a victim count or a comprehensive exploitation count.

The release question is simple: what rejects the bad request?

The useful test for a vibe-coded application is not whether the agent can explain its security choices, whether the interface hides a dangerous action, or whether the selected model scored well in another benchmark. It is whether an independent mechanism rejects the unsafe behavior at the boundary that owns the decision.

For backend data, that means direct authorization tests and row-ownership rules. For business logic, it means executable invariants. For outbound requests, it means explicit network policy. For deployment controls, it means release gates. For model changes, it means rerunning the same security regression suite against the generated runtime.

Model quality can improve the starting point, and platform defaults can remove entire classes of mistakes, but neither one proves that the deployed application is secure. The final authority still sits in the runtime controls that accept or reject the request.

Research behind this