Analysis · DiggingBeagle record

Three AI-security defenses that do not depend on perfect recognition

Prompt filters can miss hostile instructions, agents can misuse legitimate authority, and synthetic media can make identity checks less reliable. Three Habr pieces converge on a stronger defensive pattern: place deterministic policy, least privilege or independent verification between recognition and the next consequential action.

Overview

When a model reads hostile text, the defensive question is not whether it will always notice the trap. It is what happens after it fails. The same question applies when an agent has legitimate access to powerful tools, or when a convincing synthetic voice or face reaches an identity-verification process.

Three Habr pieces approach those problems from different directions. One examines prompt-injection containment, another applies Zero Trust principles to tool-using agents, and the third discusses AI-assisted phishing and deepfakes. They are not equivalent evidence, but together they expose a useful pattern: recognition is safer when it is followed by a separate boundary that does not depend on the same judgment being correct.

Put another boundary after recognition

Threat What can fail Authority at risk Independent boundary Remaining problem
Prompt injection A model or filter accepts attacker-controlled instructions Private data, tool calls, outbound communication and state changes Provenance tracking, deterministic tool policy, schemas and explicit approval Bad policy, unmodeled data flows and approval fatigue
Tool-using agent The agent behaves incorrectly or a connected component is compromised Credentials, files, network access and connected services Separate identity, short-lived credentials, least privilege and constrained egress Abuse of authority that the task genuinely requires
Deepfake-assisted social engineering A person or verification system accepts a false identity signal Recovery controls, payment approval and trusted-factor replacement Independent-channel verification and additional approval Compromise of the independent channel or deliberate approval of the wrong action

The controls differ because the threats differ. What connects them is the transition from information to authority. A document becomes dangerous when its contents can influence a consequential tool call; a compromised agent becomes more damaging when one identity reaches unrelated systems; a convincing impersonation matters when passing one check can establish the next trusted factor.

Prompt injection: assume some hostile instructions will get through

The Habr prompt-injection analysis argues that hostile-text filtering should not be the only security boundary when untrusted content shares a workflow with private data or consequential tools.

The article cites adaptive-attack research in which most of twelve evaluated prompt-injection or jailbreak defenses were bypassed at attack-success rates above 90 percent. That result does not establish that classifiers are useless. It does show why a security design becomes fragile when its decisive control is the model's ability to recognize an adversarial instruction.

The architectural alternatives move part of the decision outside that recognition step. Meta's Rule of Two focuses on the combination of untrusted input, sensitive access and an ability to communicate or change state. CaMeL-style designs track data provenance and apply deterministic policy before values can reach consequential tools. Schemas and allowlists can further restrict which actions or parameters are valid even when the model asks for something else.

DiggingBeagle's Semantic Kernel case demonstrates why the downstream path matters. Attacker-influenced model output reached a vulnerable deterministic filter path and then host code execution. Processing hostile text was only the beginning of the chain; framework code supplied the route from generated content to an execution sink.

The GTG-50020 record shows a similar boundary around credentials. Anthropic reported that malicious instructions reached an automated evaluation sandbox and caused disclosure of production API keys from multiple providers. A detector might have stopped a particular instruction, but an architectural question remains even when detection fails: why could an input-processing environment reach production credentials at all?

Detection still has a role. It can reject obvious attacks, add evidence and reduce the number of dangerous requests reaching later controls. The stronger arrangement is to make a missed detection insufficient by itself to authorize the final action.

Agent security: constrain authority before asking whether the model is trustworthy

The Habr Zero Trust overview treats an AI agent as a security principal rather than as a trusted extension of the person who launched it. That framing matters because a capable agent can be perfectly obedient and still carry too much authority for the task it is performing.

Separate identity makes permissions attributable to the agent instead of inheriting a broad human or service account. Short-lived credentials reduce the useful lifetime of a leak. Deny-by-default tool access limits unrelated capabilities, while read-only permissions and constrained network egress reduce the number of ways a compromised workflow can change state or move data. Parameter validation and complete tool-call traces provide another boundary around what an allowed tool can actually do and what operators can reconstruct afterward.

These controls do not require the model to determine whether a prompt, plugin or retrieved document is malicious. They limit the consequences when that determination is wrong.

This is also the problem examined in DiggingBeagle's transitive-authority analysis and the extension study. Plugins, MCP services and workflow components often receive secrets or network access for legitimate reasons. Once the host grants that authority, the security boundary cannot depend entirely on the component's README, prompt or claimed purpose.

Least privilege has a precise limitation: it reduces reachable authority but cannot prevent misuse of permissions that remain necessary. A support agent that legitimately needs to issue refunds may still issue the wrong refund. Zero Trust changes the blast radius and the prerequisites for abuse; it does not make model behavior correct.

Deepfakes: separate the identity claim from the action it authorizes

The Habr phishing interview discusses personalized phishing, synthetic voice and video, multi-stage social engineering and phishing-as-a-service. Among its recommended controls are independent communication channels and second-person approval for critical operations.

Those controls do something different from a deepfake detector. A detector tries to decide whether the media is genuine. An independent callback or separately established contact path changes the attacker's prerequisite: controlling the original call, video session or message is no longer enough. A second approver similarly prevents one successful impersonation from automatically becoming the final authorization.

The KZ-CERT case provides a concrete example of why the second transition matters. The reported sequence used DeepFake-assisted video verification and was followed by replacement of the trusted phone number. The important security consequence was therefore not limited to accepting synthetic media. Passing the verification step could reportedly establish a new trusted factor for later account control.

Authenticator-based two-factor authentication can protect an account login, but it does not stop a person from approving a fraudulent payment after being persuaded through another channel. Deepfake detection may contribute another signal, but the Habr interview does not provide controlled effectiveness or false-positive measurements for those products. Independent authorization changes the workflow even when recognition remains uncertain.

For the fuller recovery sequence, see When a deepfake can become an account-recovery pivot.

A practical test for AI-assisted workflows

The three examples suggest a simple review method. First identify the input whose judgment may be wrong: retrieved text, model output, a connected component or an identity signal. Then identify the authority that becomes reachable after that judgment, such as credentials, outbound communication, a write operation, payment approval or recovery-factor replacement. Finally, ask whether another independently enforced condition stands between the two.

That condition should match the threat. Tool authorization can stop a model from inventing a valid capability. Provenance policy can stop untrusted data from crossing into a privileged sink. Separate credentials and egress controls can contain a compromised agent. An independent communication path can make one convincing impersonation insufficient for a sensitive human decision.

None of these controls requires perfect recognition, which is their main defensive advantage. They assume that a model, detector or person can sometimes be wrong and make that error only one step in a longer authorization chain.

What the evidence does not establish

The Zero Trust Habr piece is a technical synthesis, not independent production validation of its complete control set. The prompt-injection article's three-model implementation and cost figures are illustrative, and adaptive-attack results against a set of defenses do not prove that every classifier or model-side technique has zero value.

The phishing source is an interview. Its 64 percent phishing statistic, Q1 fraudulent-site counts and RUB 18 billion loss estimate describe different populations and should not be merged into one trend. It also does not provide controlled evidence that commercial deepfake detectors prevent the attack paths discussed.

The narrower conclusion is better supported. AI-backed threats increasingly place uncertain model or human judgments next to real authority. Recognition can still be useful, but the consequential transition should have its own enforceable boundary whenever the cost of a wrong judgment is high.

Research behind this

Cite this record

DiggingBeagle. “Three AI-security defenses that do not depend on perfect recognition.” https://diggingbeagle.com/articles/three-habr-defenses-move-trust-outside-the-model/

Citation guidance