Analysis · DiggingBeagle record
Three AI-security defenses that do not depend on perfect recognition
Prompt filters can miss hostile instructions, agents can misuse legitimate authority, and synthetic media can make identity checks less reliable. Three Habr pieces converge on a stronger defensive pattern: place deterministic policy, least privilege or independent verification between recognition and the next consequential action.
Overview
When a model reads hostile text, the defensive question is not whether it will always notice the trap. It is what happens after it fails. The same question applies when an agent has legitimate access to powerful tools, or when a convincing synthetic voice or face reaches an identity-verification process.
Three Habr pieces approach those problems from different directions. One examines prompt-injection containment, another applies Zero Trust principles to tool-using agents, and the third discusses AI-assisted phishing and deepfakes. They are not equivalent evidence, but together they expose a useful pattern: recognition is safer when it is followed by a separate boundary that does not depend on the same judgment being correct.
Put another boundary after recognition
| Threat | What can fail | Authority at risk | Independent boundary | Remaining problem |
|---|---|---|---|---|
| Prompt injection | A model or filter accepts attacker-controlled instructions | Private data, tool calls, outbound communication and state changes | Provenance tracking, deterministic tool policy, schemas and explicit approval | Bad policy, unmodeled data flows and approval fatigue |
| Tool-using agent | The agent behaves incorrectly or a connected component is compromised | Credentials, files, network access and connected services | Separate identity, short-lived credentials, least privilege and constrained egress | Abuse of authority that the task genuinely requires |
| Deepfake-assisted social engineering | A person or verification system accepts a false identity signal | Recovery controls, payment approval and trusted-factor replacement | Independent-channel verification and additional approval | Compromise of the independent channel or deliberate approval of the wrong action |
The controls differ because the threats differ. What connects them is the transition from information to authority. A document becomes dangerous when its contents can influence a consequential tool call; a compromised agent becomes more damaging when one identity reaches unrelated systems; a convincing impersonation matters when passing one check can establish the next trusted factor.
Prompt injection: assume some hostile instructions will get through
The Habr prompt-injection analysis argues that hostile-text filtering should not be the only security boundary when untrusted content shares a workflow with private data or consequential tools.
The article cites adaptive-attack research in which most of twelve evaluated prompt-injection or jailbreak defenses were bypassed at attack-success rates above 90 percent. That result does not establish that classifiers are useless. It does show why a security design becomes fragile when its decisive control is the model's ability to recognize an adversarial instruction.
The architectural alternatives move part of the decision outside that recognition step. Meta's Rule of Two focuses on the combination of untrusted input, sensitive access and an ability to communicate or change state. CaMeL-style designs track data provenance and apply deterministic policy before values can reach consequential tools. Schemas and allowlists can further restrict which actions or parameters are valid even when the model asks for something else.
DiggingBeagle's Semantic Kernel case demonstrates why the downstream path matters. Attacker-influenced model output reached a vulnerable deterministic filter path and then host code execution. Processing hostile text was only the beginning of the chain; framework code supplied the route from generated content to an execution sink.
The GTG-50020 record shows a similar boundary around credentials. Anthropic reported that malicious instructions reached an automated evaluation sandbox and caused disclosure of production API keys from multiple providers. A detector might have stopped a particular instruction, but an architectural question remains even when detection fails: why could an input-processing environment reach production credentials at all?
Detection still has a role. It can reject obvious attacks, add evidence and reduce the number of dangerous requests reaching later controls. The stronger arrangement is to make a missed detection insufficient by itself to authorize the final action.
A practical test for AI-assisted workflows
The three examples suggest a simple review method. First identify the input whose judgment may be wrong: retrieved text, model output, a connected component or an identity signal. Then identify the authority that becomes reachable after that judgment, such as credentials, outbound communication, a write operation, payment approval or recovery-factor replacement. Finally, ask whether another independently enforced condition stands between the two.
That condition should match the threat. Tool authorization can stop a model from inventing a valid capability. Provenance policy can stop untrusted data from crossing into a privileged sink. Separate credentials and egress controls can contain a compromised agent. An independent communication path can make one convincing impersonation insufficient for a sensitive human decision.
None of these controls requires perfect recognition, which is their main defensive advantage. They assume that a model, detector or person can sometimes be wrong and make that error only one step in a longer authorization chain.
What the evidence does not establish
The Zero Trust Habr piece is a technical synthesis, not independent production validation of its complete control set. The prompt-injection article's three-model implementation and cost figures are illustrative, and adaptive-attack results against a set of defenses do not prove that every classifier or model-side technique has zero value.
The phishing source is an interview. Its 64 percent phishing statistic, Q1 fraudulent-site counts and RUB 18 billion loss estimate describe different populations and should not be merged into one trend. It also does not provide controlled evidence that commercial deepfake detectors prevent the attack paths discussed.
The narrower conclusion is better supported. AI-backed threats increasingly place uncertain model or human judgments next to real authority. Recognition can still be useful, but the consequential transition should have its own enforceable boundary whenever the cost of a wrong judgment is high.
Research behind this
Cite this record
DiggingBeagle. “Three AI-security defenses that do not depend on perfect recognition.” https://diggingbeagle.com/articles/three-habr-defenses-move-trust-outside-the-model/