Evidence · DiggingBeagle record

Prompt Injection Cannot Be Fixed with a Filter: An Architecture for Isolating Untrusted Text

This Habr analysis focuses on the condition that makes prompt injection high impact: untrusted text shares a workflow with sensitive data and tools that can communicate or change state. Rather than treating hostile-text classification as the main security boundary, it compares structural controls including Meta's Agents Rule of Two, privileged and quarantined model separation, CaMeL-style control and data-flow separation, capability checks, deterministic schemas, allowlists and human approval before consequential actions. The article cites adaptive-attack research in which most of twelve recent prompt-injection or jailbreak defenses were bypassed at attack-success rates above 90%, which supports skepticism toward static robustness evaluations but does not prove that every detection technique is useless. It also explains an important limit of CaMeL-style designs: policies still have to be specified and maintained, and repeated approval requests can create operator fatigue. The defensible conclusion is therefore narrower than the headline: filtering should not be the sole authority boundary when untrusted content can reach private data or consequential tools. The article's three-model implementation and cost figures remain illustrative rather than independently reproduced production results.

Published
Jul 24, 2026
Source role
commentary

Evidence record

This Habr analysis focuses on the condition that makes prompt injection high impact: untrusted text shares a workflow with sensitive data and tools that can communicate or change state. Rather than treating hostile-text classification as the main security boundary, it compares structural controls including Meta's Agents Rule of Two, privileged and quarantined model separation, CaMeL-style control and data-flow separation, capability checks, deterministic schemas, allowlists and human approval before consequential actions. The article cites adaptive-attack research in which most of twelve recent prompt-injection or jailbreak defenses were bypassed at attack-success rates above 90%, which supports skepticism toward static robustness evaluations but does not prove that every detection technique is useless. It also explains an important limit of CaMeL-style designs: policies still have to be specified and maintained, and repeated approval requests can create operator fatigue. The defensible conclusion is therefore narrower than the headline: filtering should not be the sole authority boundary when untrusted content can reach private data or consequential tools. The article's three-model implementation and cost figures remain illustrative rather than independently reproduced production results.

Read the original source ↗

Cite this record

DiggingBeagle. “Prompt Injection Cannot Be Fixed with a Filter: An Architecture for Isolating Untrusted Text.” Published Jul 24, 2026. https://diggingbeagle.com/sources/prompt-injection-cannot-be-fixed-with-a-filter-an-architecture-for-isolating-unt/

Citation guidance