Evidence · DiggingBeagle record
Prompt Injection Cannot Be Fixed with a Filter: An Architecture for Isolating Untrusted Text
This Habr analysis focuses on the condition that makes prompt injection high impact: untrusted text shares a workflow with sensitive data and tools that can communicate or change state. Rather than treating hostile-text classification as the main security boundary, it compares structural controls including Meta's Agents Rule of Two, privileged and quarantined model separation, CaMeL-style control and data-flow separation, capability checks, deterministic schemas, allowlists and human approval before consequential actions. The article cites adaptive-attack research in which most of twelve recent prompt-injection or jailbreak defenses were bypassed at attack-success rates above 90%, which supports skepticism toward static robustness evaluations but does not prove that every detection technique is useless. It also explains an important limit of CaMeL-style designs: policies still have to be specified and maintained, and repeated approval requests can create operator fatigue. The defensible conclusion is therefore narrower than the headline: filtering should not be the sole authority boundary when untrusted content can reach private data or consequential tools. The article's three-model implementation and cost figures remain illustrative rather than independently reproduced production results.
- Published
- Jul 24, 2026
- Source role
- commentary
Evidence record
This Habr analysis focuses on the condition that makes prompt injection high impact: untrusted text shares a workflow with sensitive data and tools that can communicate or change state. Rather than treating hostile-text classification as the main security boundary, it compares structural controls including Meta's Agents Rule of Two, privileged and quarantined model separation, CaMeL-style control and data-flow separation, capability checks, deterministic schemas, allowlists and human approval before consequential actions. The article cites adaptive-attack research in which most of twelve recent prompt-injection or jailbreak defenses were bypassed at attack-success rates above 90%, which supports skepticism toward static robustness evaluations but does not prove that every detection technique is useless. It also explains an important limit of CaMeL-style designs: policies still have to be specified and maintained, and repeated approval requests can create operator fatigue. The defensible conclusion is therefore narrower than the headline: filtering should not be the sole authority boundary when untrusted content can reach private data or consequential tools. The article's three-model implementation and cost figures remain illustrative rather than independently reproduced production results.
Cite this record
DiggingBeagle. “Prompt Injection Cannot Be Fixed with a Filter: An Architecture for Isolating Untrusted Text.” Published Jul 24, 2026. https://diggingbeagle.com/sources/prompt-injection-cannot-be-fixed-with-a-filter-an-architecture-for-isolating-unt/