AI security research / cases

Case files

Investigations, evidence and outcomes.

90 Cases · Page 5 of 5

Search & filter

On this page only.

Published cases

Case

GTG-50020 used prompt injection against an AI evaluation sandbox

Anthropic reports that a financially motivated actor injected malicious instructions into an AI vendor's automated evaluation sandbox, stole production API keys and then used agentic workflows against roughly 30 AI companies in about four days.

Setting: productionExploitation: observed liveEvidence: Grade B
incidentFirst seen May 21, 2026WITH AI · AGAINST AI4 claims
Case

MCP Python SDK tasks crossed client-session boundaries

The MCP Python SDK’s experimental task handlers let one connected client read or cancel another client’s work. Version 1.27.2 adds session binding for SDK-generated task IDs, but custom IDs and custom handlers still need an ownership policy.

Exploitation: not establishedEvidence: Grade BRemediation: patch available
vulnerabilityFirst seen Jun 5, 2026AGAINST AI4 claims
Case

Semantic Kernel prompt injection reached host code execution

CVE-2026-26030 allowed a model-controlled Search Plugin parameter to reach an unsafe eval-based filter path and execute code on the Semantic Kernel host under affected conditions.

Setting: research demonstrationExploitation: demonstratedEvidence: Grade BRemediation: patch available
vulnerabilityFirst seen Feb 19, 2026AGAINST AI4 claims
Case

The May RubyGems package flood is now linked to OpenAI agents, but attribution remains disputed

RubyGems confirms a large May spam-publishing campaign. Independent researchers attribute the activity to OpenAI agents; OpenAI confirms agent use of RubyGems but says it has not verified the malicious-package claims.

Setting: productionExploitation: observed liveEvidence: Grade BRemediation: mitigation available
incidentFirst seen May 11, 2026BY AI4 claims
Case

Windows-MCP HTTP modes exposed PowerShell behind a weak control plane

Before version 0.7.5, Windows-MCP's documented HTTP transports could expose an unauthenticated MCP control plane with wildcard CORS while the same server exposed a PowerShell execution tool.

Setting: research demonstrationExploitation: not establishedEvidence: Grade BRemediation: patch available
vulnerabilityFirst seen May 14, 2026AGAINST AI4 claims
Case

Google AI Mode identified tronify.rent as Tronify's official domain while the site was already flagged

Public evidence now shows that tronify.rent had both security warnings and at least one first-person loss report months before the reported September Google AI Mode interaction. That strengthens the evidence that the domain was already associated with alleged harm while narrowing AI Mode's role to a possible trust amplifier, not the origin of the scam or proof of the aggregate losses.

Unrated
incidentFirst seen 2026-09 (month precision)WITH AI7 claims