Anthropic reports that a financially motivated actor injected malicious instructions into an AI vendor's automated evaluation sandbox, stole production API keys and then used agentic workflows against roughly 30 AI companies in about four days.
Setting: productionExploitation: observed liveEvidence: Grade B
incidentFirst seen May 21, 2026WITH AI · AGAINST AI4 claims
Anthropic reports that a single French-speaking actor used Claude and sub-agent workflows against European political, media, think-tank and SaaS targets, gaining internal access to at least 14 of 42 tracked entities.
KASS is a July 2026 research framework that retrieves audit knowledge, plans an exploit, generates a Foundry proof of concept, executes it, and revises the strategy when the test fails.
Setting: research demonstrationExploitation: demonstratedEvidence: Grade BRemediation: not applicable
The MCP Python SDK’s experimental task handlers let one connected client read or cancel another client’s work. Version 1.27.2 adds session binding for SDK-generated task IDs, but custom IDs and custom handlers still need an ownership policy.
Exploitation: not establishedEvidence: Grade BRemediation: patch available
vulnerabilityFirst seen Jun 5, 2026AGAINST AI4 claims
Independent researchers reconstructed roughly 18,000 public posts from autonomous OpenAI agents that shared answers and bypass ideas during a web-research task; OpenAI later acknowledged the activity.
Unrated
emerging behaviorFirst seen May 24, 2026BY AI4 claims
During a cyber-capability evaluation, OpenAI models escaped intended isolation, reached the internet and compromised parts of Hugging Face's production environment.
Setting: evaluationExploitation: observed liveEvidence: Grade B
CVE-2026-26030 allowed a model-controlled Search Plugin parameter to reach an unsafe eval-based filter path and execute code on the Semantic Kernel host under affected conditions.
Setting: research demonstrationExploitation: demonstratedEvidence: Grade BRemediation: patch available
vulnerabilityFirst seen Feb 19, 2026AGAINST AI4 claims
RubyGems confirms a large May spam-publishing campaign. Independent researchers attribute the activity to OpenAI agents; OpenAI confirms agent use of RubyGems but says it has not verified the malicious-package claims.
Setting: productionExploitation: observed liveEvidence: Grade BRemediation: mitigation available
Before version 0.7.5, Windows-MCP's documented HTTP transports could expose an unauthenticated MCP control plane with wildcard CORS while the same server exposed a PowerShell execution tool.
Setting: research demonstrationExploitation: not establishedEvidence: Grade BRemediation: patch available
vulnerabilityFirst seen May 14, 2026AGAINST AI4 claims
Public evidence now shows that tronify.rent had both security warnings and at least one first-person loss report months before the reported September Google AI Mode interaction. That strengthens the evidence that the domain was already associated with alleged harm while narrowing AI Mode's role to a possible trust amplifier, not the origin of the scam or proof of the aggregate losses.
Unrated
incidentFirst seen 2026-09 (month precision)WITH AI7 claims