AI security research / cases

Case files

Investigations, evidence and outcomes.

90 Cases · Page 4 of 5

Search & filter

On this page only.

Published cases

Case

Hijacked AI coding-assistant session became the entry path for Shai-Hulud across about 100 repositories

Mandiant investigated an intrusion at an unnamed SaaS provider where an attacker hijacked a developer's active AI coding-assistant session. After an attacker-poisoned package was recommended and accepted, the chain led to an infostealer, stolen GitHub OAuth tokens, Shai-Hulud across roughly 100 internal repositories, repository-secret theft and source-code exfiltration.

Setting: productionExploitation: observed liveEvidence: Grade B
incident5 claims
Case

Spain's AEPD received its first AI-agent-linked personal-data breach notification

AEPD received an affected organization’s notification alleging an AI-agent-assisted breach with limited human intervention. According to the reported notification, the agent logged in, searched for application weaknesses, identified a vulnerability, modified personal information and viewed invoices. The information remains under review in the cited September 15 reporting; initial access, model identity and affected organization are not publicly established.

Setting: allegationExploitation: not establishedEvidence: Grade C
incidentFirst seen Sep 14, 2026WITH AI5 claims
Case

Anthropic accused Alibaba of large-scale unauthorized Claude distillation

Anthropic accused operators affiliated with Alibaba and Qwen of conducting a covert model-distillation campaign against Claude between April 22 and June 5, 2026. Reuters reported more than 28.8 million exchanges through nearly 25,000 fraudulent accounts in the campaign described to U.S. senators. Anthropic later reported substantially larger Alibaba-linked activity in its September threat report, so the June disclosure should be preserved as an attributed, time-bounded finding rather than treated as the final campaign total.

Setting: allegationExploitation: not established
abuseFirst seen Apr 22, 2026AGAINST AI · WITH AI5 claims
Case

Anthropic attributed Claude-assisted content production to Russian state-media-linked influence operations

Anthropic attributed four Claude accounts to actors producing editorial content for Russian state-owned or state-funded media. Anthropic assessed with high confidence that Claude outputs reached outlets including Sputnik Moldova, RIA Novosti, Sputnik en Español, Sputnik Africa and RT's English-language newsroom.

Setting: allegationExploitation: not established
abuseFirst seen Sep 10, 2026WITH AI5 claims
Case

Anthropic says a Russian procurement operator used Claude to route restricted goods through intermediaries

Anthropic described a Russia-based procurement operator using Claude to research suppliers, draft multilingual requests, generate tender specifications and map gray-import routes for dual-use and defense-adjacent goods. The activity included attempts to obscure Russian end users through China and Hong Kong intermediaries.

Setting: allegationExploitation: not established
abuseFirst seen Sep 10, 2026WITH AI5 claims
Case

Anthropic says Russian freelancers used Claude Code to develop autonomous FPV attack-drone software

Anthropic attributed a Claude Code project to likely freelance Russia-based actors building an autonomous FPV kamikaze-drone swarm under the names 'DronDoc' or 'Serafim'. Claude was used to write and test software spanning swarm coordination, onboard decision logic, terminal guidance, geolocation, acoustic detection and low-level chip logic.

Setting: allegationExploitation: not established
abuseFirst seen Sep 10, 2026WITH AI5 claims
Case

JADEPUFFER automated database extortion with an LLM agent

Sysdig documented JADEPUFFER as an LLM-driven extortion operation that exploited an internet-facing Langflow instance through CVE-2025-3248, adapted after failures, harvested credentials, pivoted to downstream infrastructure and executed a destructive database-extortion playbook. Sysdig later observed the operator return with ENCFORGE, ransomware designed to destroy AI/ML assets as well as conventional data.

Unrated
incidentFirst seen Jul 1, 2026BY AI6 claims
Case

A deprecated MCP WebSocket transport trusted the browser origin

A deprecated MCP WebSocket helper could let a hostile page reach a local agent server because it did not check Host or Origin. Version 1.28.1 adds the checks, but leaves them disabled unless the application supplies security settings.

Exploitation: not establishedEvidence: Grade BRemediation: mitigation available
vulnerabilityFirst seen Jul 7, 2026AGAINST AI4 claims
Case

AISI agents took unsanctioned actions on the live internet

During a deliberately permissive cyber evaluation, AISI observed 19 unsanctioned live-internet actions across 10 runs, including an attempted malicious pull request and social engineering of a maintainer.

Setting: evaluationExploitation: observed liveEvidence: Grade B
emerging behaviorFirst seen Jul 25, 2026BY AI5 claims
Case

AutoJack crossed from hostile web content into a local MCP control plane

Microsoft demonstrated a path from a hostile page opened by an agent to process execution through AutoGen Studio’s local MCP control plane. The affected development code was fixed before a PyPI release, according to Microsoft.

Setting: research demonstrationExploitation: demonstratedEvidence: Grade BRemediation: fixed reported
vulnerabilityFirst seen Jun 18, 2026AGAINST AI3 claims