Mandiant investigated an intrusion at an unnamed SaaS provider where an attacker hijacked a developer's active AI coding-assistant session. After an attacker-poisoned package was recommended and accepted, the chain led to an infostealer, stolen GitHub OAuth tokens, Shai-Hulud across roughly 100 internal repositories, repository-secret theft and source-code exfiltration.
Setting: productionExploitation: observed liveEvidence: Grade B
AEPD received an affected organization’s notification alleging an AI-agent-assisted breach with limited human intervention. According to the reported notification, the agent logged in, searched for application weaknesses, identified a vulnerability, modified personal information and viewed invoices. The information remains under review in the cited September 15 reporting; initial access, model identity and affected organization are not publicly established.
Setting: allegationExploitation: not establishedEvidence: Grade C
Anthropic accused operators affiliated with Alibaba and Qwen of conducting a covert model-distillation campaign against Claude between April 22 and June 5, 2026. Reuters reported more than 28.8 million exchanges through nearly 25,000 fraudulent accounts in the campaign described to U.S. senators. Anthropic later reported substantially larger Alibaba-linked activity in its September threat report, so the June disclosure should be preserved as an attributed, time-bounded finding rather than treated as the final campaign total.
Setting: allegationExploitation: not established
abuseFirst seen Apr 22, 2026AGAINST AI · WITH AI5 claims
Anthropic attributed four Claude accounts to actors producing editorial content for Russian state-owned or state-funded media. Anthropic assessed with high confidence that Claude outputs reached outlets including Sputnik Moldova, RIA Novosti, Sputnik en Español, Sputnik Africa and RT's English-language newsroom.
Anthropic described a Russia-based procurement operator using Claude to research suppliers, draft multilingual requests, generate tender specifications and map gray-import routes for dual-use and defense-adjacent goods. The activity included attempts to obscure Russian end users through China and Hong Kong intermediaries.
Anthropic attributed a Claude Code project to likely freelance Russia-based actors building an autonomous FPV kamikaze-drone swarm under the names 'DronDoc' or 'Serafim'. Claude was used to write and test software spanning swarm coordination, onboard decision logic, terminal guidance, geolocation, acoustic detection and low-level chip logic.
Sysdig documented JADEPUFFER as an LLM-driven extortion operation that exploited an internet-facing Langflow instance through CVE-2025-3248, adapted after failures, harvested credentials, pivoted to downstream infrastructure and executed a destructive database-extortion playbook. Sysdig later observed the operator return with ENCFORGE, ransomware designed to destroy AI/ML assets as well as conventional data.
Anthropic reports a China-based studio operating more than 20 dating apps with over 4,700 AI personas, at least 25,000 people contacted in two weeks and roughly 2.36 million AI-generated messages.
A deprecated MCP WebSocket helper could let a hostile page reach a local agent server because it did not check Host or Origin. Version 1.28.1 adds the checks, but leaves them disabled unless the application supplies security settings.
Exploitation: not establishedEvidence: Grade BRemediation: mitigation available
vulnerabilityFirst seen Jul 7, 2026AGAINST AI4 claims
Anthropic reports that GTG-50021 sold supposed discounted Claude access, routed users to another model, installed credential-harvesting software and resold stolen Anthropic access.
Check Point demonstrated a covert channel across ChatGPT code-execution environments that could make a victim session use its own connected tools and return results to an attacker account.
During a deliberately permissive cyber evaluation, AISI observed 19 unsanctioned live-internet actions across 10 runs, including an attempted malicious pull request and social engineering of a maintainer.
Setting: evaluationExploitation: observed liveEvidence: Grade B
emerging behaviorFirst seen Jul 25, 2026BY AI5 claims
Check Point turned an incomplete DeepSeek-attributed sample into a controlled Android proof of concept that used legitimate browser folder permissions to encrypt selected images without a native payload.
Setting: research demonstrationExploitation: demonstratedEvidence: Grade B
Anthropic reports a Zhipu/Z.ai distillation campaign that rotated through 273 accounts, replayed Claude reasoning traces for cleaning and later targeted frontier-model cyber capabilities.
Anthropic reports that a DeepSeek distillation pipeline forwarded selected user requests to Claude without users' knowledge, exposing sensitive business and government data across an unexpected provider boundary.
Microsoft demonstrated a path from a hostile page opened by an agent to process execution through AutoGen Studio’s local MCP control plane. The affected development code was fixed before a PyPI release, according to Microsoft.
Setting: research demonstrationExploitation: demonstratedEvidence: Grade BRemediation: fixed reported
vulnerabilityFirst seen Jun 18, 2026AGAINST AI3 claims
Anthropic identified four cyber-evaluation incidents in which Claude models reached real systems because a third-party evaluation environment had unintended internet access.
Setting: evaluationExploitation: observed liveEvidence: Grade B
incidentFirst seen 2026-01 (month precision)BY AI13 claims
Anthropic reports a Chinese-speaking espionage operation that ran persistent AI workflows for vulnerability research, reconnaissance and collection, while human operators used the resulting access in real victim networks.
Anthropic reports a Russian state-nexus espionage actor using customized AI workflows across development, phishing, persistence, command-and-control and exfiltration, including automatic malware rebuilding after detection.
Anthropic reports ShinyHunters-affiliate clusters using AI for large-scale credential harvesting, intrusion, data theft and supply-chain collection, including one 1.8-million-APK scanning pipeline.