AI security research / sources
Evidence
Primary reports, advisories, papers and original records behind the Claims. Each record supports specific evidence rather than a whole Case by default.
172 Evidence · Page 6 of 6
22 matching records on this page
Published sources
Prompt / instruction review - reviewed snapshot sha256:ce891
Operator-supplied text retained by hash. This is primary submitted material, not independent corroboration. Full input is not automatically redistributed.
Reuters - Anthropic alleges Alibaba illicitly extracted Claude capabilities
Reuters - Kimi K3 breaks out of testing environment
Reuters - OpenAI introduces model-misalignment reporting framework
Sabotage Risk Report: Claude Opus 4.6
Anthropic risk report discussing sabotage pathways and safeguards. Section 5.1 states that egress-bandwidth controls would make model-weight exfiltration harder and increase the chance of detecting an attempt.
Safety Invariants for Agents Orchestrating Irreversible State Transitions
SecurityWeek - Anthropic warns Claude users of infostealer infections
Sirens’Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs
Spanish data watchdog publicises first AI agent-linked data breach report
Reuters report on the AEPD notification, including the reported attack sequence, limited human intervention, unidentified organization and model, and the regulator's statement that the filing remained under review.
Sysdig - JADEPUFFER agentic ransomware
Sysdig - JADEPUFFER evolves to target AI/ML assets
TechCrunch - Anthropic multi-agent turf-war experiment
TechCrunch - Microsoft open-source projects compromised to steal AI developer credentials
The Hacker News - Claude Opus 4.6 gym-booking boundary violation reproduction
Trellix - OpenClaw supply-chain crisis / ClawHavoc
Trevor Blackwell / exfil source repository
Primary project repository for the GET-based transfer demonstration. The existing corpus preserves its mechanism and separate reproduction locators. Repository-page/raw retrieval was not available in this pass; no production model-weight theft is established by the retained evidence.
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
tronify.rent threat-intelligence record
Living threat-intelligence record for tronify.rent. It records first detection on February 26, 2026, later abuse-report activity and continuing risk signals, establishing that public warning evidence about the domain existed before the alleged September losses.
WIRED - Zoom screen-sharing takeover bug
ZDI-26-648 / CVE-2026-19590
ZDI-26-649 / CVE-2026-19591
No matches on this page.