Case · DiggingBeagle record

Anthropic attributes a large reasoning-trace extraction pipeline to Zhipu/Z.ai

Anthropic reports a Zhipu/Z.ai distillation campaign that rotated through 273 accounts, replayed Claude reasoning traces for cleaning and later targeted frontier-model cyber capabilities.

Anthropic's September 2026 attribution. The current source set does not include a Zhipu response.

First seen
Sep 10, 2026
Case kind
incident
AI role
WITH AI
Claims
3

Reconstruction

Anthropic describes a production-scale extraction pipeline rather than occasional prompt copying. Zhipu allegedly rotated through 273 fraudulent accounts, captured Claude reasoning traces, replayed them through Claude to clean the transcripts and used the outputs in post-training work.

Over a ten-day period in June, Anthropic counted 770,609 exchanges through the chain-of-thought cleaner and more than three million exchanges attributed to Zhipu overall. The same report says Claude was used to judge outputs, normalize traces, score training data and implement testing.

The security relevance increased when the target moved toward cyber capability. Anthropic says Zhipu used public vulnerability datasets to construct capture-the-flag challenges and then ran distillation attacks against frontier models to acquire stronger cyber performance. That makes the boundary between model extraction and capability transfer explicit.

Mechanism & boundary

  1. 01

    Rotate fraudulent access

    Hundreds of accounts distribute requests intended to bypass provider restrictions.

    Boundary: account boundary / model access

  2. 02

    Capture reasoning traces

    The reported pipeline records model reasoning material.

    Boundary: model output / extraction store

  3. 03

    Replay traces through a cleaner

    Captured traces are sent back through Claude for normalization.

    Boundary: raw trace / training-ready data

  4. 04

    Target cyber capability

    The operator builds CTF tasks and seeks stronger frontier-model cyber performance.

    Boundary: general distillation / capability-specific transfer

Timeline

  1. Sep 10, 2026

    Anthropic publishes GTG-16006 case study

    report

    The report describes account rotation, CoT cleaning and later cyber-capability targeting.

Claims & evidence

reported findingsupported

Anthropic says Zhipu later used public vulnerability datasets to build capture-the-flag challenges and targeted the cyber capabilities of frontier models during distillation work.

  • supports
    Countering misuse of AI: September 2026

    Locator: GTG-16006 cyber-capability targeting

    Anthropic describes CTF construction and later targeting of another frontier lab's top model for cyber-capability distillation.
reported findingsupported

Anthropic reports 770,609 exchanges through the chain-of-thought cleaner and more than three million Zhipu-attributed exchanges over the same 10-day period.

Measured value
770609 cleaner exchanges
Method
Anthropic threat-intelligence attribution
Period
10 days in June 2026
reported findingsupported

Anthropic reports that Zhipu/Z.ai rotated through 273 fraudulent accounts during a 10-day chain-of-thought extraction campaign against Claude Opus 4.8.

Evidence visuals

chart

Zhipu/Z.ai distillation activity reported by Anthropic

Anthropic September 2026 threat report. The total is a lower bound because the report says over 3 million.

Measureexchanges
Total attributed3000000
CoT cleaner770609
The cleaner count is a subset/workstream inside the broader attributed traffic, not an independent total. · Source: Anthropic attributes a large reasoning-trace extraction pipeline to Zhipu/Z.ai

diagram

Zhipu reasoning-trace extraction pipeline

  1. 273 fraudulent accounts

    Rotated to evade restrictions

  2. Capture reasoning traces

    Record Claude reasoning material

  3. CoT cleaner

    Replay traces through Claude for normalization

  4. Post-training pipeline

    Use cleaned outputs for judging, filtering and training work

  5. Cyber-capability tasks

    Later CTF work targets frontier-model cyber performance

  • 273 fraudulent accounts Capture reasoning traces: query
  • Capture reasoning traces CoT cleaner: replay
  • CoT cleaner Post-training pipeline: normalized traces
  • Post-training pipeline Cyber-capability tasks: capability targeting
Project-authored reconstruction from Anthropic's GTG-16006 case study. · Source: Anthropic attributes a large reasoning-trace extraction pipeline to Zhipu/Z.ai

Implications

Model-access abuse can become a data pipeline. Detection needs to look for account rotation, replay signatures, structured extraction workflows and capability-specific task distributions rather than only individual suspicious prompts.

Controls & mitigations

  • Detect coordinated account rotation and repeated trace-replay patterns.
  • Rate-limit and investigate structured extraction workloads spanning many accounts.
  • Treat capability-specific distillation against cyber tasks as a distinct high-risk pattern.

What remains unknown

  • The current source set contains Anthropic's attribution but no Zhipu response.
  • The report does not establish how much downstream model capability was gained from the campaign.

Cite this record

DiggingBeagle. “Anthropic attributes a large reasoning-trace extraction pipeline to Zhipu/Z.ai.” First seen Sep 10, 2026. https://diggingbeagle.com/cases/anthropic-attributes-a-large-reasoning-trace-extraction-pipeline-to-zhipu-z-ai/

Citation guidance

Why this archive exists

The source matters after the headline fades.

DiggingBeagle is a non profit research project documenting AI security incidents, agent failures, vulnerabilities and AI-assisted operations. A case keeps its claims beside the sources that support, contest or limit them. Later updates stay visible, so a reader can see when the account changed.

We publish case reconstructions, dated reporting and analysis across records. Each has a different evidentiary role. About the project and our methodology explain how the work is reviewed.