Practical Guides

Anthropic Found Seven Misuse Categories—Its Cases Were Not Typical

|Author: QUASA Editorial Team|5 min read| 7
Anthropic Found Seven Misuse Categories—Its Cases Were Not Typical

On September 10, Anthropic published an official threat-intelligence report covering activity it disrupted from December 2025 through August 2026 across seven harm areas. The selected cases involved Claude Haiku-, Sonnet- and Opus-class models.

The disclosure documents serious but deliberately exceptional incidents, not ordinary Claude use or the prevalence of abuse. AP’s account of the disclosure includes Anthropic’s warning that the cases were notable and novel rather than typical misuse.

Seven categories require seven different readings

Anthropic’s documented Claude misuse cases separated by actor, task, harm, intervention and evidentiary limits

The categories combine different actors, AI-assisted tasks and levels of demonstrated harm. The resulting case matrix shows what Anthropic observed and interrupted, but it cannot establish how often comparable conduct occurs on Claude or elsewhere.

  • Cyber operations: suspected state-linked groups, criminals and individual attackers used Claude for reconnaissance, phishing, exploit work, malware modification and data processing. Linked accounts were banned and detections were updated, but provider records do not establish the outcome of every intrusion or precisely quantify the advantage produced by AI.
  • Influence operations: government-linked, commercial and political operators generated personas, political material and content for fabricated or established outlets. Anthropic removed associated access and developed behavioral detections. Several operations attracted little authentic engagement, so content production does not by itself demonstrate successful persuasion.
  • Surveillance: state-aligned actors, contractors and commercial vendors used Claude to profile people, query monitoring systems or help design surveillance software. Responses included account bans and intelligence sharing. In one case, enforcement interrupted development activity but could not disable a locally deployed platform that used another model.
  • Scams and fraud: the selected cases included deceptive dating applications supported by automated conversations and related infrastructure. Accounts and organizations tied to the operation were removed. The evidence documents a provider-observed system, not a complete victim count or verified financial loss.
  • Biological misuse: working scientists used Claude for research that could support dangerous work involving pathogens, venoms or toxins. Their accounts were banned and the findings informed safeguards, but identities were withheld and harmful intent was not established because the research was dual-use.
  • Conventional weapons: actors linked to China, Russia and northern Yemen used Claude for weapons software, technical research, procurement or intelligence. Linked accounts were banned and information was shared with partners. Some actor claims remained unverifiable, while a guided-rocket field test did not establish an operational weapon.
  • Illicit distillation: China-based laboratories were accused of using fraudulent accounts, proxies and large-scale prompting to extract Claude capabilities. Anthropic blocked accounts and changed access controls. The attribution and measurements are company findings, not the result of an independent audit.

Reuters’ independent coverage summarized the biological, weapons, cyber-espionage and distillation allegations, while noting that the named Chinese companies did not immediately respond and China’s foreign ministry was unaware of the report.

The affected model classes set a narrow boundary

Model-safety review distinguishing affected Claude classes from newer systems with stronger safeguards

The selected incidents involved Haiku-, Sonnet- and Opus-class systems. Fable- and Mythos-class models appeared in none of the cases except one illicit-distillation campaign. That boundary applies only to this collection and does not demonstrate that newer systems are immune to abuse.

The biological cases expose a further distinction. Evaluations placed the older Claude Opus 4 and Sonnet 4.5 systems below the threshold for meaningfully assisting a sophisticated user with dangerous biological research, allowing narrower safeguards focused largely on preventing novice uplift.

That assurance no longer extends confidently to more capable systems, which can assist with complex scientific tasks. Anthropic therefore applied broader restrictions to dual-use biological queries on newer models. This does not prove that an older model produced a biological weapon or that the newer controls eliminate risk; it shows that capability and access restrictions must be assessed together.

Defenders should look for workflows, not an AI fingerprint

Security team correlating behavioral indicators and interrupting an AI-assisted intrusion workflow

For model providers, recurring indicators include account clusters sharing infrastructure, fraudulent or stolen payment methods, stolen API credentials, reseller traffic, geographic proxies, repeated attempts to extract reasoning and projects divided across sessions so that no single request exposes the full objective. Automated tool use becomes more significant when it coincides with targeting, credential theft or exfiltration.

Enterprise teams can correlate the corresponding activity on their own systems. The cyber cases paired familiar access routes—phishing, exposed services and stolen credentials—with rapid reconnaissance, bulk mailbox access, automated data processing and repeated malware rebuilding after detection. Those combined events provide stronger evidence than an attempt to identify an AI-generated signature in a single file.

The disclosure also supplies conventional indicators of compromise for particular campaigns. Individual domains and infrastructure can change quickly; identity, account, payment, tool-call and network telemetry may reveal the longer workflow after one indicator has expired.

Disruption did not necessarily prevent downstream harm

In this report, “disrupted” primarily means that Anthropic banned every account it could connect to an actor. Some cases also produced new classifiers, stronger monitoring, partner notifications or coordinated action against infrastructure.

Removing Claude access does not recover stolen data, disable every external system or prove that downstream harm was prevented. The surveillance example involving an already deployed local platform makes that limit especially clear: provider enforcement can stop further use of one service without neutralizing the broader operation.

The disclosure provides no denominator for total Claude activity, detected misuse or abuse that escaped detection. It supports the finding that Anthropic encountered activity in seven harm areas, but not a claim that any category is common, increasing at a measurable rate or representative of Claude users. Independent validation of individual operations, their ultimate outcomes and the effectiveness of the resulting safeguards remains incomplete.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0