Anthropic’s AI Misuse Report Shows Claude in Surveillance and Influence Operations
The cases show AI compressing technical and operational work for threat actors—and why account bans alone cannot substitute for tool controls, monitoring, and cross-platform response.

Bottom line
Anthropic says it disrupted Claude misuse across seven harm areas, including mass surveillance and influence campaigns. Here is what the cases establish and what remains unknown.
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 4
- Products covered
- 1
- Last checked
- 2026-09-14
Important limits
- • Features, availability, and pricing can change after publication; confirm consequential details with the provider.
In this guide
*This analysis covers Anthropic's September 10, 2026 threat-intelligence report and related policies. The case studies are based primarily on Anthropic's internal visibility and attribution; they are not an independent census of AI-enabled abuse.*
The short answer
Anthropic says it identified and disrupted malicious use of Claude between December 2025 and August 2026 across cyber operations, influence campaigns, non-consensual surveillance, scams and fraud, biological misuse, conventional-weapons development, and illicit model distillation. The report describes actors using AI not only to write content but to build software, analyze targets, translate and impersonate, prepare technical documents, and coordinate longer operational workflows.
The central trend is labor compression. Anthropic reports cases where a single operator used Claude to perform work that previously required engineering or analyst teams. But the document covers activity Anthropic detected on its own service; it cannot show the total prevalence of misuse, everything actors did elsewhere, or how often safeguards stopped harmful requests before a case escalated.
What kinds of misuse did Anthropic report?
The report spans seven harm categories. Examples include fake social profiles and news sites used in influence operations, software for mass surveillance and identity harvesting, fraud workflows, prohibited biological research assistance, weapons-related engineering and documentation, cyber activity, and attempts to extract model behavior through distillation.
In surveillance cases, Anthropic says actors used Claude to turn bulk social posts into structured profiles containing inferred location, demographics, and political views. In influence cases, operators produced coordinated personas and deceptive content for audiences across multiple continents. Some weapons cases involved simulation, firmware, targeting, or procurement work, though Anthropic explicitly notes important uncertainty about attribution and whether several systems became operational.
Did Claude independently carry out these operations?
No. Human actors supplied objectives, domain knowledge, data, access, hardware, and external infrastructure. Claude accelerated portions of planning, coding, analysis, translation, documentation, or iteration. The report does describe more agentic use, including writing directly to project files and supporting multi-step workflows, but that is different from an autonomous system choosing the operation.
This distinction does not make the risk trivial. Reducing the people, time, language ability, or specialist knowledge needed for harmful work can change who is capable of attempting it. It does mean readers should avoid turning vendor case descriptions into claims that a chatbot independently ran a state operation.
What did Anthropic do in response?
Anthropic says it banned associated accounts, strengthened detections using observed behavioral signatures, and shared indicators with authorities and industry partners where appropriate. Its Usage Policy prohibits areas including non-consensual surveillance, weapons development, fraud, and violations of civil liberties.
The report also exposes a control gap: in one described cluster, Claude refused explicit profiling and propaganda prompts but did not refuse many software-building requests that supported the same surveillance operation. Intent can be fragmented across innocuous-looking subtasks, so filters at the prompt level are not enough.
What should organizations learn from the cases?
Govern the whole agent system. Give models distinct identities, least-privilege credentials, allow-listed tools and destinations, limits on steps and spending, logging that captures actions as well as chat, and human approval before bulk outreach, account creation, surveillance-like profiling, code deployment, or physical-world changes.
Security teams should look for behavior across sessions: unusual volumes, repeated persona creation, bulk targeting, attempts to evade geographic controls, sensitive tool combinations, and workflows deliberately divided into small pieces. Providers and customers also need incident-response paths that can revoke credentials and share indicators without exposing unrelated user data.
The verdict
Anthropic's report provides unusually concrete evidence that AI assistants can compress parts of real malicious workflows across several domains. Its most useful lesson is not that one model is uniquely dangerous; it is that general coding, language, research, and tool-use capabilities become dual-use when connected to data and action.
The evidence also has limits. This is a vendor-authored sample of detected activity, not a denominator-based measurement of misuse or an independent audit of Anthropic's controls. Organizations should use it to improve system-level defenses while demanding comparable reporting, clearer metrics, and external scrutiny across the industry.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
What did Anthropic's September 2026 threat report find?
Anthropic documented and disrupted Claude misuse across cyber operations, influence, surveillance, fraud, biological misuse, conventional weapons, and illicit model distillation between December 2025 and August 2026.
Did Claude autonomously run the reported operations?
No. Human actors provided goals, access, data, infrastructure, and expertise. Claude accelerated tasks such as coding, analysis, translation, documentation, and iteration within their workflows.
How did Anthropic respond to the misuse?
Anthropic says it banned linked accounts, incorporated observed behaviors into detections, and shared relevant threat information with authorities and industry partners where appropriate.
What is the main limitation of the report?
It is based largely on activity Anthropic detected through its own service. It does not measure all AI misuse, provide a reliable denominator, or independently validate every attribution and downstream outcome.
Found this useful?
Get the next one in your inbox.
One five-minute briefing a week: a meaningful change, a practical workflow, and a clearer tool decision—already filtered for lean teams.
Free · one email a week · unsubscribe any time
Tools mentioned in this article
Claude
Anthropic's thoughtful, safety-focused AI with exceptional long-form reasoning
Claude excels at deep analysis, long-form writing, and nuanced reasoning. Built by Anthropic with a focus on safety and helpfulness.
Read next
Recommended for you

JobCopilot Review 2026: Is AI Auto-Apply Worth It?
JobCopilot can remove hours of repetitive application work, but match quality, truthful answers, and human review matter more than raw submission volume.
A research-based JobCopilot review covering automated applications, current pricing, privacy, risks, alternatives, and a practical buyer test.
Read guide
GPT-6 Astra vs Claude for Business: Which Should Your Team Choose?
DocsBot AI Review 2026: Pricing, Accuracy, Privacy, and Fit
TypingMind Review 2026: Multi-Model Chat, Pricing, and Privacy