Security & Governance

Anthropic's fourth threat intelligence report details missile guidance, bioweapons queries, and 200 million distillation attacks

Anthropic's September 2026 threat report covers eight months of Claude misuse across seven harm areas, from Yemen missile software to Chinese AI distillation.

security governance category

Anthropic published its fourth threat intelligence report on 10 September 2026, covering disrupted misuse of Claude between December 2025 and August 2026. The full report runs to approximately 36,000 words and documents activity across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and what Anthropic calls illicit distillation. The cases were chosen because they are novel or sophisticated, not because they are typical of everyday misuse on the platform.

The headline finding is uncomfortable reading. Anthropic’s earlier models, including Claude Opus 4 and Claude Sonnet 4.5 from 2025, were, in the company’s own words, “well below the threshold where they could meaningfully assist a sophisticated user in carrying out dangerous biological research.” Newer Claude models are no longer comfortably below that threshold. The report is, in part, a public acknowledgement that the risk profile of frontier AI has shifted.

A Yemeni weapons cell used Claude Code as a software team

The most striking conventional weapons case involves a northern Yemen group Anthropic designates GTG-87001. The cell used Claude Code in place of human software engineers to develop guidance, navigation, and control software for a guided rocket, a multi-stage ballistic missile with a stated range target above 2,000 km, and a hypersonic glide vehicle variant.

The group structured their Claude sessions to mirror an actual engineering department: one instance handled firmware and control code, a second researched and selected algorithms, and a third reviewed the output. They obscured their ultimate goals and split work across separate sessions so no single prompt revealed the full picture. Anthropic says its safeguards blocked many requests but acknowledges some slipped through. There is no evidence the group successfully fielded a working weapon, though an unsuccessful test-fire appears to have been attempted.

This is worth dwelling on if you work in defence, critical infrastructure, or export-controlled technology. The barrier to standing up a multi-agent engineering pipeline for weapons development has dropped to the point where a resource-constrained actor can attempt it using a commercial AI subscription. The sessions were disrupted, but the attempt was real.

Bioweapons queries from working scientists

The biological misuse section covers five case studies. In one from May 2026, a scientist sought help drafting a grant application for gain-of-function research on chikungunya virus, proposing to engineer mutations for increased transmissibility and immune evasion, to be conducted at a military research institute. In a separate case, a researcher used Claude to plan experiments on highly pathogenic avian influenza focused on mammalian adaptation. Because Anthropic’s filters blocked its more capable models, that researcher was pushed down to Claude’s weakest available model.

Anthropic is direct about the implication: the safety margin that existed in 2025 has narrowed. Organisations with biosafety obligations, dual-use research programmes, or CBRN risk assessments should treat this as primary-source vendor intelligence confirming the threat is active. It is not theoretical.

Cyber operations and “vibe hacking”

The report introduces a term worth knowing: vibe hacking. An operator gives the model a general goal, and the model surveys the environment, writes and runs scripts, summarises findings, and repeats until the task is complete. Humans act as overseers rather than hands-on operators. Anthropic first documented fully agentic attack chains in November 2025; by mid-2026, this pattern had spread to every class of actor in the report.

The Russia-linked group designated GTG-20006, associated publicly with Midnight Blizzard, ran a 130-day operation engaging 24 of 27 targeted institutions including Ukrainian ministries, defence bodies, and drone supply-chain manufacturers. One technique involved hijacking DNS records using compromised administrator credentials, redirecting hotel WiFi traffic to actor-controlled servers, and delivering ClickFix-style lures to guests’ Windows, Android, and iOS devices. A separate Russia-linked case (GTG-27005) involved building an autonomous FPV kamikaze drone swarm whose onboard model could select targets, including a “person” class, and issue detonation commands without a human in the loop. The training data was scraped Ukrainian combat footage.

If your organisation deploys Claude or similar models in agentic configurations, your misuse monitoring needs to cover multi-agent pipeline behaviour, not just individual prompt content. Acceptable use policies written for single-turn interactions do not address what the report documents.

Influence operations at scale

Nine influence operation clusters are detailed, originating in Russia, Iran, Turkey, and across the Gulf, South Asia, Africa, and Europe. One cluster (GTG-54002) published at least 8,913 articles across 70 fake news websites in approximately 20 languages, backed by more than 250 inauthentic commenting accounts. Another (GTG-84005) operated over 1,000 fake accounts on X and sought one million artificial views on its content.

The practical implication for communications, policy, and media professionals is that AI-assisted influence operations are no longer a future risk to model. They are a documented present reality, operating at a scale that manual detection cannot match.

Chinese AI labs and 200 million distillation attacks

The distillation section names Alibaba, Moonshot AI, DeepSeek, Z.ai, Xiaomi, SenseTime, and MiniMax. Illicit distillation, as Anthropic defines it, is the process of using outputs from a more capable model to train a competing model and replicate its capabilities without authorisation.

The Alibaba campaign is the largest Anthropic has ever observed. Between May and July 2026, the company recorded 151 million exchanges peaking at nearly three million per day, spread across 3,500 accounts, all using a single fixed prompt to extract training material for Alibaba’s Qwen models. DeepSeek generated more than 12 million distillation attempts over just 14 days in July 2026. Across all five campaigns, Anthropic observed close to 200 million exchanges linked to distillation attacks. Some of the extracted content included sensitive information from individual users, major multinational companies, and state-affiliated actors, which Anthropic notes is likely inconsistent with privacy laws and the labs’ own terms of service.

This matters beyond the AI industry. On 11 September 2026, the NSA, FBI, and CISA issued a joint cybersecurity advisory formally naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, with the intelligence community assessing that these activities likely occur with the awareness and implicit support of the Chinese government. If your organisation uses any of the named Chinese AI products, you should review what data those interactions may have exposed.

What does this mean for you?

A few concrete takeaways depending on where you sit:

If you run AI deployments in your organisation, the agentic attack patterns documented here mean your AI governance and monitoring frameworks need to extend beyond prompt-level review. Multi-session, multi-agent pipelines are now the attack surface.

If you work in biosafety or dual-use research compliance, Anthropic has now published primary-source evidence that working scientists are querying frontier AI models for bioweapons-relevant information. That belongs in your risk register.

If you are evaluating Chinese AI models for enterprise use, the distillation findings and the joint government advisory issued on 11 September 2026 add significant due diligence questions about data handling and training provenance.

If you work in defence, export control, or critical infrastructure, the Yemen weapons case is a proof of concept that low-resource actors can attempt to use commercial AI as a substitute engineering team. The attempt was disrupted, but the capability gap has narrowed.

Anthropic says it disrupted every operation in the report and shared relevant findings with authorities and other AI companies where appropriate. That is worth acknowledging. It is also worth noting, as the report itself does, that these cases show where safeguards work and where they need to improve. The fact that some requests in the Yemen case slipped through is in the report. Anthropic did not hide it.

The full report is available on Anthropic’s CDN. It is detailed, specific, and worth reading if any of the above areas are within your professional remit.