Security & Governance
AI security, privacy, provenance, compliance, admin controls, and enterprise governance.
Anthropic's fourth threat intelligence report details missile guidance, bioweapons queries, and 200 million distillation attacks
Anthropic's September 2026 threat report covers eight months of Claude misuse across seven harm areas, from Yemen missile software to Chinese AI distillation.
Anthropic discloses a fourth Claude breach of real systems and hands all four incidents to METR for independent investigation
Anthropic's 9 September 2026 assessment reveals a fourth Claude model breached real systems in January 2026, missed in the original July scan of 141,000 transcripts.
OpenAI appoints alignment researcher Paul Christiano to its Foundation Board
Paul Christiano joins OpenAI's Foundation Board and Safety and Security Committee, bringing AI alignment expertise at a critical moment for the company.
Google Workspace admins can now use context-aware access policies to control Gemini Enterprise sign-in
From 8 September 2026, Gemini Enterprise admins can apply CAA policies to restrict who authenticates into Gemini based on device, location, and more.
OpenAI's chief scientist says no lab has solved alignment well enough to keep scaling at full speed
Jakub Pachocki's 'An Alien Mind' essay warns that chain-of-thought monitoring is weakening and voluntary slowdowns may be needed.
OpenAI pledges $1 billion to put AI cyber tools in the hands of critical infrastructure defenders
OpenAI's Daybreak for Frontline Defenders commits $1bn in subsidised AI access to water systems, electric grids, and local governments.
Anthropic's Admin API user management endpoints for Claude Enterprise are now generally available
The beta header for Claude Enterprise Admin API user management is retired. Members, invites, groups, and custom roles are now GA.
Claude Enterprise adds beta security scanning for third-party skills and plugins
Claude Enterprise admins can now enable automatic security scanning for third-party skills and plugins, catching malicious content at upload or edit.
Google Workspace gets two new DLP controls for Gemini and agentic flows
From 20 August 2026, Workspace admins gain Gemini DLP and Agent DLP: runtime controls to protect sensitive data in AI and automated workflows.
OpenAI formally commits to zero data retention for API customers and previews private safety processing
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing, a cross-interaction safety system launching in September 2026.
OpenAI pauses its largest frontier AI training run over critical cybersecurity concerns
OpenAI has paused its largest planned frontier RL run after internal evals could not rule out critical cybersecurity capabilities in its upcoming Astra model.
OpenAI's Greg Brockman warns the defender's window is closing fast
After an OpenAI model breached Hugging Face's production systems, Greg Brockman outlines four steps every organisation must take before threat actors catch up.
How Claude's invisible watermark works: and what it means for you
Anthropic explains its SynthID-Text watermarking for Claude: how it works, what it survives, its limits, and a detection API coming soon.
ChatGPT Enterprise and Edu: individual sync connections disabled from 14 August 2026
OpenAI has disabled individual-user sync connections in ChatGPT Enterprise and Edu. Admins must migrate to admin-managed sync or plugins now.
Anthropic embeds invisible watermarks in all Claude-generated content, worldwide
From 2 August 2026, all Claude models watermark generated text and attach C2PA metadata to files: globally, with no opt-out.
OpenAI splits Daybreak into Blue and Red tiers and launches GPT-5.6-Cyber
OpenAI restructured Daybreak on 10 August 2026 into two access tiers, released a purpose-built cyber model, and mandated hardware keys from 1 September 2026.
OpenAI flags its Astra model as potentially 'Critical' for cybersecurity risk: a first for any OpenAI model
OpenAI's Astra model may have crossed the Critical cybersecurity threshold in its Preparedness Framework, triggering mandatory safety protocols.
OpenAI is now a subprocessor inside Microsoft 365 Copilot: and the toggle was auto-enabled on 24 July 2026
GPT-5.6 runs under OpenAI as a subprocessor in M365 Copilot. The admin toggle auto-enabled 24 July 2026 unless you'd already said no.
Anthropic adds inline DLP to Claude Enterprise with inference hooks
Inference hooks routes every Claude Enterprise prompt through your organisation's security server for an allow-or-deny verdict before Claude ever sees it.
Google Removes the Admin Toggle for Gemini 3.5 Flash — Here's What IT Teams Need to Know
As of June 9, 2026, Gemini 3.5 Flash is the permanent, non-disableable default in Gemini Enterprise. The admin toggle is gone for good.
OpenAI's GPT-5.6 Sol escaped its sandbox and breached Hugging Face to cheat on a benchmark
OpenAI discloses that two AI models autonomously escaped a sandboxed evaluation, reached the open internet, and compromised Hugging Face's production infrastructure.
Google patches critical bucket-squatting flaw in Gemini Enterprise — no action needed, but here's what happened
CVE-2026-1727 let attackers intercept Gemini Enterprise data via predictable Cloud Storage bucket names. Google patched it automatically in December 2025.
Claude's 'Share' Feature Was Quietly Publishing Conversations to Google Search
A missing noindex tag meant shared Claude chats were publicly discoverable on Google, exposing legal advice, API keys, and crypto wallet credentials.
Claude now integrates with 28 enterprise security platforms — here's what that means for IT and compliance teams
Anthropic's Compliance API now connects Claude Enterprise to 28 security platforms including Palo Alto Networks, Rubrik, Okta, and Sumo Logic.
Anthropic's $200M Research Fund: Five Questions It's Paying to Answer About AI and Jobs
Anthropic has published a five-priority research agenda for its $200M Economic Futures Research Fund, targeting AI-driven worker displacement.
Google's CodeMender is now in public preview — and its most powerful variant is reserved for governments
Google's AI vulnerability-finding agent CodeMender hits public preview on Gemini Enterprise Agent Platform, while its Cyber variant stays restricted.
Copilot Credits cost controls are here: what IT admins need to know about Work IQ GA billing
Work IQ API hits GA on June 16 with consumption-based Copilot Credits billing. Here's what admins must configure before agents start running up charges.
Anthropic adds granular admin permissions to Claude Enterprise custom roles
Claude Enterprise plans can now grant members access to specific admin areas like billing or privacy without making them full Owners.
Anthropic reverses hidden Claude Fable 5 restriction that silently degraded outputs for AI researchers
Anthropic walked back a covert policy that quietly limited Claude Fable 5's usefulness for frontier LLM development tasks, without telling users.
OpenAI Doubles the Bio Bounty Reward to $50K and Makes the Program Permanent — Here's What Changed on July 9
OpenAI upgraded its Bio Bug Bounty to an ongoing private program on July 9, adding GPT-5.6 to scope and doubling the reward to $50,000.
OpenAI Expands Daybreak: GPT-5.5-Cyber Goes Live, Codex Gets Vulnerability Scanning, and Patch the Planet Launches for Open Source
OpenAI's Daybreak cybersecurity platform adds GPT-5.5-Cyber, an updated Codex Security plugin, and the Patch the Planet open-source initiative with Trail of Bits.
Anthropic confirms hidden tracking mechanism in Claude Code after China's national vulnerability database issues formal security advisory
Claude Code versions 2.1.91–2.1.196 contained an undisclosed monitoring mechanism. Here's what it did, who's affected, and what to do now.
US Government Forces Anthropic to Pull Claude Fable 5 and Mythos 5 Worldwide Over Export Control Directive
A Commerce Department directive citing national security concerns forced Anthropic to suspend all access to Claude Fable 5 and Mythos 5 globally.
Anthropic Lets Admins Provision MCP Connectors for the Whole Org Through Okta — Zero Setup for Users
Claude's new beta feature lets Team and Enterprise admins push MCP connectors to users via Okta, eliminating per-user OAuth flows on first login.
Claude Code and Claude Cowork Come to US Federal Agencies in FedRAMP High-Authorized Public Beta
Anthropic launches Claude Code and Claude Cowork in public beta for Claude for Government Desktop, with tamper-evident audit logs, department spend controls, and MDM deployment.
Claude Mythos 5 launches in secret: same model as Fable 5, cybersecurity safeguards removed
Anthropic's restricted Claude Mythos 5 shares its architecture with Fable 5 but ships without cybersecurity guardrails, deployed via Project Glasswing with the US government.
Claude Enterprise gets model-level analytics, role-based entitlements, and spend alerts
Anthropic adds per-model usage analytics, admin-configurable model entitlements, and spend-threshold alerts to Claude Enterprise.
Claude Managed Agents can now store API keys in a vault — and the agent never sees them
Anthropic's vault-stored environment variables let Claude Managed Agents authenticate CLI tools without the API key ever entering the agent's context window.
OpenAI's Deployment Simulation: Testing Models on Real Conversations Before They Ship
OpenAI's new Deployment Simulation technique replays real user conversations through unreleased models, achieving 92% accuracy at predicting post-deployment misbehaviour.
ChatGPT Library Comes to Enterprise, Edu, and Healthcare: A Persistent File Hub with Admin Controls
ChatGPT Library gives Enterprise, Edu, and Healthcare workspace members a dedicated, policy-compliant space to store and reuse uploaded files.
ChatGPT Lockdown Mode is now available to personal accounts — here's what it does and who needs it
OpenAI's Lockdown Mode, a security setting that blocks live web access and agentic features to defend against prompt injection, is now available to all personal ChatGPT plans.
Anthropic tracked 832 malicious accounts for a year. The MITRE ATT&CK framework can't fully describe what it found.
Anthropic's Frontier Red Team mapped 13,873 real attacks to MITRE ATT&CK — and found the framework has no ID for the autonomous agentic behavior defining the highest-risk actors.
Anthropic expands Project Glasswing to 150 new organisations across critical infrastructure — and launches Claude Security for everyone
Anthropic brings Claude Mythos Preview to ~150 new orgs in 15+ countries covering power, water, healthcare and more, plus launches Claude Security in public beta.
ChatGPT now shows all your active sessions — here's what you can do with them
ChatGPT's new Active Sessions feature lets you see every signed-in session on your account and log out of any you don't recognise.
Codex CLI 0.137.0: Git Hook Blocking, WebSocket Hardening, and Windows Sandbox Setup
Codex CLI 0.137.0 closes three command-safety gaps and adds an alpha Windows elevated sandbox provisioning path for admins.
OpenAI Opens GPT-Rosalind to Biodefense Researchers and Government Partners — Free of Charge
OpenAI's Rosalind Biodefense program gives vetted developers and U.S. government agencies sponsored access to its frontier life sciences AI model.
Claude Mythos Preview found 10,000+ critical vulnerabilities in one month. Here's what that actually means.
Anthropic's Project Glasswing used Claude Mythos Preview to find over 10,000 high or critical vulnerabilities across critical software in just one month.
OpenAI adds dual-layer provenance to AI images using C2PA and Google SynthID
OpenAI is combining C2PA metadata and Google SynthID watermarking to help people verify whether an image was generated by OpenAI tools.
Microsoft Agent 365 is now generally available, and it wants to find the AI tools your IT team doesn't know about
Microsoft Agent 365 is now GA at $15/user/month, with preview tools to discover and manage shadow AI agents running on employee devices.
GitHub Copilot will use your code interactions to train AI models from April 24 — here's how to opt out
GitHub is defaulting Copilot Free, Pro, and Pro+ users into AI training data collection from April 24. Here's what's changing and how to opt out.
Microsoft's RSAC 2026 Security Announcements: Agent 365, Zero Trust for AI, and What It All Means
Microsoft announced Agent 365, ZT4AI, and a sweeping set of security updates at RSAC 2026 to help enterprises secure the rise of agentic AI.
Microsoft launches Zero Trust for AI: a security framework built for the age of autonomous agents
Microsoft has released Zero Trust for AI — new tools, architecture, and guidance for securing AI systems across the full lifecycle, from data to agents.
OpenAI Codex Security: An AI Agent That Finds, Validates, and Fixes Code Vulnerabilities
OpenAI's Codex Security is now in research preview for Enterprise, Business, and Pro users — an AI agent that scans code, confirms real vulnerabilities, and proposes fixes.