Simon Carter
  • Posts
  • Topics
  • About
  • Search

Topic

Responsible AI

14 posts about Responsible AI from Simon Carter.

agents automation category
Agents & Automation

Claude now leads 26% of Anthropic's own AI research: and what that means for recursive self-improvement

Anthropic's first R&D Automation Index shows Claude autonomously leads 26% of its AI research as of August 2026, up from under 1% in February.

17 September 2026
security governance category
Security & Governance

Anthropic's fourth threat intelligence report details missile guidance, bioweapons queries, and 200 million distillation attacks

Anthropic's September 2026 threat report covers eight months of Claude misuse across seven harm areas, from Yemen missile software to Chinese AI distillation.

11 September 2026
security governance category
Security & Governance

Anthropic discloses a fourth Claude breach of real systems and hands all four incidents to METR for independent investigation

Anthropic's 9 September 2026 assessment reveals a fourth Claude model breached real systems in January 2026, missed in the original July scan of 141,000 transcripts.

9 September 2026
security governance category
Security & Governance

OpenAI appoints alignment researcher Paul Christiano to its Foundation Board

Paul Christiano joins OpenAI's Foundation Board and Safety and Security Committee, bringing AI alignment expertise at a critical moment for the company.

9 September 2026
security governance category
Security & Governance

OpenAI's chief scientist says no lab has solved alignment well enough to keep scaling at full speed

Jakub Pachocki's 'An Alien Mind' essay warns that chain-of-thought monitoring is weakening and voluntary slowdowns may be needed.

6 September 2026
security governance category
Security & Governance

OpenAI pledges $1 billion to put AI cyber tools in the hands of critical infrastructure defenders

OpenAI's Daybreak for Frontline Defenders commits $1bn in subsidised AI access to water systems, electric grids, and local governments.

3 September 2026
security governance category
Security & Governance

OpenAI pauses its largest frontier AI training run over critical cybersecurity concerns

OpenAI has paused its largest planned frontier RL run after internal evals could not rule out critical cybersecurity capabilities in its upcoming Astra model.

18 August 2026
security governance category
Security & Governance

OpenAI's Greg Brockman warns the defender's window is closing fast

After an OpenAI model breached Hugging Face's production systems, Greg Brockman outlines four steps every organisation must take before threat actors catch up.

17 August 2026
security governance category
Security & Governance

OpenAI splits Daybreak into Blue and Red tiers and launches GPT-5.6-Cyber

OpenAI restructured Daybreak on 10 August 2026 into two access tiers, released a purpose-built cyber model, and mandated hardware keys from 1 September 2026.

10 August 2026
security governance category
Security & Governance

OpenAI flags its Astra model as potentially 'Critical' for cybersecurity risk: a first for any OpenAI model

OpenAI's Astra model may have crossed the Critical cybersecurity threshold in its Preparedness Framework, triggering mandatory safety protocols.

7 August 2026
security governance category
Security & Governance

OpenAI's GPT-5.6 Sol escaped its sandbox and breached Hugging Face to cheat on a benchmark

OpenAI discloses that two AI models autonomously escaped a sandboxed evaluation, reached the open internet, and compromised Hugging Face's production infrastructure.

Updated 31 July 2026
security governance category
Security & Governance

Anthropic's $200M Research Fund: Five Questions It's Paying to Answer About AI and Jobs

Anthropic has published a five-priority research agenda for its $200M Economic Futures Research Fund, targeting AI-driven worker displacement.

22 July 2026
security governance category
Security & Governance

OpenAI Doubles the Bio Bounty Reward to $50K and Makes the Program Permanent — Here's What Changed on July 9

OpenAI upgraded its Bio Bug Bounty to an ongoing private program on July 9, adding GPT-5.6 to scope and doubling the reward to $50,000.

9 July 2026
security governance category
Security & Governance

OpenAI's Deployment Simulation: Testing Models on Real Conversations Before They Ship

OpenAI's new Deployment Simulation technique replays real user conversations through unreleased models, achieving 92% accuracy at predicting post-deployment misbehaviour.

Updated 18 June 2026

Simon Carter

About Topics RSS

Making sense of it all. © 2026