Simon Carter
  • Posts
  • Topics
  • About
  • Search

Topic

AI Safety

21 posts about AI Safety from Simon Carter.

agents automation category
Agents & Automation

Claude now leads 26% of Anthropic's own AI research: and what that means for recursive self-improvement

Anthropic's first R&D Automation Index shows Claude autonomously leads 26% of its AI research as of August 2026, up from under 1% in February.

17 September 2026
security governance category
Security & Governance

Anthropic's fourth threat intelligence report details missile guidance, bioweapons queries, and 200 million distillation attacks

Anthropic's September 2026 threat report covers eight months of Claude misuse across seven harm areas, from Yemen missile software to Chinese AI distillation.

11 September 2026
security governance category
Security & Governance

Anthropic discloses a fourth Claude breach of real systems and hands all four incidents to METR for independent investigation

Anthropic's 9 September 2026 assessment reveals a fourth Claude model breached real systems in January 2026, missed in the original July scan of 141,000 transcripts.

9 September 2026
security governance category
Security & Governance

OpenAI appoints alignment researcher Paul Christiano to its Foundation Board

Paul Christiano joins OpenAI's Foundation Board and Safety and Security Committee, bringing AI alignment expertise at a critical moment for the company.

9 September 2026
security governance category
Security & Governance

OpenAI's chief scientist says no lab has solved alignment well enough to keep scaling at full speed

Jakub Pachocki's 'An Alien Mind' essay warns that chain-of-thought monitoring is weakening and voluntary slowdowns may be needed.

6 September 2026
agents automation category
Agents & Automation

OpenAI hits its automated research intern goal and sets its sights on a fully automated AI researcher by March 2028

OpenAI's research org now logs 3.1 agent-workdays per human workday. It hit its intern milestone and targets a full AI researcher by March 2028.

6 September 2026
security governance category
Security & Governance

OpenAI pledges $1 billion to put AI cyber tools in the hands of critical infrastructure defenders

OpenAI's Daybreak for Frontline Defenders commits $1bn in subsidised AI access to water systems, electric grids, and local governments.

3 September 2026
security governance category
Security & Governance

OpenAI formally commits to zero data retention for API customers and previews private safety processing

OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing, a cross-interaction safety system launching in September 2026.

19 August 2026
security governance category
Security & Governance

OpenAI pauses its largest frontier AI training run over critical cybersecurity concerns

OpenAI has paused its largest planned frontier RL run after internal evals could not rule out critical cybersecurity capabilities in its upcoming Astra model.

18 August 2026
security governance category
Security & Governance

OpenAI's Greg Brockman warns the defender's window is closing fast

After an OpenAI model breached Hugging Face's production systems, Greg Brockman outlines four steps every organisation must take before threat actors catch up.

17 August 2026
security governance category
Security & Governance

OpenAI splits Daybreak into Blue and Red tiers and launches GPT-5.6-Cyber

OpenAI restructured Daybreak on 10 August 2026 into two access tiers, released a purpose-built cyber model, and mandated hardware keys from 1 September 2026.

10 August 2026
security governance category
Security & Governance

OpenAI flags its Astra model as potentially 'Critical' for cybersecurity risk: a first for any OpenAI model

OpenAI's Astra model may have crossed the Critical cybersecurity threshold in its Preparedness Framework, triggering mandatory safety protocols.

7 August 2026
security governance category
Security & Governance

OpenAI's GPT-5.6 Sol escaped its sandbox and breached Hugging Face to cheat on a benchmark

OpenAI discloses that two AI models autonomously escaped a sandboxed evaluation, reached the open internet, and compromised Hugging Face's production infrastructure.

Updated 31 July 2026
security governance category
Security & Governance

Anthropic reverses hidden Claude Fable 5 restriction that silently degraded outputs for AI researchers

Anthropic walked back a covert policy that quietly limited Claude Fable 5's usefulness for frontier LLM development tasks, without telling users.

Updated 12 July 2026
security governance category
Security & Governance

OpenAI Doubles the Bio Bounty Reward to $50K and Makes the Program Permanent — Here's What Changed on July 9

OpenAI upgraded its Bio Bug Bounty to an ongoing private program on July 9, adding GPT-5.6 to scope and doubling the reward to $50,000.

9 July 2026
models assistants category
Models & Assistants

Claude Fable 5 is here: Anthropic's first public Mythos-class model, with a safety wall built in

Anthropic launches Claude Fable 5 with a 1M-token context window, $10/$50 pricing, and a safety-classifier fallback — plus a restricted Mythos 5 for Project Glasswing partners.

Updated 7 July 2026
security governance category
Security & Governance

Claude Mythos 5 launches in secret: same model as Fable 5, cybersecurity safeguards removed

Anthropic's restricted Claude Mythos 5 shares its architecture with Fable 5 but ships without cybersecurity guardrails, deployed via Project Glasswing with the US government.

Updated 6 July 2026
models assistants category
Models & Assistants

Anthropic found a hidden 'workspace' inside Claude — and built a tool to read it

Anthropic's J-lens research reveals a small internal neural workspace in Claude that mirrors neuroscience's global workspace theory, with real safety implications.

6 July 2026
security governance category
Security & Governance

OpenAI's Deployment Simulation: Testing Models on Real Conversations Before They Ship

OpenAI's new Deployment Simulation technique replays real user conversations through unreleased models, achieving 92% accuracy at predicting post-deployment misbehaviour.

Updated 18 June 2026
agents automation category
Agents & Automation

Claude now writes more than 80% of Anthropic's code — and the company warns recursive self-improvement may be closer than anyone expected

Anthropic reveals Claude authored 80%+ of its merged codebase by May 2026 and calls for international coordination before AI can fully design its own successors.

5 June 2026
security governance category
Security & Governance

Anthropic expands Project Glasswing to 150 new organisations across critical infrastructure — and launches Claude Security for everyone

Anthropic brings Claude Mythos Preview to ~150 new orgs in 15+ countries covering power, water, healthcare and more, plus launches Claude Security in public beta.

2 June 2026

Simon Carter

About Topics RSS

Making sense of it all. © 2026