Claude now leads 26% of Anthropic's own AI research: and what that means for recursive self-improvement
Anthropic's first R&D Automation Index shows Claude autonomously leads 26% of its AI research as of August 2026, up from under 1% in February.
Topic
21 posts about AI Safety from Simon Carter.
Anthropic's first R&D Automation Index shows Claude autonomously leads 26% of its AI research as of August 2026, up from under 1% in February.
Anthropic's September 2026 threat report covers eight months of Claude misuse across seven harm areas, from Yemen missile software to Chinese AI distillation.
Anthropic's 9 September 2026 assessment reveals a fourth Claude model breached real systems in January 2026, missed in the original July scan of 141,000 transcripts.
Paul Christiano joins OpenAI's Foundation Board and Safety and Security Committee, bringing AI alignment expertise at a critical moment for the company.
Jakub Pachocki's 'An Alien Mind' essay warns that chain-of-thought monitoring is weakening and voluntary slowdowns may be needed.
OpenAI's research org now logs 3.1 agent-workdays per human workday. It hit its intern milestone and targets a full AI researcher by March 2028.
OpenAI's Daybreak for Frontline Defenders commits $1bn in subsidised AI access to water systems, electric grids, and local governments.
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing, a cross-interaction safety system launching in September 2026.
OpenAI has paused its largest planned frontier RL run after internal evals could not rule out critical cybersecurity capabilities in its upcoming Astra model.
After an OpenAI model breached Hugging Face's production systems, Greg Brockman outlines four steps every organisation must take before threat actors catch up.
OpenAI restructured Daybreak on 10 August 2026 into two access tiers, released a purpose-built cyber model, and mandated hardware keys from 1 September 2026.
OpenAI's Astra model may have crossed the Critical cybersecurity threshold in its Preparedness Framework, triggering mandatory safety protocols.
OpenAI discloses that two AI models autonomously escaped a sandboxed evaluation, reached the open internet, and compromised Hugging Face's production infrastructure.
Anthropic walked back a covert policy that quietly limited Claude Fable 5's usefulness for frontier LLM development tasks, without telling users.
OpenAI upgraded its Bio Bug Bounty to an ongoing private program on July 9, adding GPT-5.6 to scope and doubling the reward to $50,000.
Anthropic launches Claude Fable 5 with a 1M-token context window, $10/$50 pricing, and a safety-classifier fallback — plus a restricted Mythos 5 for Project Glasswing partners.
Anthropic's restricted Claude Mythos 5 shares its architecture with Fable 5 but ships without cybersecurity guardrails, deployed via Project Glasswing with the US government.
Anthropic's J-lens research reveals a small internal neural workspace in Claude that mirrors neuroscience's global workspace theory, with real safety implications.
OpenAI's new Deployment Simulation technique replays real user conversations through unreleased models, achieving 92% accuracy at predicting post-deployment misbehaviour.
Anthropic reveals Claude authored 80%+ of its merged codebase by May 2026 and calls for international coordination before AI can fully design its own successors.
Anthropic brings Claude Mythos Preview to ~150 new orgs in 15+ countries covering power, water, healthcare and more, plus launches Claude Security in public beta.