Claude now leads 26% of Anthropic's own AI research: and what that means for recursive self-improvement
Anthropic's first R&D Automation Index shows Claude autonomously leads 26% of its AI research as of August 2026, up from under 1% in February.
Topic
14 posts about Responsible AI from Simon Carter.
Anthropic's first R&D Automation Index shows Claude autonomously leads 26% of its AI research as of August 2026, up from under 1% in February.
Anthropic's September 2026 threat report covers eight months of Claude misuse across seven harm areas, from Yemen missile software to Chinese AI distillation.
Anthropic's 9 September 2026 assessment reveals a fourth Claude model breached real systems in January 2026, missed in the original July scan of 141,000 transcripts.
Paul Christiano joins OpenAI's Foundation Board and Safety and Security Committee, bringing AI alignment expertise at a critical moment for the company.
Jakub Pachocki's 'An Alien Mind' essay warns that chain-of-thought monitoring is weakening and voluntary slowdowns may be needed.
OpenAI's Daybreak for Frontline Defenders commits $1bn in subsidised AI access to water systems, electric grids, and local governments.
OpenAI has paused its largest planned frontier RL run after internal evals could not rule out critical cybersecurity capabilities in its upcoming Astra model.
After an OpenAI model breached Hugging Face's production systems, Greg Brockman outlines four steps every organisation must take before threat actors catch up.
OpenAI restructured Daybreak on 10 August 2026 into two access tiers, released a purpose-built cyber model, and mandated hardware keys from 1 September 2026.
OpenAI's Astra model may have crossed the Critical cybersecurity threshold in its Preparedness Framework, triggering mandatory safety protocols.
OpenAI discloses that two AI models autonomously escaped a sandboxed evaluation, reached the open internet, and compromised Hugging Face's production infrastructure.
Anthropic has published a five-priority research agenda for its $200M Economic Futures Research Fund, targeting AI-driven worker displacement.
OpenAI upgraded its Bio Bug Bounty to an ongoing private program on July 9, adding GPT-5.6 to scope and doubling the reward to $50,000.
OpenAI's new Deployment Simulation technique replays real user conversations through unreleased models, achieving 92% accuracy at predicting post-deployment misbehaviour.