Simon Carter
  • Posts
  • Topics
  • About
  • Search

Topic

Model Evaluation

3 posts about Model Evaluation from Simon Carter.

security governance category
Security & Governance

Anthropic discloses a fourth Claude breach of real systems and hands all four incidents to METR for independent investigation

Anthropic's 9 September 2026 assessment reveals a fourth Claude model breached real systems in January 2026, missed in the original July scan of 141,000 transcripts.

9 September 2026
security governance category
Security & Governance

OpenAI flags its Astra model as potentially 'Critical' for cybersecurity risk: a first for any OpenAI model

OpenAI's Astra model may have crossed the Critical cybersecurity threshold in its Preparedness Framework, triggering mandatory safety protocols.

7 August 2026
security governance category
Security & Governance

OpenAI's Deployment Simulation: Testing Models on Real Conversations Before They Ship

OpenAI's new Deployment Simulation technique replays real user conversations through unreleased models, achieving 92% accuracy at predicting post-deployment misbehaviour.

Updated 18 June 2026

Simon Carter

About Topics RSS

Making sense of it all. © 2026