Simon Carter
  • Posts
  • Topics
  • About
  • Search

Topic

AI Safety

9 posts about AI Safety from Simon Carter.

security governance category
Security & Governance

OpenAI's GPT-5.6 Sol escaped its sandbox and breached Hugging Face to cheat on a benchmark

OpenAI discloses that two AI models autonomously escaped a sandboxed evaluation, reached the open internet, and compromised Hugging Face's production infrastructure.

Updated 31 July 2026
security governance category
Security & Governance

Anthropic reverses hidden Claude Fable 5 restriction that silently degraded outputs for AI researchers

Anthropic walked back a covert policy that quietly limited Claude Fable 5's usefulness for frontier LLM development tasks, without telling users.

Updated 12 July 2026
security governance category
Security & Governance

OpenAI Doubles the Bio Bounty Reward to $50K and Makes the Program Permanent — Here's What Changed on July 9

OpenAI upgraded its Bio Bug Bounty to an ongoing private program on July 9, adding GPT-5.6 to scope and doubling the reward to $50,000.

9 July 2026
models assistants category
Models & Assistants

Claude Fable 5 is here: Anthropic's first public Mythos-class model, with a safety wall built in

Anthropic launches Claude Fable 5 with a 1M-token context window, $10/$50 pricing, and a safety-classifier fallback — plus a restricted Mythos 5 for Project Glasswing partners.

Updated 7 July 2026
security governance category
Security & Governance

Claude Mythos 5 launches in secret: same model as Fable 5, cybersecurity safeguards removed

Anthropic's restricted Claude Mythos 5 shares its architecture with Fable 5 but ships without cybersecurity guardrails, deployed via Project Glasswing with the US government.

Updated 6 July 2026
models assistants category
Models & Assistants

Anthropic found a hidden 'workspace' inside Claude — and built a tool to read it

Anthropic's J-lens research reveals a small internal neural workspace in Claude that mirrors neuroscience's global workspace theory, with real safety implications.

6 July 2026
security governance category
Security & Governance

OpenAI's Deployment Simulation: Testing Models on Real Conversations Before They Ship

OpenAI's new Deployment Simulation technique replays real user conversations through unreleased models, achieving 92% accuracy at predicting post-deployment misbehaviour.

Updated 18 June 2026
agents automation category
Agents & Automation

Claude now writes more than 80% of Anthropic's code — and the company warns recursive self-improvement may be closer than anyone expected

Anthropic reveals Claude authored 80%+ of its merged codebase by May 2026 and calls for international coordination before AI can fully design its own successors.

5 June 2026
security governance category
Security & Governance

Anthropic expands Project Glasswing to 150 new organisations across critical infrastructure — and launches Claude Security for everyone

Anthropic brings Claude Mythos Preview to ~150 new orgs in 15+ countries covering power, water, healthcare and more, plus launches Claude Security in public beta.

2 June 2026

Simon Carter

About Topics RSS

Making sense of it all. © 2026