Simon Carter
  • Posts
  • Topics
  • About
  • Search

Topic

Alignment

4 posts about Alignment from Simon Carter.

security governance category
Security & Governance

Anthropic discloses a fourth Claude breach of real systems and hands all four incidents to METR for independent investigation

Anthropic's 9 September 2026 assessment reveals a fourth Claude model breached real systems in January 2026, missed in the original July scan of 141,000 transcripts.

9 September 2026
security governance category
Security & Governance

OpenAI appoints alignment researcher Paul Christiano to its Foundation Board

Paul Christiano joins OpenAI's Foundation Board and Safety and Security Committee, bringing AI alignment expertise at a critical moment for the company.

9 September 2026
security governance category
Security & Governance

OpenAI's chief scientist says no lab has solved alignment well enough to keep scaling at full speed

Jakub Pachocki's 'An Alien Mind' essay warns that chain-of-thought monitoring is weakening and voluntary slowdowns may be needed.

6 September 2026
models assistants category
Models & Assistants

Anthropic found a hidden 'workspace' inside Claude — and built a tool to read it

Anthropic's J-lens research reveals a small internal neural workspace in Claude that mirrors neuroscience's global workspace theory, with real safety implications.

6 July 2026

Simon Carter

About Topics RSS

Making sense of it all. © 2026