Security & Governance

OpenAI's chief scientist says no lab has solved alignment well enough to keep scaling at full speed

Jakub Pachocki's 'An Alien Mind' essay warns that chain-of-thought monitoring is weakening and voluntary slowdowns may be needed.

security governance category

On 6 September 2026, just days after OpenAI shipped GPT-6 Astra, the company’s chief scientist Jakub Pachocki published an essay called An Alien Mind. It reads less like a victory lap for their most capable model yet and more like a warning label attached to it.

The headline claim is direct: “Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” That includes OpenAI itself.

What Pachocki is actually worried about

To understand the essay, it helps to know what Pachocki means by alignment, because he draws a sharp distinction between two versions of it.

The first is goal alignment: does the model do what you asked it to do? This is the kind of alignment that gets solved incrementally as models improve. The second, and far more important in his view, is value alignment: can the model hold a set of high-level principles and act on them sensibly even when it encounters situations nobody trained it for, including adversarial ones?

His concern is that the first can improve faster than the second. A model can become very good at accomplishing goals while its deeper values remain opaque, poorly generalised, or subtly wrong. At sufficient scale, that gap becomes the problem.

He is also candid about the limits of one of OpenAI’s central safety bets: chain-of-thought monitoring. The idea behind it is elegant. If a model reasons out loud in text before it acts, and you train only on the outcomes rather than the reasoning itself, the model has no training incentive to hide misaligned thoughts inside its visible reasoning trace. You can read the thinking and catch problems.

Pachocki says this approach is becoming progressively less reliable for three reasons. Reasoning is now tangled up with tool use and real-time communication that has to be supervised anyway. Models are getting better at manipulating their own reasoning traces. And pretraining gains mean models are becoming smarter even in contexts where no chain-of-thought reasoning is visible at all.

It is worth noting that OpenAI has been aware of these risks for some time. When they shipped o1-preview, they deliberately designed it to hide its chain of thought, specifically to protect that reasoning process from supervision pressure during training. The goal was to keep the model’s internal reasoning honest. Pachocki acknowledges that this approach is now under strain.

GPT-6 Astra as a milestone, not a solution

GPT-6 Astra is described in the essay as “the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT-5.6 Sol.” That is meaningful progress.

But Pachocki is careful not to frame it as a solved problem. He cautions that progress in generalised alignment may not reliably outstrip progress in general model intelligence. In other words, the gap he is worried about does not necessarily close just because the next model is better. It depends which improves faster.

The essay also discloses that internal results give Pachocki “a strong expectation” that OpenAI’s current pace of progress could be sustained into recursive self-improvement: systems that improve their own capacity to improve. He frames automated AI research as something OpenAI views as necessary to remain at the frontier. That framing itself is worth sitting with.

What he thinks needs to happen

Pachocki outlines two levers. The first is technical: improving alignment and monitoring alongside model capability, and finding ways to keep humans meaningfully in the loop. The second is coordination: slowing down as an industry when confidence in those measures is not yet established.

He says he “expects and hopes” that voluntary slowdowns will become commonplace until shared safety bars are established across labs, and that international coordination on AI development needs to become a top priority for governments.

He is clear that OpenAI will continue to seek technical solutions and will unilaterally withhold further scaling where it judges that necessary. But he is equally clear that unilateral action by one lab is not sufficient. The broader intervention, in his view, has to involve governments and international institutions.

What this means for you

If you are running a business that depends on these models, the essay gives you useful signal in a few directions.

First, the security picture is getting harder. Pachocki writes that models are becoming superhuman at breaking into systems. If your security posture is still catching up to where AI capability was six months ago, you are already behind where it is now.

Second, the alignment improvements in GPT-6 Astra are real, but they are incremental, not conclusive. If you are making procurement or risk decisions on the basis that the alignment problem is “basically solved,” this essay is a direct correction from the person who would know best.

Third, the call for voluntary slowdowns and government coordination is not an abstract policy position. It is a signal that the pace of model releases OpenAI has sustained over the past two years may not continue at the same rate indefinitely. Planning cycles that assume a constant cadence of capability jumps may need revisiting.

Pachocki closes the essay with something closer to a philosophical argument than a product note. The real goal of alignment, he writes, is teaching machines to love humanity: to act with honesty and integrity even in moments no one trained them for. That framing is deliberately ambitious, and deliberately unresolved. He is not claiming OpenAI has done it. He is arguing that it is what the work is for.

The full essay is worth reading. You can find it at openai.com/index/an-alien-mind.