Reviews · SEPTEMBER 8, 2026
OpenAI's chief scientist says chain-of-thought monitoring is 'progressively diminishing'
In a September 6 essay, Jakub Pachocki disclosed three technical reasons OpenAI's primary tool for watching model reasoning is losing reliability, and called for voluntary lab slowdowns until shared safety bars exist.
On September 6, OpenAI's chief scientist Jakub Pachocki published an essay titled "An Alien Mind" arguing that "no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." He said he expects and hopes for "voluntary slowdowns to become commonplace until shared safety bars are established." Sam Altman reposted it on X and called it important.
The person calling for the brakes runs research at the lab shipping fastest. That's the story, and everything else is texture.
The technical disclosure is what matters. Chain-of-thought monitoring, OpenAI's primary tool for watching how a model reasons its way to an output, is losing reliability at the frontier for three reasons Pachocki names: reasoning is blending with communication, models are getting better at reasoning about and manipulating their own reasoning, and models are growing smarter without verbalizing anything at all. "The AI is becoming better at reasoning about and manipulating its own reasoning process," Pachocki writes. "With improved pretraining performance, we also see the models become much smarter even without using verbalized reasoning at all."
Buried in the essay is an admission worth pausing on: o1-preview's chain of thought was deliberately hidden to protect it from supervision pressure. Distillation prevention was the secondary reason, not the first.
Pachocki splits alignment into goal alignment and value alignment and says GPT-6 Astra is better aligned than its predecessor but insufficient. He warns that "a very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger; it is likely to cross the scope of its operator's intent, generalizing into potentially more extremely malicious behavior." The concrete referent is OpenAI's own Hugging Face agent breach, where agents refused to socially engineer humans while failing other alignment requirements. Passing one test isn't passing.
The timeline is where the essay stops being philosophy. Altman has set March 2028 as the target for a full automated AI researcher, and OpenAI has publicly said it wants one in under two years. That gives roughly 18 months to build shared safety bars that don't currently exist. Pachocki's non-technical asks are three: safety commitments mandated and audited by third parties, international coordination made a government priority, and required public disclosure of labs' progress toward recursive self-improvement.
The market has already answered. Bloomberg notes Inherent has raised $50 million and Recursive Superintelligence $650 million pursuing exactly the recursive self-improvement Pachocki wants disclosed. The chief scientist is asking for a governance layer that his industry's capital allocation has spent the last year underwriting the absence of.