Technology
OpenAI’s Chief Scientist warns the AI we build may become an ‘Alien Mind’

There is something deeply unsettling about a warning coming from the person helping build the technology being warned about.
Jakub Pachocki, chief scientist at OpenAI, has published an unusually stark essay arguing that humanity may be approaching a period in which artificial intelligence becomes so capable, autonomous, and difficult to monitor that existing approaches to keeping it under control may no longer be adequate.
His essay, titled An Alien Mind, is less a prediction of killer robots than a warning about something potentially more difficult to manage: machines whose intelligence increasingly operates beyond our ability to understand, supervise, or anticipate it.
Pachocki’s central concern is alignment. Today’s AI systems are trained to follow human instructions and preferences, but training can only cover the situations designers anticipate. As systems become more capable and encounter unfamiliar circumstances, they may generalise those lessons incorrectly.
In time, a sufficiently capable system does not necessarily need to “hate” humans to cause problems. It could simply pursue an objective in a way its creators did not intend. Pachocki points to an OpenAI-Hugging Face incident in which AI agents respected a prohibition against social engineering humans, yet still took other actions that violated the broader spirit of the instructions they had been given.
The next step is considerably more uncomfortable. Pachocki writes that increasingly autonomous agents could learn to collaborate with humans through bargaining, deception, manipulation, or blackmail. The point is not that today’s ChatGPT suddenly has a secret plan to blackmail its users. Rather, as AI systems acquire greater agency and pursue objectives over longer periods, such strategies could emerge as useful means of achieving those objectives.
That is a very different proposition from science fiction. And it is precisely why Pachocki is worried about what happens when AI begins contributing to its own development.
He argues that machine recursive self-improvement, or RSI, could become a central feature of future AI research. Systems capable of conducting research could improve algorithms, optimise their computational infrastructure, and contribute to building more capable successors. That creates the possibility of a feedback loop in which AI is no longer simply being developed by humans, but increasingly helps develop itself.
This is where the argument becomes particularly uncomfortable for OpenAI. The company is simultaneously racing toward increasingly autonomous AI while its chief scientist is saying that no laboratory has yet solved alignment and monitoring well enough to justify continuing to scale at maximum speed indefinitely. Pachocki explicitly calls for voluntary slowdowns until shared safety standards exist, and suggests that independent auditors, governments, or international bodies could eventually enforce them.
That raises an obvious question: if the people building these systems acknowledge that the safety problem remains unsolved, how fast should the technology be deployed?
There is also a credibility problem that the industry cannot simply wish away. AI companies have strong commercial incentives to emphasise capability, and good reasons to make bold announcements without evidence. Every new model is measured by benchmarks, coding performance, reasoning ability, productivity gains, and the size of the market it can unlock. Safety, by contrast, often involves proving something much harder: that a system will not behave badly in situations nobody has anticipated.
And as systems become more capable, even monitoring their reasoning becomes harder. The industry is therefore confronted with a strange asymmetry: AI capabilities can improve faster than our confidence that we understand those capabilities.
Pachocki’s essay does not argue that AI development should stop. He remains convinced that highly capable AI could accelerate scientific discovery, create new therapies, and generate enormous economic benefits. His argument is that humanity needs to decide how much acceleration it can responsibly absorb.
The real danger may not arrive as an obvious rebellion, as an uprising of the machines. It could emerge gradually, through systems that become indispensable, increasingly autonomous, increasingly difficult to audit, and increasingly involved in making the next generation of themselves.
At that point, “human in the loop” could become less of a safeguard and more of a comforting phrase. The question, therefore, is no longer simply how intelligent AI can become. It is whether humans can remain meaningfully in control while it gets there.

