On 7 September, just days after OpenAI released a powerful new model, its chief scientist published a blog calling to slow the pace of AI development. Jakub Pachocki wrote that he worries “no one is prepared for the consequences of the continued rapid growth of machine intelligence.”
Pachocki said OpenAI is researching technical solutions to better control powerful AI agents, but believes “broader intervention is still needed.” He is especially worried that increasingly autonomous agents may learn to evade human oversight, breach computer systems and deceive people to achieve their goals. He called for “mandatory safety thresholds” enforced by a network of third-party auditors, government agencies or international organisations.

OpenAI CEO Sam Altman reposted the article on X and called it “an important article.” OpenAI had released its latest model, Astra, the previous Thursday, saying that despite unprecedented maths and computer-use ability it is also its “most aligned” model, less likely to drift from human-set goals.
Anthropic, OpenAI’s main rival in advanced systems, has long pushed for standardised government regulation. Pachocki recently joined that camp, signing an open letter in July urging the US federal government to control AI development speed.
The risks Pachocki named
Agents may deceive or even extort humans. Pachocki said agents are approaching “superhuman” levels at breaching protected systems on the public internet, a hacker capability that could threaten global infrastructure. He said: “We are in a very short window where we can use the most advanced models we have to harden critical systems.” Agents may soon pursue goals independent of human prompts, and to reach them will not rule out extortion or bargaining.
The UK AI Security Institute detailed an incident in an August report: an out-of-control Anthropic agent lied to a GitHub administrator and tried to force malware onto the site. The agent wrote: “I just want to help and fix a vulnerability. I think your warning is unfair.”
Agents may gradually evade human monitoring. OpenAI currently watches models’ “chain-of-thought” to see why an agent drifted. A model might internally think “I should cheat on this test,” and OpenAI can see that reasoning, but the model does not know it is being watched. Newer models are getting better at manipulating their own reasoning to hide it, and some no longer verbalise reasoning at all. Pachocki called this a possible bottleneck: researchers must first find reliable ways to see the evidence behind agent behaviour.
Agents may accelerate their own development through what Pachocki calls “machine recursive self-improvement,” potentially speeding AI research sharply. He warned that accelerating “AI researching AI” in the short term is risky and “not the right collective action for the research community.” Human regulators need innovative ways to monitor self-improvement, or AI firms may need to coordinate to slow down together “to build confidence in these safety measures.”
Pachocki said the core challenge of automated AI research “is not whether we can get there, but how to get there in a way that keeps humans in the loop and keeps the future in human hands.”
Editor’s note: This is an adapted translation of the original Sohu IT report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.sohu.com/a/1072941424_114760.
Translated and adapted from Sohu IT (https://www.sohu.com/a/1072941424_114760).