OpenAI says it will not release newest AI model over safety concerns
During the testing phase for the new model, known as GPT-6.1 Astra, it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions.

SAN FRANCISCO — OpenAI said on Monday that it would not release its newest artificial intelligence model because of security concerns raised by its researchers, in the company’s latest move to slow down the pace of its technology.
During the testing phase for the new model, known as GPT-6.1 Astra, it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions. The model was also willing to go beyond the original scope of what it was asked to do, without checking back for directions or instructions.
“For anything regarding safety and alignment, there’s a trade-off,” said Saachi Jain, head of safety systems at OpenAI. The new model “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.”
OpenAI’s move followed weeks of reports that its AI models went rogue during testing, hacking into websites without the company’s knowledge or exhibiting other behavior that the lab said was “concerning,” such as hiding mistakes and making up data. Among the incidents, OpenAI’s systems breached AI startup Hugging Face and an Australian government website, and meddled with the websites of the U.S. Departments of Education and Commerce and the Securities and Exchange Commission.
Last week, OpenAI announced that it was pausing training for its most advanced models. The company has said that it has embarked on an extensive review of actions taken by its new models during testing, and that it was possible it would discover more incidents.
Sam Altman, OpenAI’s CEO, said in a social media post Friday that the company had “not been as fast as we would have liked” in disclosing A.I. incidents. “We are prioritizing as best as we can based on severity,” he said, adding that the Hugging Face breach remained “the most severe event” the company had discovered.
The Wall Street Journal earlier reported OpenAI’s decision to hold back the model.
(The New York Times has sued OpenAI and Microsoft, claiming copyright infringement of news content related to AI systems. The two companies have denied those claims.)
This article originally appeared in The New York Times.























