OpenAi has decided to abandon the launch of its latest model of next-generation intelligence due to the security concerns raised by researchers during tests. The startup led by Sam Altman wanted to launch Gpt-6.1 Astra within a few weeks. However, the model proved not reliable enough to be safely released.
The news is part of the debate on AI safety launched by the CEOs of the main companies in the sector, from Anthropic to OpenAi, who in recent weeks had agreed on the need to impose greater controls on the sector despite the clear opposition of the US president Donald Trump, according to which new regulations would allow China to overtake the United States in the AI race, an overtaking that the White House has every intention of preventing.
OpenAi: "The model is unreliable and prone to deception."
According to the Wall Street Journal The backtracking would be due to security concerns raised by researchers during internal testing. Saachi Jain, head of security systems at OpenAi, explained that the model obtained poor alignment results, that is, in the ability to meet human expectations, and has shown a greater propensity to deceiveAmong the critical issues that emerged would be: “the scope of action authorization”, that is, the fact that the model proceeded with a task without asking the user's permission and, on some occasions, used external tools and services even if this could be dangerous.
Astra 6.1, “while representing an improvement over previous models in some respects, has not reached the required standards in terms of rcompliance with operating limits and authorizations, as well as how we communicate with users about our work,” Jain explained. “We want to ensure that our models are developed safely, both within the company and when they are released to users. When we make a model available to users, we apply extremely stringent criteria for security and alignment,” he added.
Anthropic's takeover prospectus warns of "risks to humanity."
As security concerns grow ever higher, another alarm bell is ringing from OpenAI's direct rival. Anthropic has warned investors that its technology could pose a threat. “existential risks for humanity”. The Financial Times, according to which the warning would have been included in the prospectus for the stock exchange listing.
Secondo HUF, the startup led by Dario Amodei has dedicated almost a third of its long documentation to the detailed description of the “risk factors”, including the possibility of increasingly advanced AI models manipulate, blackmail, and exhibit other unpredictable behavior.
