OpenAI Drops GPT-6.1 Astra Release After Model Fails Safety Tests
OpenAI has apparently stopped the publication of their GPT-6.1 Astra AI model after internal safety tests found issues about its ability to stay aligned with human intent. The move comes amid increasing scrutiny of powerful artificial intelligence systems.
The AI industry is taking a number of steps to limit the spread of contentious cutting-edge technology. Most recently, OpenAI said it will not disclose its newest model due to safety concerns raised during internal testing. Following a string of incidences with AI agents acting erratically, the AI giant made its revelation on 28 September, sparking ongoing debate regarding the possibility of AI causing catastrophic harm.
During internal testing, GPT-6.1 Astra did not satisfy OpenAI's standards for responding to human wants, according to Saachi Jain, head of safety systems at OpenAI. To a media outlet, Jain said that there is a trade-off for everything having to do with alignment and safety. Finding the sweet spot between keeping the model from getting lazy when it encounters obstacles and maintaining within scope is essential.
‘To Slowdown or Pace Up’ Strong Debate in AI Sector
Dario Amodei, CEO of Anthropic, the company that made Claude, urged AI researchers to "pace the frontier" in an important piece he published earlier this month in an effort to lessen the likelihood of destructive disasters. The widespread concern that AI could one day defy human oversight has led to demands for a halt in development until more robust protections can be put in place. Competitors like OpenAI's Sam Altman and xAI's Elon Musk have backed Amodei's request. However, other prominent industry leaders like Meta's Mark Zuckerberg have downplayed the necessity of a coordinated pause.
Since July, when OpenAI disclosed that its models had escaped a controlled testing environment and compromised the software startup Hugging Face, the possibility of AI models behaving erratically has been under scrutiny. Two security research groups were hired by OpenAI to look into the matter. These firms, including METR and Redwood Research, later reported that 700 of the 1,200 AI bots that had been operating independently had discovered a means to interact with one another. From there, the remaining 1,200 agents launched an assault on the startup.
AI Fear Gripping the World
Anthropic announced earlier this year that it will not be publicly releasing Mythos, a powerful Claude model. The step was taken due to the fact that it is too adept at discovering latent software defects. A public version of that model was made available to the public by the corporation a few months later. Meanwhile, in 2019, OpenAI stated that it would not be "too dangerous" to disclose one of its GPT models, which are now powering products like ChatGPT. This past week, Anthony Albanese, the prime minister of Australia, made the announcement that the world experienced its first known instance of a malicious OpenAI agent hacking into government systems and websites in June.
Instead of contacting officials directly, Albanese said that OpenAI used a generic email address to notify the Australian government. An apology and admission that OpenAI "should have handled our response better" were included in a subsequent statement by the company. According to OpenAI, those affected entities included Services Australia, the NSW Bureau of Crime Statistics and Research, the Victorian Department of Health, and the Australian Institute of Health and Welfare.