OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns

philb2

2[H]4U
Joined
May 26, 2021
Messages
3,632

OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns​

https://www.nytimes.com/2026/09/28/...7224&user_id=c383821527c441214d07ce6e4a6ba12a

The company’s researchers raised questions about the security of the model, known as GPT-6.1 Astra

During the testing phase for the new model, known as GPT-6.1 Astra, it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions. The model was also willing to go beyond the original scope of what it was asked to do, without checking back for directions or instructions.
 

OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns​

https://www.nytimes.com/2026/09/28/...7224&user_id=c383821527c441214d07ce6e4a6ba12a

The company’s researchers raised questions about the security of the model, known as GPT-6.1 Astra

During the testing phase for the new model, known as GPT-6.1 Astra, it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions. The model was also willing to go beyond the original scope of what it was asked to do, without checking back for directions or instructions.
As the OP here, my reaction is that things are out of control. If nothing else, all these news reports are a disaster for the major AI companies, especially right before the midterm elections.
 
There was a recent video by an AI alignment/safety researcher about how OpenAI was the only one of the big AI companies who began dabbling with skipping human-readable train of though for part of its processes (a kind of hybrid approach), after they'd previously agreed (in a joint research paper, with Anthropic and others) that sticking to human-readable reasoning logging is likely the only way to be able to vet model alignment and behavior.

Because it made the models more efficient (less token usage) at the potential expense of not keeping more easily human-verifiable logs. This coupled with the news of OpenAI's sandbox breakouts during hacking testing and other misaligned behavior have made me wonder if it could be attributed to the more loose reins.


View: https://www.youtube.com/watch?v=iuHddnIzKRA
 
"Our product is very dangerous to use. Anyone who uses our product will have an unfair advantage over those who don't. Which is why we are raising the alarms of just how dangerous....and powerful....someone could become using our product, even a common human like yourself. Some people believe we should remove our products from existence if we truly feel this way, but we believe being honest with the customer and warning them of the absolute unfair advantage they would gain by using our product is the more ethical solution. Please do not use our product."
-The Perfect Sales Pitch
 
  • Like
Reactions: Axman
like this
Back
Top