OpenAI’s Astra model has AI researchers spooked. Here’s why

0



Arrived just days after the launch of Anthropic’s Claude Fable 5.1 and Mythos 5.1, Astra marks a “jump” in AI capabilities, OpenAI president Greg Brockman said, boasting that the new model “can really do anything a human can do with a computer.”

Astra is also OpenAI’s first model to reach the “critical” threshold of the company’s “preparedness framework” due to its extreme cybersecurity skills, meaning it could carry out “end-to-end” attacks on “hardened targets” on its own, among other capabilities. 

OpenAI previously paused work on Astra to bolster its safeguards before announcing earlier this week that the model is “consistently more likely to respect explicit safety restrictions and warnings” than GPT-5.6 Sol, the OpenAI model involved in the now infamous Hugging Face attack.

Despite OpenAI’s assurances, AI experts remain worried about Astra. The new model is said to employ a reasoning technique known variously as “recurrent depth” or “opaque recurrence,” which (as TechCrunch describes) makes its “chain of thought” much harder to read.

Keeping tabs on a frontier AI model’s thinking is, obviously, a big deal when it comes to preventing the kinds of rogue AI hacks we’ve been hearing about over the past several weeks, and the potential of losing that kind of surveillance has spooked top AI researchers.

“If this is true, OpenAI seems to be violating one of the few redlines that exist in the AI community,” wrote Steven Adler, a former OpenAI safety lead, on X.



Source link