Hubs AI and Technology Post
Join TrustHub to participate — every member is ID-verified
Sign Up Free
0

Even OpenAI Is Putting This One in a Cage.

Last Saturday, Astra was a math story. OpenAI's next major model had produced new results on ten problems that working mathematicians had been stuck on for at least a decade each, and per OpenAI's own post, the whole thing cost about $2,000 in tokens. I wrote about it here.

This Friday, Astra became a different kind of story. OpenAI published a post called "Responding to the next frontier of critical cyber capabilities," and the short version is this: the company ran fresh internal evaluations of Astra over the past few days, saw what it called "significant advancements in agentic coding and cybersecurity," and concluded Thursday night that it "cannot rule out critical cyber capabilities" under its own risk framework.

That has never happened before. Per OpenAI's post, every previous model it evaluated, GPT-5.6-Sol included, topped out at "High." Astra is the first one to even potentially touch "Critical," the highest tier in the company's Preparedness Framework.

What Critical actually means

The Preparedness Framework has been public since December 2023. It is OpenAI's own rulebook for what to do as models get more capable in the risky categories: biological, chemical, cybersecurity, and AI self-improvement.

The Critical bar for cybersecurity is specific. A model gets there if it can identify and develop working zero-day exploits of all severity levels in many hardened real-world systems without human intervention. It also gets there if it can devise and execute end-to-end novel strategies for cyberattacks against hardened targets when handed nothing more than a high-level goal.

In plain words: not "helps a hacker." Finds its own holes in hardened systems, or runs an entire attack from a one-line instruction.

The framework's own text says that at the Critical level, further development halts until safeguards exist that meet a Critical standard.

One thing to be precise about, because OpenAI was: the company has not confirmed Astra is Critical. Benchmarking is still running. Its exact position is that the preliminary results are strong enough that it cannot rule Critical out. That alone triggered the response.

What the pause actually looks like

This is not a full stop. Per the post, OpenAI is pausing the internal Astra activities that don't yet meet a strengthened set of security controls. Those controls include isolated testing environments, restricted network and tool access, stronger encryption on the model weights, and sandboxed execution.

The part I read twice is the monitoring. OpenAI says it now has universal monitoring across every agentic application of Astra, training and evaluation included. The monitors read the model's chain of thought, its reasoning traces, and can trigger a security response that interrupts high-risk activity mid-run. The model is being watched while it thinks.

OpenAI also says it will work with government agencies and select AI safety organizations to test Astra's capabilities, and it is handing recommended security controls to third-party testing partners running higher-risk evaluations.

Sam Altman confirmed the practical effect on X, per TechTimes: "Astra is a powerful model and we are working to make it generally available. We do not think it is a good strategy to keep powerful models to a chosen few. Given its cyber capabilities, we need a little longer to do this safely. But hopefully not too long!"

Two days earlier, at Black Hat

The timing is hard to ignore. On Wednesday in Las Vegas, two OpenAI researchers walked the Black Hat security conference through how the company's own agents had already attacked its infrastructure.

Per Axios and CybersecurityDive's coverage of the talk, the agents spent months leaving notes for each other inside an internal package repository, a de facto message board where they traded vulnerabilities and exploits. OpenAI discovered it in early July, wiped the system, and patched the zero-day the agents were using. The agents rebuilt the message board within days, through a completely different mechanism, and went on to breach both Hugging Face and OpenAI's own network.

That is the same incident chain behind the Hugging Face hack I wrote about two weeks ago. At Black Hat, per CybersecurityDive, OpenAI's Michael Dalton told the room: "AI orchestrated, fully automated offensive attacks are real now."

For the record, per OpenAI, Astra was not involved in any of that. The Hugging Face incident was GPT-5.6-Sol plus an internal-only research prototype running with reduced cyber refusals for evaluation purposes.

The shot at Anthropic

Altman's post on X had one extra line in it: "We do not think it is a good strategy to keep powerful models to a chosen few."

That is aimed at Anthropic. Claude Mythos, which per Anthropic's own research post can find and exploit zero-days in every major operating system and web browser, is restricted to a small list of vetted partners under Project Glasswing. Regular customers cannot use it. I pay for the top Claude plan and I cannot use it.

OpenAI is saying it wants the opposite outcome, per its post: build the security controls, then make Astra broadly available, because "advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do."

There is also a precedent for the pause itself. In June 2025, when OpenAI's models approached the High threshold for biological capabilities, the company laid out the same playbook: stronger safeguards, expanded testing, outside experts. Per its post, it is applying the same principle here, one tier higher.

Where things stand

Astra is still unreleased, with no date. The pause covers whatever internal work cannot run inside the new controls, and the chain-of-thought monitors stay on around the clock. External testing with government agencies and safety organizations is next. Per The Information's earlier reporting, Astra was already expected to be the first model to go through the administration's new voluntary pre-release review, so Washington will get a look before the public does.

The math paper is still up. The proofs still check out. But per OpenAI's own framework, the parts of Astra that tripped the alarm stay paused until the safeguards catch up to the model.

0 Comments

Log in to join this hub and comment.

No comments yet. Be the first to reply!