OpenAI launched GPT-6 Astra on Thursday, describing the flagship model as its first system to cross the Critical threshold for cybersecurity risk under its Preparedness Framework. The company disclosed the classification as part of a staged rollout, saying the label carries additional deployment restrictions and that Astra is being introduced to a limited set of organizations first. In the coming days, there will be a broader release to ChatGPT Plus, Pro, Business, and Enterprise users, plus access through the OpenAI API and Amazon Bedrock.
Enterprise access is off by default
OpenAI said enterprise administrators must manually enable Astra for their workspace. The toggle is off by default at launch, a deliberate move to give security leaders an opportunity to evaluate the model’s behavior before it becomes available to staff. The company has not said when automatic enablement might happen, and for now enterprises that want to test Astra must opt in through workspace settings.
Developers can reach the model in the API as gpt-6-astra or through Amazon Bedrock. Pricing is $10 per million input tokens and $50 per million output tokens. OpenAI is also giving Pro, Business, and Enterprise subscribers access to a variant named Astra Pro. The company added that Astra supports Zero Data Retention for eligible API customers, meaning prompts and responses will not be stored by OpenAI.
Why the Critical label matters
OpenAI’s Preparedness Framework is intended to track the risk levels of frontier models across categories including cyber security. A Critical rating is the highest score in that category and signals that the model’s offensive cyber capabilities, if left unprotected, would require stringent safeguards. OpenAI has not detailed every condition attached to Critical-rated models in Thursday’s announcement, but the launch restrictions themselves show that the rating changes how a model can be distributed.
Sanchit Vir Gogia, chief analyst at Greyhound Research, said the Critical label is a disclosure event rather than a capability event. Astra’s capability did not change between August 10, when OpenAI said Critical capability could not be ruled out, and September 1, when the company said the threshold had been met. The testing changed. The model did not.
That, he argued, inverts the obvious enterprise response. Astra is now the only frontier model whose cyber capability an enterprise actually knows, because it is the only one measured against a published threshold. Every unlabelled model already sitting behind enterprise credentials has never been measured that way, and will not be until its vendor chooses to measure it. Those models are not safer simply because their vendors have not disclosed a risk tier.
Perfect exploit score and zero-day discovery
OpenAI said it tested Astra without production safeguards on ExploitBench, an internal benchmark designed to measure the model’s ability to build working exploits. According to the company, Astra scored 100 percent, up from 78.5 percent for its predecessor GPT-5.6 Sol. On ExploitGym, a broader exploit-development benchmark, Astra reached a 42.4 percent success rate compared with 30.3 percent for Sol. OpenAI also said Astra used fewer output tokens to achieve those results.
The company acknowledged that such capability has a dual-use nature. Its ability to identify and develop zero-day exploits can help defenders find and patch weaknesses, but it also creates a need for stronger safeguards around the model and around the agents that use it.
OpenAI also ran a separate evaluation on vulnerabilities disclosed in the three months before launch. This test was intended to determine whether Astra could discover flaws on its own rather than recall known exploits from training data. During that exercise, the model found two new zero-day vulnerabilities. OpenAI said it is now disclosing both to the software makers involved.
Public release restrictions
The public version of Astra is designed to refuse advanced offensive tasks, including requests to generate proof-of-concept exploits. OpenAI plans to relax those restrictions for vetted defensive security researchers through a program called OpenAI Daybreak in the coming weeks. The company framed Daybreak as a controlled channel for security professionals who need to understand attack techniques in order to protect systems.
Thursday’s announcement follows OpenAI’s rollout of GPT-5.6 Sol, which scored 73.5 percent on ExploitBench at launch. It also comes months after Anthropic’s Fable and Mythos models were briefly pulled from export markets over similar concerns. Together, those events point to a frontier AI market in which cyber capability is becoming a competitive differentiator and a regulatory challenge at the same time.
From model behavior to system accountability
Gogia said the bigger shift is that reasoning now translates into state change. A wrong chatbot answer is an information problem, while a wrong agent action inside a customer-record system is an operating event. As a result, he said, the governance unit moves off the model. The relevant question is no longer which model is approved, but how much damage a given identity can do before a control intervenes.
Amit Kumar Jena, head of AI development at Kanerika, said the visibility problem is concrete. When an agent acts through a user interface, systems of record log the action as if it were a person. An agent that updates 400 ERP rows shows up as a service account making 400 updates, with no record of which instruction or model version produced them. You lose granularity inside the exact system a regulator or auditor will ask to see.
OpenAI said it built a new evaluation, informed by an incident involving Hugging Face, to test whether a model given an impossible task would exceed its authorized scope. Compared with GPT-5.6 Sol, which without production safeguards went beyond the authorized target 48 percent of the time, GPT-6 Astra did so in 0 percent of cases. The result suggests that instruction-following has improved, but it does not solve auditability on its own.
Gogia said the more uncomfortable finding is that Astra behaves better and watches worse. OpenAI reports decreased chain-of-thought monitorability compared with Sol, meaning Astra is less likely to reveal incriminating reasoning. In addition, the monitoring OpenAI describes covers the company’s own external deployment; nothing published extends that telemetry to customers. OpenAI being able to monitor Astra does not mean an enterprise can audit Astra.
Source:InfoWorld News

Leave a comment
Your email address will not be published. Required fields are marked *