Connect with us
In focus Magazine June 2026 advertise

Technology

OpenAI’s Astra raises a bigger question: Who watches the watchers? 

karan Karayi PP

Published

on

OpenAI’s Astra Raises a Bigger Question: Who Watches AI?

OpenAI’s next major AI model, codenamed Astra, is arriving with an unusual warning attached to it: the company itself says the system is powerful enough to require stronger safety measures before it can be released. 

Also read: OpenAI Safety Team Disbanded, Raising AI Safety Concerns 

Astra has reportedly demonstrated cybersecurity capabilities beyond those of OpenAI’s existing public models. Under the company’s own Preparedness Framework, it has reached the “Critical” cybersecurity capability threshold. With the right tools and access, OpenAI says Astra can identify previously unknown security vulnerabilities and develop exploits across well-protected systems without requiring a human to guide every step.  

In other words, the concern is no longer hypothetical. An AI model is becoming capable of doing some of the work that highly skilled cyber attackers do today, potentially at greater speed, scale, and consistency. 

That creates an obvious dual-use problem. The same capability that could help defenders discover vulnerabilities before criminals do could also lower the barrier to launching sophisticated cyberattacks. Hospitals, banks, governments, industrial systems, cloud infrastructure, and critical utilities could all become potential targets if powerful models fall into the wrong hands. 

The timing makes the issue more uncomfortable. OpenAI has already had to slow aspects of Astra’s development after internal evaluations raised concerns about its agentic coding and cybersecurity abilities. The company has said it is strengthening monitoring, alignment, and security measures, and that Astra will initially be made available to a limited audience.  

Yet there is another development that deserves scrutiny: the apparent weakening and restructuring of some of the very institutional structures designed to think about catastrophic AI risk. 

Reports in August said OpenAI had disbanded its Preparedness team, which assessed whether models could pose severe or catastrophic risks. OpenAI disputed the characterization, saying the work had been redistributed among existing teams. The company has also previously dismantled its mission alignment team, while its head of ethics, Chloé Bakalar, recently left the company.  

It would be unfair to suggest that OpenAI has simply abandoned AI safety. The company is doing the opposite in some respects, adding safeguards precisely because Astra has become more capable. 

But institutional design matters as much as technical safeguards. If responsibility for safety is distributed across teams whose primary mandates include building, deploying, and commercialising increasingly powerful systems, who has the authority to say, “Stop”? Who independently challenges the assumptions made by the people developing the technology? And who has the organisational power to slow a release when commercial and competitive pressures are pushing in the opposite direction? 

There is an additional concern around Astra’s reported use of a technique called “recurrent depth”, which allows the model to perform additional internal processing without producing the same kind of visible reasoning trace associated with conventional reasoning models. Safety researchers worry that reduced visibility could make it harder to monitor what a model is doing and why.  

That is an uncomfortable combination: more capability, more autonomy, potentially less visibility, and questions about the strength of independent safety oversight. 

The stakes extend beyond OpenAI. The Financial Stability Board has already identified AI-driven cyber risk as the most immediate AI-related threat to global financial stability, warning that advanced models could accelerate vulnerability discovery faster than institutions can test and recover their systems.  

The industry is entering a phase where “move fast and fix the safeguards” becomes increasingly difficult to defend. Once an AI system can discover vulnerabilities, write exploits, operate autonomously, and potentially evade conventional monitoring, the margin for error shrinks dramatically. 

Astra may ultimately prove to be safely deployed. OpenAI may well succeed in building effective layers of containment and monitoring. 

But the larger question is harder to engineer away: as AI systems become powerful enough to require extraordinary guardrails, are the institutions responsible for providing those guardrails becoming stronger at the same pace? That is the question worth asking before Astra becomes widely available.