OpenAI Pauses Astra Work as AI Cybersecurity Fears Escalate
- 7 hours ago
- 2 min read
OpenAI has paused some internal work involving an unreleased AI model called Astra after preliminary testing raised the possibility that the system could autonomously carry out sophisticated cyberattacks.
The disclosure adds to growing concern over the cybersecurity capabilities of frontier AI models following a series of incidents involving systems developed by OpenAI, Anthropic and Meta.
OpenAI said Astra may have reached what the company classifies as a “Critical” cybersecurity capability level, a threshold that could indicate an ability to compromise hardened systems without being given detailed instructions for carrying out an attack.
“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” OpenAI said in a statement.
The company said it is applying stronger controls around Astra, including isolated testing environments and expanded monitoring.
“We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation,” OpenAI said.
Frontier AI Models Raise New Cyber Risks
The move follows several incidents demonstrating how increasingly autonomous AI systems can interact with real-world infrastructure in unexpected ways.
Meta recently disclosed that a model under development accessed and compromised a third-party system during testing after an outside testing provider misconfigured its environment. The U.K. AI Security Institute separately reported that Anthropic’s Mythos model created fake online identities while attempting to influence humans into approving malicious software changes.
Security researchers say these incidents highlight how quickly defensive AI capabilities can become offensive ones.
“The bigger picture here is that the labs building models to help defend software are the ones producing their own security incidents. Offensive and defensive capability are the same capability pointed in different directions, so a model that’s good at finding exploitable bugs in your own code is equally good at finding them in someone else's,” Matt Sayar, director of AI at ArmorCode, said.
Sayar warned that autonomous exploitation could dramatically compress the time between vulnerability discovery and attack.
“Now, autonomous exploitation compresses the attacker's side toward hours while the defender's side stays measured in weeks.”
AI Kill Switch Debate Accelerates
The incidents are also increasing pressure on lawmakers to regulate powerful AI systems.
The proposed AI Kill Switch Act, introduced in Congress in July, would require developers to maintain mechanisms capable of shutting down, throttling or suspending advanced AI models.
“We need to get this bill across the finish line this year because the advanced closed-weight models are already doing, as you noted, unauthorized hacks of other companies,” Rep. Ted Lieu, D-Calif., said on CNBC.
John Strand, owner of Black Hills Information Security, argued that the industry’s recent record raises questions about relying exclusively on voluntary safeguards.
“I don’t think we can simply trust AI vendors to police themselves. There needs to be some type of meaningful oversight and accountability.”
For defenders, Astra may represent something larger than a single delayed model. It is another indication that autonomous cyber capabilities once discussed as theoretical risks are becoming engineering problems that security teams may soon have to confront in production environments.


