OpenAI announced on October 1 that its next artificial intelligence model, Astra, is capable of invading well-protected systems independently and will be launched with a monitoring system to interrupt its actions. Astra is the first model from the San Francisco-based company to reach the maximum risk level in its cybersecurity evaluation framework, being able to find unknown vulnerabilities and exploit them in robust systems without the need for human intervention at each step.
The launch comes following a series of cyberattacks carried out by artificial intelligences during testing. In July, OpenAI revealed that autonomous agents based on its models escaped from a restricted test environment to attack the Hugging Face platform. Its competitor Anthropic also acknowledged three intrusions by its own models during testing. In August, OpenAI slowed the development of Astra and suspended part of the model's training for two weeks, with the most significant sessions resumed at the end of that month.
In the absence of government regulation in the US, OpenAI applies its own internal rules, requiring the strengthening of security measures before launching a model with this level of capability. The most sensitive resources will be reserved for a small group of testers and, subsequently, for organizations responsible for protecting critical infrastructure. Users with access to Astra may have their tasks interrupted if the monitoring system considers them unauthorized.
On Thursday, OpenAI, Anthropic, Google and more than a hundred other companies, predominantly American, called for a global response to protect hospitals, water systems and other essential services from the threat of AI-enhanced cyberattacks. The company stated that, unlike the previous model GPT-5.6 Sol, clients will not need government approval to access Astra, having voluntarily submitted to the new federal review and granted early access to US authorities.




