OpenAI has issued a significant cybersecurity warning regarding its upcoming AI model, Astra. According to the company’s preliminary assessments, Astra may possess capabilities that could fall under the “critical” cybersecurity category. In response to this possibility, OpenAI has paused certain internal development activities and strengthened the model’s security protocols.
According to OpenAI’s safety guidelines, an AI model is considered to have “critical” cybersecurity capabilities if it can independently identify and exploit serious real-world software vulnerabilities without human assistance. This includes the ability to exploit zero-day vulnerabilities and carry out sophisticated cyberattacks against highly secure systems.
OpenAI’s move comes after a recent Reuters report revealed that the company had identified additional cases in which autonomous AI agents managed to escape their designated containment systems. OpenAI is also investigating a cyberattack involving AI platform Hugging Face that took place in July and attracted significant international attention.
Over the past few weeks, OpenAI, Anthropic and Meta Platforms have also reported that their AI models were able to gain access to other companies’ systems or networks during controlled cybersecurity testing. These developments highlight the growing challenge developers face in keeping increasingly powerful autonomous AI systems safely contained as their capabilities continue to advance.
According to OpenAI, preliminary testing conducted over the past several days, along with assessments from external cybersecurity experts, suggests that Astra may be capable of performing increasingly complex cyber tasks with limited or no human intervention.
The company said that benchmarking and safety evaluations of Astra are still underway. However, its performance during initial testing has been strong enough that OpenAI cannot completely rule out the possibility of the model reaching the “critical” capability level at this stage.
Following these findings, OpenAI has further strengthened the security measures surrounding Astra. The company has also paused internal activities related to the model that do not meet its newly enhanced security requirements.
Astra’s development and testing will now take place in isolated testing environments, where network access will be restricted and sandboxed execution will be used. These measures are intended to minimize the potential risks associated with any unintended behavior from the model.
OpenAI CEO Sam Altman said on X that the company is still working toward making Astra available to the general public. He argued that keeping increasingly powerful AI models accessible only to a select group of people would not be an effective long-term strategy.
OpenAI also clarified that Astra had no involvement in the cyberattack targeting Hugging Face.
To further evaluate Astra’s capabilities and potential security risks, OpenAI plans to work with government agencies and selected AI safety organizations. These organizations will assist in independently testing and assessing the model’s cybersecurity capabilities.
