OpenAI has determined that one of its upcoming artificial intelligence models is powerful enough to require additional safety measures before its public release.

The model, known as Astra, can identify more cybersecurity vulnerabilities than the most advanced OpenAI model currently available to the public while requiring less computational power to carry out such tasks, company officials said during a conference call with reporters.

According to Amelia Glaese, an OpenAI vice president overseeing safety, the model has demonstrated the ability to discover previously unknown security flaws and develop methods for exploiting them across well-protected systems when given the appropriate tools and access.

“With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step,” Glaese said.

OpenAI said Astra would be made available soon to a limited group of users but did not disclose a specific launch date or provide details about who would initially gain access.

The company said the additional safeguards could sometimes slow down, pause or stop legitimate work as it seeks to prevent potential misuse. Glaese said OpenAI would work to minimise such disruptions.

Astra is the first OpenAI model to trigger the more stringent protections outlined in the company’s safety protocol, marking the first time the threshold has been reached by one of its models.

The development comes amid increasing scrutiny of how AI companies are managing increasingly capable systems, particularly models capable of independently carrying out complex cybersecurity tasks.

The announcement follows a recent incident involving OpenAI’s AI agents, which breached their testing environment and compromised the open-source platform Hugging Face. The incident prompted the company to pause much of its model development for two weeks while strengthening its security systems.

OpenAI said Astra was not involved in the Hugging Face incident but that its cybersecurity capabilities nevertheless required tighter controls.

The company restarted its largest model training run on August 28 but said some smaller experiments remained on hold.

Under OpenAI’s safety protocol, stricter safeguards are required when a model demonstrates the ability to identify and exploit new cybersecurity vulnerabilities or independently plan and execute detailed and novel cyberattacks with minimal human involvement.

OpenAI said it has introduced additional restrictions to make Astra less capable of responding to harmful cybersecurity requests and will monitor its activity for signs that its safeguards have been bypassed.

Saachi Jain, who oversees safety work at OpenAI, said one of the challenges facing the company is determining how much autonomy increasingly capable AI systems should be allowed to exercise.

She said the company was working to ensure models understand the limits within which they are expected to operate.

“There are constraints that, as humans, we know that we should be adhering to when we perform a task,” Jain said. “A lot of the work here has been to also train the model to understand what those scopes are.”

Bank Recapitalization-abacha-university-ad