Can governments control an AI that breaks its cage?

Experts warn that advanced AI models with cyberattack capabilities should not be made open to the public. With defence windows down from weeks to hours in an AI cyberattack, Lianhe Zaobao journalist Tan Jet Min finds out what regulators need to do to mitigate the risks.

A person holds up a sign on a pedestrian bridge during a nationwide protest against AI data center expansion in Berkeley, California, on 18 July 2026.
A person holds up a sign on a pedestrian bridge during a nationwide protest against AI data center expansion in Berkeley, California, on 18 July 2026. (Josh Edelson/AFP)

(Edited and refined by Josephine Hong, with the assistance of AI translation.)

OpenAI, the US artificial intelligence (AI) giant, revealed that one of its advanced AI models autonomously escaped a sandbox testing environment and not only gained access to the internet, but also breached another company’s platform to steal information in order to complete its test task. Experts warn that the window for responding to AI-driven attacks has shrunk from weeks to between 12 and 72 hours, and in extreme cases may be only a few hours.

According to a statement published on OpenAI’s website, the company was testing the cyber capabilities of its GPT-5.6 Sol model and another pre-released model in a sandbox environment. Exploiting vulnerabilities in third-party software, the two models broke out of the sandbox, which had been highly isolated from the external internet, infiltrated the prominent open-source AI community Hugging Face, and stole confidential information from its database to forcibly complete their test tasks.

OpenAI subsequently disclosed that the models had also attacked four publicly available online services in the same incident.

Citing people familiar with the matter, Bloomberg reported that the attack, which would have taken human hackers weeks to carry out, was completed by OpenAI’s models in mere hours.

Claude app icon in this illustration taken 5 June 2026.
Claude app icon in this illustration taken 5 June 2026. (Dado Ruvic/Illustration/Reuters)

Another AI giant, Anthropic, encountered a similar incident during cybersecurity testing. After a testing environment was mistakenly connected to the public internet, three models, including Claude Mythos 5, mistook the open internet for a simulated environment, autonomously searched for vulnerabilities and breached three external organisations.

Given Mythos 5’s powerful capabilities and higher risks, Anthropic ultimately decided not to make it broadly available for now, allowing only vetted cyber-defence partners limited access for testing.

Vulnerabilities that could bring down the whole system 

Speaking to Lianhe Zaobao (LHZB), Gerald Mako, a University of Cambridge researcher and an expert in AI and cybersecurity, said frontier AI models are moving beyond simply assisting operators towards acting with greater autonomy in complex cybersecurity tasks.

He explained that rather than only generating code or answering technical questions, frontier models are increasingly capable of identifying vulnerabilities, using external tools, and planning multistep workflows. They can adapt to feedback, significantly compressing the time between vulnerability discovery and potential exploitation.

He added the enhanced offensive capabilities meant that the challenge to defenders is quickly shifting from detecting known malicious code towards anticipating systems that can reason and act within digital environments.

OpenAI logo is seen in this illustration created on 11 June 2026.
OpenAI logo is seen in this illustration created on 11 June 2026. (Dado Ruvic/Illustration/Reuters)

Hannah Lim, a business law associate professor at Nanyang Technological University (NTU)’s Nanyang Business School and an expert in AI and law, cautioned that such incidents should not be simply interpreted as evidence that AI are “super-intelligent”, as many cyber systems themselves contain vulnerabilities.

She explained that if a system is well configured and cybersecurity is well defended, none of the tests so far have indicated that the AI agents can successfully attack such systems. But for weak systems, such tools can quickly and successfully search for vulnerabilities and even bring the whole system down.

Defence response window shrinks from weeks to hours

Gabriela Ramos, a researcher at the US-based Institute for Artificial Intelligence Policy and Strategy, cautioned that governments and businesses could no longer measure response times in terms of weeks or months.

She said that in 2022, defenders typically had weeks between the disclosure of a vulnerability and its exploitation. Today, close to a third of new vulnerabilities are exploited within 24 hours, and realistic response windows are measured in 12 to 72 hours.

OpenAI CEO Sam Altman arrives for a meeting at the White House in Washington, DC, US, on 30 July 2026.
OpenAI CEO Sam Altman arrives for a meeting at the White House in Washington, DC, US, on 30 July 2026. (Nathan Howard/Reuters)

US authorities have required government agencies to patch the highest-risk vulnerabilities within three days. India’s Computer Emergency Response Team has also recommended that patching be done within 12 hours for critical systems.

Simon Chesterman, a David Marshall professor of law at the National University of Singapore (NUS) and AI governance and policy lead at the NUS AI Institute, said bluntly, “In some cases, we may no longer be talking about months or weeks, but hours.”

Calls for reporting mechanisms and stronger risk assessments

While Mako was more cautious about specifying a precise timeframe, he stressed that governments and businesses could no longer assume they had years to adjust gradually. The new capabilities demonstrated by advanced AI models could spread quickly through research publications, open-source implementations, fine-tuning or independent replication.

There is still no comprehensive understanding of the true scale of AI involvement in cyberattacks. Microsoft’s 2024 Digital Defense Report showed that its customers face more than 600 million attacks daily, while it blocks about 7,000 password attacks every second — up sharply from fewer than 600 per second in 2021. Research by IBM found that about 16% of data breaches now involve attackers using AI.

Get the ThinkChina Weekly Newsletter

Insights on China, right in your mailbox. Sign up now.

Ramos called for mandatory incident reporting with a harmonised taxonomy capturing how AI was used, to move beyond the limitations of assessing risks based only on scattered cases.

AI models have frequently demonstrated the ability to bypass security barriers and infiltrate networks on their own, prompting profound reflection within the international community on existing regulatory mechanisms. Several leading legal and tech experts have called on governments to establish a comprehensive regulatory system covering the entire model development process as soon as possible, and to promote cross-border regulatory coordination, avoiding the reliance on the self-discipline of tech giants for public safety.

The logo of Hugging Face Inc. AI startup on a smartphone and laptop, on 9 August 2026.
The logo of Hugging Face Inc. AI startup on a smartphone and laptop, on 9 August 2026. (Andrey Rudakov/Bloomberg)

Although Hugging Face reported the hacking by OpenAI’s GPT-5.6 Sol and other models to law enforcement, it is unclear whether the FBI has launched an investigation; publicly available information does not indicate that the US authorities have taken any punitive measures against OpenAI.

On 29 July, US President Donald Trump said, “We’re looking at AI, we’re looking at controls.” However, he also said that he does not want to “restrict” AI developers from creating new products.

Regulators struggle to keep pace with AI giants

This incident raises a series of regulatory questions: when should regulators intervene? Who should be responsible for assessing advanced AI models? And what responsibilities and obligations should AI developers bear?

Mako said that high-risk capabilities can emerge as early as the laboratory and internal evaluation stages, so regulation cannot focus solely on how AI systems are used after deployment. Instead, oversight must extend across the entire lifecycle — from internal evaluation and capability testing, through deployment decisions to continued monitoring after release.

Yet, current AI safety regulation suffers from a clear imbalance between regulatory authority and technical expertise.

Demonstrators participate in the "Stop the AI Race" protest march in San Francisco, California, on 11 July 2026.
Demonstrators participate in the "Stop the AI Race" protest march in San Francisco, California, on 11 July 2026. (Karl Mondon/AFP)

Ramos highlighted that advanced AI models are currently assessed almost entirely by the companies that develop them. Governments generally lack the technical expertise, auditing authority and access to models and computing resources needed for independent oversight. Until public institutions develop the capacity to conduct their own testing and audits, effective regulation will remain heavily dependent on trust in AI companies.

NTU’s Lim noted that most regulators are legally trained but do not necessarily have a deep understanding of how AI models and network systems operate. Even when legal and technical experts work together, differences in professional language and conceptual frameworks can result in risks going unnoticed. Strengthening interdisciplinary qualifications in the regulatory space is therefore an urgent priority.

Enforcing accountability and global standards 

On the implementation of regulation, Chesterman stressed that advanced AI model developers should be subject to rigorous capability and containment testing, stage access to potentially dangerous capabilities, maintain secure evaluation environments, and report serious incidents. These responsibilities should not depend solely on the goodwill of a handful of companies.

Lim further argued that governments should establish minimum cybersecurity standards for companies and clearly define the minimum level of responsibilities of AI vendors when things go wrong, rather than leaving it to litigation in contract law or negligence law to settle the matter.

Michael Trazzi, founder and CEO of The Inside View media company producing documentaries about AI Safety, speaks out against the AI race in front of the offices of Open AI during the "Stop the AI Race" protest march in San Francisco, California, on 11 July 2026.
Michael Trazzi, founder and CEO of The Inside View media company producing documentaries about AI Safety, speaks out against the AI race in front of the offices of Open AI during the "Stop the AI Race" protest march in San Francisco, California, on 11 July 2026. (Karl Mondon/AFP)

She also believes that AI models with powerful offensive cyber capabilities should not be made publicly available, but should instead be restricted to organisations for the purpose of finding and fixing security flaws.

However, even if a single country succeeds in building a comprehensive regulatory framework, it cannot address cross-border risks on its own.

Chesterman warned, “National regulation will never be sufficient on its own. The models, developers and attacks all cross borders. If regulation is to be effective, it will require substantial global coordination.”

Popular This Month

Culture

Society

Society

Politics

Technology