We’ve now seen three disclosures from OpenAI, Anthropic, and Meta about models hacking third party systems. Four takeaways so far.
Self-awareness exhibited by models calls into question model test results, especially on alignment and safety
Focusing on the OpenAI/Hugging Face incident disclosed July 21, the model correctly reasoned that:
It was being tested
That the test it was taking was ExploitGym (a test of offensive cybersecurity capabilities)
That it could pass the test by cheating and find the answers on Hugging Face (a large repository of models tests and answer keys)
That it could collaborate with other models, dividing and conquering work, and sharing information
The result of this reasoning chain was the hack to get out of its test environment and into Hugging Face to get the test answers.
The problem (beyond the escape, which I’ll get to in a second) is that self-awareness of AI can lead to the release of misaligned and unexpectedly powerful models. If models can recognize when they are being tested, then they can cheat when they don’t have the answer, lie when asked about alignment, or play dumb to hide much stronger capabilities - all so that they can “pass the test” and get released. This is a longstanding problem that OpenAI called “scheming” in a research post last year.
Labs do not have a good way to detect when the models are scheming (or these issues would have come to light in a different way - and much sooner). So how can we be sure current and future models have legitimately passed?
The world’s leading AI labs are unable to contain their models when the guardrails are off. (But first, are these companies just pulling a PR stunt?)
While there is an enormous amount of speculation that these hacks were faked or part of a PR stunt, I believe these models are capable of these hacks because:
They all involve a corroborating third party. For OpenAI, it was Hugging Face; for both Anthropic and Meta, it was Irregular.
Companies are already using less capable models for cyber defense. It’s not a stretch that a more capable model without guardrails could go on offense.
I spoke to someone involved in Anthropic’s Project Glasswing that got early access to Mythos to identify and patch vulnerabilities. While they didn’t share anything proprietary, this very serious person got a bit glassy-eyed talking about their experience, like they’d seen something otherworldly.
OpenAI gave a detailed talk at the Black Hat conference and shared significant details about the incident include the collaboration between models that led to the escape. You can judge for yourself but it comes across as a well documented series of lapses and a new capability they are very worried about.
If the world’s best labs are unable to contain their models with the guardrails off, should they be testing the models with the guardrails off? Should they even be training models with these baseline capabilities in the first place? And why aren’t models tested individually in isolated, air-gapped instances?
The debate is similar to gain-of-function research that intentionally increases virulence and transmissibility of pathogens. While there are some benefits to the research - like understanding disease adaption - it is also very risky. The US banned this research in 2014, rescinded the ban in 2017, and as recently eliminated funding for it July of this year. The similarities don’t end there.
Biological agents and nuclear enrichment technology are useful analogues for frontier models
As a caution against overzealous regulation, some analysts have pointed to encryption as a good analogue for advanced models with cyber capabilities. Encryption was once subject to export controls because the government worried about the “bad guys” getting access to it. Ultimately it was more important that as many of the “good guys” adopted it to protect consumers and businesses. Indeed the OpenAI-Hugging Face incident showed the value of defenders having access to advanced models with cyber defense capabilities.
This analogy breaks down because encryption is only a defensive capability. These models clearly represent an offensive capability and can cause damage.
A better analogue would be biological agents which are important to research but can be extremely dangerous. Anthropic actually modeled their AI Safety Levels (ASL) after the US government’s Biosafety Levels (BSL) for handling of dangerous biological materials.
Released in September 2023, it’s telling how fast the space has accelerated (emphasis mine):
ASL-1 refers to systems which pose no meaningful catastrophic risk, for example a 2018 LLM or an AI system that only plays chess.
ASL-2 refers to systems that show early signs of dangerous capabilities – for example ability to give instructions on how to build bioweapons – but where the information is not yet useful due to insufficient reliability or not providing information that e.g. a search engine couldn’t. Current LLMs, including Claude, appear to be ASL-2.
ASL-3 refers to systems that substantially increase the risk of catastrophic misuse compared to non-AI baselines (e.g. search engines or textbooks) OR that show low-level autonomous capabilities.
ASL-4 and higher (ASL-5+) is not yet defined as it is too far from present systems, but will likely involve qualitative escalations in catastrophic misuse potential and autonomy.
Based on these definitions, model capability is already at the maximum defined risk. Containment procedures have not kept pace.
Oversight is needed
There are lots of examples for how to mitigate catastrophic risk and continuously improve safety.
The US Government’s Federal Select Agent Program (FSAP) enforces the BSL standards that Anthropic modeled its safety system after, registering labs and levying penalties when standards are not met and when agents are mishandled.
The National Transportation and Safety Board (NTSB) investigates incidents such as airplane crashes without assigning blame or liability in order to constantly improve a culture of safety and reduce harm. They publish detailed findings and recommendations which are adopted by industry and regulators. The world would benefit from shared best practices in model evaluation, containment, and detection of scheming.
The detailed disclosures from OpenAI and Anthropic are an important start, but can’t be the end. Concern about AI impact on society is a bipartisan issue. These incidents must serve as a catalyst for oversight that improves safety and ensures progress - especially on cybersecurity defense from rogue models. Or we are just going to see more and bigger incidents.


