The headlines have been filled recently with accounts of AI agents acting autonomously and attacking third-party businesses. The issue became public when OpenAI published a statement describing how an agent they had been testing had attacked Hugging Face (an online model repository). They described the incident as an “unprecedented cyber incident” involving “state of the art cyber capabilities” and have engaged two external research organisations to assist in the investigation.
Current findings indicate that OpenAI was evaluating two models – one pre-release model and one existing model (GPT 5.6 Sol) – for the purpose of testing the models’ “cyber capabilities”. To achieve this, some of the default safety settings for the models were disabled. These settings are normally in place to prevent a model being used for cyber-attacks. The models were then run in a sandbox test environment with no external internet connection.
Unexpectedly, the models discovered a zero-day exploit (an unknown security flaw) in the test environment, which enabled them to gain external internet access. To meet their objective, the models selected Hugging Face as a data source that would hold the answers they needed to pass their test. They proceeded to use publicly available user credentials for Hugging Face to gain access to the information they needed. The agents performed in the region of 17,000 actions over two days, leading Hugging Face to believe it was subject to a sophisticated nation-state or cyber-criminal attack. OpenAI later acknowledged that the attack was generated by one of its test models. OpenAI was not initially aware of the event and did not know the models had escaped the sandbox. Hugging Face is not pursuing legal action but is working with OpenAI to investigate the incident and restore its systems.
The importance of appropriate safeguards
This chain of events was remarkable as it was considered the first time an AI Agent had autonomously hacked a third party and was also believed to be the first time an advanced AI model had escaped from a sandbox. The agent was not adversarial by design and did not have specific instructions to attack any third parties; rather, it was using all available means to achieve its test objective. The incident has led cyber security professionals and AI research teams to reconsider the controls that should be applied when developing and testing AI models.
It also highlights a new set of risks relating to unintended consequences in AI development and use. The revelation has since led other AI providers to share that they have experienced similar incidents, such as Anthropic, which had an AI model go rogue in April 2026. This was only discovered following checks they ran in response to the Hugging Face situation. Anthropic found that its models had hacked three different systems in a similar way, by identifying an external internet connection outside their sandbox. This is believed to have be made possible by a misconfiguration by a third-party provider supporting the test environment. Most recently, Meta has also announced that it had seen an advanced model escape a sandbox through a third-party misconfiguration and then go on to behave in unintended ways.
While one perspective is that this demonstrates the capability of advanced models, it also emphasises the importance of ensuring released models are subject to appropriate safeguards and that test environments are properly secured. Other questions have been raised in relation to transparency, as these failings were not originally announced publicly.
Assessing potential risks
However, further work by the AI Security Institute has involved testing AI agents to assess the potential risk posed by uncontrolled agents that have been jailbroken by a hacker (had their safeguards removed). The Institute found that the model tested, Anthropic’s Mythos, identified GitHub users, impersonated those users and attempted to contact GitHub directly to gain access to live accounts. GitHub staff recognised the attempts as social engineering and blocked the agent.
For organisations that are not frontier AI providers, it may seem that these events do not have a direct impact. However, they highlight the potential harm arising from use of agents without suitable controls. Given the proliferation of agent offerings now available as features in many SaaS (software-as-a-service) and AI productivity tools, it is critical that organisations carefully assess the risks and control their use in operations.
Preparing for AI-generated cyber attacks
The UK’s NCSC (National Cyber Security Centre) issued a statement in June 2026 urging businesses to prepare for the risk from AI-generated cyber attacks. The statement was issued in response to a briefing from the Five Eyes partnership (the UK, US, Canada, Australia and New Zealand intelligence alliance), which issued a warning regarding the use of advanced AI models to accelerate the speed, scale and sophistication of attacks. The NCSC has also published guidance that should be applied when developing AI.
From a business perspective, due diligence is essential for any AI tool, and special consideration is needed when configuring agents. Much of the existing security approach relies on controls based on of user identity, which are effective in the context of a known employee, but may not be applicable to an ‘unknown’ agent. Many free-to-use AI agents are, by default, able to access all company data and systems in order to perform administrative tasks on behalf of their creator. Often, only tools offered under a paid licence have settings to manage what is accessible. For a business striving for efficiency gains, permitting staff to use free AI agents in the workplace, or unconfigured licensed agents, represents a significant risk. In the drive to improve performance, shortcuts in compliance can have grave consequences. Having an AI policy is simply not enough to govern responsible AI use in a business, and using autonomous agents amplifies these risks.
While few organisations will be testing AI models with reduced guardrails in a sandbox, the press releases from OpenAI, Anthropic and Meta serve as a stark reminder that uncontrolled use has consequences. One approach to ensuring your business is following best practice is to select a governance framework and conduct regular testing of all AI systems in use.
How we can help
GRC Solutions has consultants who can advise on selecting an appropriate framework and can also offer ISO 42001 readiness assessments for organisations that want to either align with, or prepare to certify to, this international AI management standard.
For organisations that are already operating to ISO standards such as the ISO 27001 information security management standard, or the ISO 27701 privacy information management standard, then the ISO 42001 AI management system framework is designed to harmonise with these standards, creating efficiencies in implementation.
Where AI systems are already deployed, we can provide specialist LLM red team-assessments to check the security of the AI you have in use.
Where you are using AI for services involving personal data, we have a team of DPO consultants and data protection lawyers who can provide tailored advice on the implications of AI for data protection and assess your operations against regulations such as the EU AI Act.