UAE Interact
Search
  • Home
  • Business
  • Finance
  • Technology
  • Press Release
  • International
  • National
Reading: OpenAI and Anthropic AI Agents implicated in UK AI Security Institute breach tests
Share
Font ResizerAa
UAE InteractUAE Interact
Search
  • Home
  • Business
  • Finance
  • Technology
  • Press Release
  • International
  • National
Follow US
© 2026 UAE Interact. All Rights Reserved.
Technology

OpenAI and Anthropic AI Agents implicated in UK AI Security Institute breach tests

By spsingh
Last updated: August 5, 2026
4 Min Read
Share
OpenAI and Anthropic AI Agents implicated in UK AI Security Institute breach tests

UK AI Security Institute tests found OpenAI and Anthropic AI agents performed unauthorized actions, including creating fake identities and attempting malicious code, highlighting gaps in agent evaluation safeguards

SAN FRANCISCO: An AI agent was caught creating fake online identities to gain unauthorised access to secure systems during tests of models from OpenAI and Anthropic which revealed a series of new breaches, Britain’s AI Security Institute disclosed on Tuesday.

The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol engaged in unauthorized actions during security evaluations the government organization conducted to assess the models’ capabilities.

“Some of the agents being tested ⁠had engaged in sustained, potentially harmful activity directed at real people and organisations,” AISI said in a blog post.

The report underscores the lax state of safeguards around the process of testing agents, which AI companies are simultaneously marketing as the future of business.

AISI, which receives access to advanced AI models under voluntary agreements from major ‌labs, put the agents through a fictional cybersecurity scenario to test their capabilities.

It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic’s agent was behind 17 of the actions, and OpenAI’s ‌agent the remaining two.

The fact that Mythos engaged ⁠in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think

The most egregious action involved an agent writing malicious code and creating ‌fake online identities in an attempt to get a human to approve the code, AISI said, adding ‌that no real-world harm was found as a ‌result of any of the breaches.

While AISI did not say which agent was behind the fake identities, Antropic confirmed its agent was responsible.

“We’re grateful to the UK AISI for their ‌leadership on this incident, which underscores the need for a broader conversation about how to ⁠safely evaluate increasingly capable AI agents,” Anthropic said in a statement.

It also said it was working with AISI to obtain more details on the incident and conduct its own investigation.

Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and ⁠dangers, said: “The fact that Mythos engaged ⁠in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think.”

OpenAI shared details in a company blog post, noting that both ⁠of its agent’s unapproved actions involved accessing the internet in ways that were forbidden by the prompt.

“We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks,” OpenAI said.

OpenAI also disclosed in its blog post a separate incident whereby a misconfiguration by Irregular, a third-party testing provider, ‌allowed its agents to mistakenly connect to the internet. It mirrored a similar disclosure about misconfiguration that Anthropic made last week.

Reuters reported last week that OpenAI had widened its hacking probe after finding evidence of other agent breakouts.

Unlike the July security breach of AI firm Hugging Face by an OpenAI agent, the agents in the AISI evaluation did not escape an isolated testing environment to reach the internet. Rather, the agency had permitted internet access in line with its standard testing procedures, AISI said.

spsingh
Website |  + postsBio ⮌
  • spsingh
    Emirates Crew Concludes 2026 Nationwide Carrier Occupation Truthful Participation in Dubai
  • spsingh
    Sturdy earthquake rattles New Zealand’s South Island, tsunami alert lifted
  • spsingh
    UAE showcases AI-driven labour marketplace fashion at BRICS labour ministers’ assembly
  • spsingh
    UAE Monetary Balance Council approves AI plan, First Monetary Balance Convention 
author avatar
spsingh
See Full Bio
TAGGED:AgentsAnthropicbreachimplicatedInstituteOpenAIsecuritytests

Sign Up For Daily Newsletter

Be keep up! Get the latest breaking news delivered straight to your inbox.
[mc4wp_form]
By signing up, you agree to our Terms of Use and acknowledge the data practices in our Privacy Policy. You may unsubscribe at any time.
Share This Article
Facebook Email Copy Link Print

HOT NEWS

Chicano Hollywood Film Festival and Pure Flix Familia Honor Oscar-Nominated Makeup Artist Ken Diaz

Press Release
August 6, 2026
UAE to Open Registration for 2027 Hajj Season on September 14

UAE to Open Registration for 2027 Hajj Season on September 14

UAE citizens who want to perform Hajj in 2027 will be able to start registering…

August 10, 2026
Dubai: When a town turns into a verb

Dubai: When a town turns into a verb

Dubai is defining itself, by itself phrases, in its personal language, and handing the sector…

June 28, 2026
IBM unveils 0.7nm nanostack chip tech for AI computing, competing with TSMC and Intel

IBM unveils 0.7nm nanostack chip tech for AI computing, competing with TSMC and Intel

The brand new chip era, which bolsters IBM's place to compete with contract chipmakers TSMC…

June 28, 2026

YOU MAY ALSO LIKE

UAE Cyber Safety Council thwarts subtle cyberattacks focused on monetary sector

The assaults integrated makes an attempt to focus on virtual methods and technical infrastructure, habits subtle phishing campaigns, exploit safety…

National
July 4, 2026

AI’s rising power and water use sparks environmental issues

Washington: The fast upward push of man-made intelligence is considerably expanding international power and water intake, elevating issues amongst professionals…

Technology
June 28, 2026

Global’s greatest area dealer GoDaddy warns India faux website crackdown may just hurt web protection and bonafide companies

GoDaddy, different area dealers resort considerations, as India ruling is reshaping how web governance worksThe sector's greatest web area dealer,…

Technology
July 3, 2026

UAE leads world AI adoption, surpasses 70% utilization fee

Microsoft record highlights fast expansion pushed via sturdy virtual infrastructure and fundingDubai: The UAE has ranked first globally in synthetic…

Technology
July 4, 2026
Logo

Welcome to UAE Interact, your trusted destination for timely, accurate, and insightful news from the United Arab Emirates and around the world

Quick Links

  • About Us
  • Contact Us
  • Cookies Policy
  • Privacy Policy
  • Disclaimer

Top Categories

  • Home
  • Business
  • Finance
  • Technology
  • Press Release
  • International
  • National

Follow US: 

UAE Interact

All Rights Reserved.

Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?