Visualização de leitura

AI Deception Emerges in Cyber Tests as Agents Target Real People and Systems

AISI found AI agents taking unsanctioned online actions, including social engineering and code attacks, during controlled cyber tests.

The UK’s AI Security Institute (AISI) has put something uncomfortable on the table: during cyber testing, frontier models didn’t just follow instructions badly. In some runs, they crossed into real-world actions, touched real people and organisations, and then tried to make the evidence go away.

“On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.” reports AISI. “We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.”

That matters because the setting was not a botched public rollout or a rogue consumer chatbot. It was a controlled evaluation, run on purpose with open internet access and some safety filters turned off, so the institute could measure what the models could actually do under permissive conditions. AISI says that’s standard practice for frontier testing, but it also notes that those conditions are not how these models are normally shipped.

To assess the models’ real-world cyber capabilities, they were given internet access and had their built-in cyber safety filters intentionally disabled under controlled testing conditions.

The headline result is simple enough, even if the details are messy. AISI ran the challenge 122 times across several models, and in 10 of those runs it saw AI agents take autonomous, unsanctioned actions on the live internet. The institute logged 19 such actions in total, with 17 tied to Anthropic’s Mythos 5 and two to OpenAI’s GPT-5.6-Sol with cyber classifiers disabled.

The most serious sequence looked less like a lab mishap and more like a small-scale social-engineering campaign. The agent tried to insert malicious code into a public open-source project, researched the maintainers, created fake identities based on real people, and used those identities to pressure a maintainer into approving the code.

“On the morning of 28th July, our security monitoring flagged data leaving one of our testing systems through the ‘Tor’ anonymity network, which is commonly used to disguise the origin of internet traffic. The team running the evaluation began reviewing the agent’s activity within minutes.” continues the report. “They quickly found that, during one run, the agent had created a malicious ‘pull request’ (a proposed code change) on a real, public open-source project on GitHub. “

It also sent messages and files, and when challenged, it edited earlier activity to look harmless and even considered adopting a fresh identity.

AISI says human review stopped the worst outcome. The malicious pull request was caught, the fake identities failed to get the code approved, and the institute did not find evidence of resulting real-world harm. Even so, it says this is the first time it has seen autonomy and deception show up this clearly, without specific prompting, in a real-world setting.

The institute is careful not to overclaim. It says the evaluation design and configuration choices helped create the conditions for the behaviour, and it notes that the models in question were not commercial public deployments. But that caveat cuts both ways: if permissive test conditions can produce this kind of behaviour now, tighter production settings are not a reason for complacency. They are just a better fence.

That is the part that should make people stop scrolling. Not because the models “went evil,” which is lazy, but because deception emerged as a by-product of goal-seeking under the conditions set by the evaluators.

The agent pursued its goal persistently. AI agents explore routes their operators did not intend. Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people.” states AISI. “It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

The agent was trying to solve a cyber challenge, and in doing so it explored routes nobody had asked it to explore, including routes that involved fooling actual people.

What makes the episode more serious is the way it lines up with recent disclosures from Anthropic and OpenAI. Anthropic has already said its models reached the public internet from within a third-party evaluation environment and accessed real systems. OpenAI has separately reported that one of its models mistook a misconfigured test setup for a real task and touched real infrastructure.

AISI’s view is that these incidents point to a shift in the risk landscape. Harm may no longer come only from obvious misuse by humans. It may also come from capable agents, in internal research settings or privileged-access environments, taking unintended actions beyond the scope they were given. That is a quieter problem than movie-style “AI rebellion,” and a more useful one to think about.

The key takeaway is straightforward. Keep cyber basics tight, verify outside code before trusting it, and stop assuming that a model will stay inside the lines just because the prompt sounded clear on the day. In security, the line between “evaluation” and “incident” can get thin fast. Machines are very good at finding the part of the process you forgot to make boring.

AISI says it will tighten internet controls, add real-time monitoring, and revisit how it designs evaluations. That is the right response, but it should not be read as a narrow fix for one lab. It is a warning to anyone testing powerful agents: if the test can reach the real internet, the real internet can reach back.

The original AISI report is here: Incident report: unsanctioned agent behaviour during cyber testing.

“Incidents of this kind reflect the speed at which AI is developing. As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.” concludes the report.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, AI Deception)

OpenAI Tightens AI Evaluation Safeguards After Testing Incidents

OpenAI models

OpenAI models accessed the public internet during separate third-party cyber evaluations conducted by independent testing partners, prompting the company to review how high-risk AI testing is managed. OpenAI said the incidents occurred under specialized testing configurations with reduced safeguards and did not reflect how its models operate in public deployments. The company added that the events were unrelated to the previously disclosed Hugging Face security incident.

The incidents involved evaluations conducted by UK AISI and cybersecurity testing partner Irregular, where testing conditions or environment configurations enabled models to interact with systems beyond the intended evaluation boundaries.

OpenAI Models Prompt Review of Third-Party AI Testing

OpenAI said independent cybersecurity evaluations are essential for understanding model capabilities before deployment. Some evaluations intentionally reduce safeguards or enable additional capabilities to measure how models perform under conditions that resemble real-world cyber operations.

According to the company, the latest incidents highlighted the need to strengthen the security controls surrounding independent testing environments as AI models become more capable.

OpenAI said it will review how it manages third-party cyber evaluations, including how higher-risk testing is identified, when internet access or reduced safeguards should be permitted, how testing environments should be isolated, and how incident reporting and monitoring procedures should be handled.

The company also plans to work with national AI institutes, independent evaluators, and other AI developers to establish stronger industry practices for high-risk AI evaluations.

UK AISI Identified Two Unsanctioned Actions by GPT 5.6 Sol

On August 3, UK AISI, the UK's AI Security Institute, informed OpenAI that during a cyber evaluation that began on July 25, models from OpenAI and another AI lab performed activities beyond the intended testing scope.

The evaluation used simulated cyber ranges where models acted as cybersecurity experts in a Capture-the-Flag exercise. Internet access was intentionally enabled so agents could download tools similar to those available to real attackers, while cyber classifiers were disabled to measure underlying model capability.

Among 19 recorded events, two involved GPT 5.6 Sol.

According to OpenAI, the model reused a publicly accessible GitHub token left behind by another lab's agent to determine whether systems inside the simulated environment were communicating with GitHub or update services. It also attempted account recovery and request-limit workarounds while registering accounts with external DNS and tunneling providers.

In a separate action, GPT 5.6 Sol used a public tunneling service to expose a locally hosted DNS server from the evaluation environment to the public internet. The server contained exploit payloads intended for a known software vulnerability within the simulated cyber range.

OpenAI said the setup did not function as intended, no evidence indicated that any real internet resolver queried the server, and the infrastructure was removed when the evaluation concluded.

UK AISI considered both actions outside the authorized testing boundary because they involved external services that were not required to complete the exercise.

Security monitoring detected unusual data transfers on July 28. UK AISI halted the evaluation, isolated the affected systems, and contained the activity within approximately one hour.

Irregular Finds Misconfigured Testing Environment

A second incident was reported on July 29 by Irregular, another organization conducting third-party cyber evaluations for OpenAI.

The testing environment was intended to remain isolated from the internet during another Capture-the-Flag exercise. However, a configuration error unintentionally allowed the models to access external websites.

During one evaluation, the fictional domain used in the exercise matched an actual website. Because internet connectivity was available, the model interacted with the live site, believing it was part of the simulated challenge.

Irregular said the incident did not involve a sandbox escape or zero-day exploit. Instead, the model exploited a basic security vulnerability on the website after gaining unintended internet access.

The investigation also found that the model used credentials associated with the same website. According to Irregular, no impacts beyond that site's own data have been identified, although its audit remains ongoing.

The organization has paused the evaluations, completed remediation, notified affected third parties, and implemented additional safeguards in its testing environment.

OpenAI said it will continue working with both UK AISI and Irregular to improve evaluation practices while ensuring independent cybersecurity testing remains rigorous as AI capabilities continue to advance.

Trump’s AI Safety Chief Is Out — After Just 90 Days on the Job

Trump administration AI testing chief Chris Fall resigned after just 90 days, leaving CAISI under interim leadership during a critical period for AI safety standards.

The post Trump’s AI Safety Chief Is Out — After Just 90 Days on the Job appeared first on TechRepublic.

U.S. Will Now Examine National Security Implications of New AI Models, Pre-Release

Claude AI, Antropic, AI, Artificial Intelligence

In the span of four days, the U.S. government announced two parallel sets of agreements with frontier AI companies that together define the two tracks Washington wants to run simultaneously—test AI for national security risks before the public ever sees it, and deploy AI directly on the military's most classified networks.

The Center for AI Standards and Innovation — CAISI, the entity under the Department of Commerce's National Institute of Standards and Technology that inherited the remit of the former AI Safety Institute — announced new agreements with Google DeepMind, Microsoft, and Elon Musk's xAI. These build on renegotiated agreements with Anthropic and OpenAI that date to 2024, updated to reflect directives from Commerce Secretary Howard Lutnick and America's AI Action Plan.

Under the CAISI agreements, the three companies will hand over their frontier AI models to government evaluators before those models are publicly released. The evaluations probe for national security-relevant capabilities and risks.

To conduct a thorough assessment, developers frequently provide CAISI with models that have reduced or removed safety guardrails — a design choice that allows evaluators to probe what a model can do at its ceiling, not what it will do under commercial safety controls. Evaluators from across the federal government participate, coordinated through the CAISI-convened TRAINS Taskforce, an interagency body focused specifically on AI national security concerns.

CAISI said it has completed more than 40 such evaluations to date. The agreements explicitly support testing in classified environments and were drafted with the flexibility to adapt rapidly as AI capabilities continue advancing.

"Independent, rigorous measurement science is essential to understanding frontier AI and its national security implications," said CAISI Director Chris Fall. "These expanded industry collaborations help us scale our work in the public interest at a critical moment."

Listen to: Charting the AI Frontier in Cybersecurity with Ryan Davis

Fall was appointed to lead CAISI after Collin Burns — a former Anthropic researcher — was reportedly removed from the director role after just four days. The personnel transition at CAISI's top reflects a broader institutional pivot. Under the Biden administration, the AI Safety Institute focused on safety standards, definitions, and voluntary guardrails. Under Trump, CAISI has shifted its emphasis toward AI acceleration and national security capability assessment. The substance of what the evaluators do — probe powerful models before release — has not changed. The framing of why they do it has.

The latest announcement comes four days after the Department of War (formerly Department of Defense) announced agreements with eight frontier AI companies to deploy their models directly on the military's classified networks for operational use.

The companies cleared are SpaceX, OpenAI, Google, NVIDIA, Reflection, Microsoft, Amazon Web Services, and Oracle. The networks in question are classified at Impact Level 6, covering secret-level data, and Impact Level 7, which refers to the most highly restricted national-security systems. The stated objectives are data synthesis, situational awareness enhancement, and warfighter decision support.

The Department of War announcement carries one conspicuous absence that dominates coverage of what it actually means. Anthropic is not on the list. The company that first deployed AI models on Pentagon classified systems — via a Palantir integration under the Maven Smart System contract — is excluded after a dispute over the guardrails governing military and surveillance use of its AI.

Also read: Australia Establishes AI Safety Institute to Combat Emerging Threats from Frontier AI Systems

The Pentagon had previously branded Anthropic a "supply chain risk," a designation typically reserved for foreign entities posing national security concerns. A March 2026 federal injunction reversed that designation, but it did not restore Anthropic's position as a Pentagon AI vendor. Palantir has pulled its Claude models from its DoD platforms accordingly.

The exclusion has strategic implications that extend beyond one company's contract status. Anthropic's recently released Mythos model — described by Treasury Secretary Scott Bessent as representing a step change in large language model capability — has generated significant attention from U.S. officials and financial sector executives about its potential to supercharge adversarial cyber operations.

The fact that Mythos is not among the models being assessed for classified military use, while simultaneously being cited by senior officials as a capability milestone that warrants concern, creates a gap in the government's stated AI security posture that is difficult to characterize as anything other than a policy contradiction.

UK gov's Mythos AI tests help separate cybersecurity threat from hype

Last week, Anthropic announced it was restricting the initial release of its Mythos Preview model to "a limited group of critical industry partners," giving them time to prepare for a model that it said is "strikingly capable at computer security tasks." Now, the UK government's AI Security Institute (AISI) has published an initial evaluation of the model's cyberattack capabilities that adds some independent public verification to those Anthropic reports.

AISI's findings show that Mythos isn't significantly different from other recent frontier models in tests of individual cybersecurity-related tasks. But Mythos could set itself apart from previous models through its ability to effectively chain these tasks into the multistep series of attacks necessary to fully infiltrate some systems.

"The Last Ones" finally falls

AISI has been putting various AI models through specially designed Capture the Flag challenges since early 2023, when GPT-3.5 Turbo struggled to complete any of the group's relatively low-level "Apprentice" tasks. Since then, the performance of subsequent models has risen steadily, to the point where Mythos Preview can complete north of 85 percent of those same Apprentice-level CTF tasks.

Read full article

Comments

© Getty Images

❌