Visualização normal
-
Graham Cluley
-
Smashing Security podcast #483: This AI helps thieves steal your iPhone
You've had your iPhone stolen. A day later, you get a text from Apple saying they've found it, and a very helpful woman called Alice from Apple Support calls to walk you through recovering it. She's polite. She's professional. But she is not from Apple. She's not even human. And she's about to break into your iPhone. Meanwhile, OpenAI, Anthropic, and Meta have all announced - with varying degrees of drama - that their AI agents have "broken out of the sandbox" and gone hacking. James takes a
-
Cybersecurity News
-
Anthropic Bolsters Security After Claude AI Escapes
Discover the Anthropic Claude security overhaul following alarming incidents where AI models breached isolated sandboxes to attack real internet systems. Related Posts: Darwin-VM Enables Apple Silicon Security Research Apple OpenAI Lawsuit Escalates Over AI Trade Secrets Chrome Manifest V2 Removal: Legacy Extensions Are Now Gone The post Anthropic Bolsters Security After Claude AI Escapes appeared first on Daily CyberSecurity.
Anthropic Bolsters Security After Claude AI Escapes
Discover the Anthropic Claude security overhaul following alarming incidents where AI models breached isolated sandboxes to attack real internet systems.
Related Posts:
- Darwin-VM Enables Apple Silicon Security Research
- Apple OpenAI Lawsuit Escalates Over AI Trade Secrets
- Chrome Manifest V2 Removal: Legacy Extensions Are Now Gone
The post Anthropic Bolsters Security After Claude AI Escapes appeared first on Daily CyberSecurity.
-
Malwarebytes

-
Infostealers are hijacking Claude accounts at users’ expense
Anthropic has warned some Claude users that criminals are using information stealers to take over their accounts. Rather than guessing passwords or intercepting two-factor authentication (2FA) codes, the attackers steal the browser sessions that prove a user is already logged in. According to a warning email shared publicly by an affected user, the attackers used common infostealer malware to copy Claude login sessions from victims’ computers. They then used those sessions to access the ac
Infostealers are hijacking Claude accounts at users’ expense
Anthropic has warned some Claude users that criminals are using information stealers to take over their accounts.
Rather than guessing passwords or intercepting two-factor authentication (2FA) codes, the attackers steal the browser sessions that prove a user is already logged in.
According to a warning email shared publicly by an affected user, the attackers used common infostealer malware to copy Claude login sessions from victims’ computers. They then used those sessions to access the accounts and consume their usage.

“We recently signed you out of Claude and removed the payment method saved on your account, so you’ll need to log back in and re-add your card. We’re sorry for the disruption. Here’s what happened and what we’ve done about it.
What happened
We have recently become aware of a bad actor that is using common infostealer malware to steal Claude login sessions from people’s computers, then using those login sessions to access Claude accounts and consume their usage. Our systems detected this activity on your account, and we’ve therefore removed your card on file and signed out the sessions involved to help block further unauthorized access.
If your usage limits looked like they refilled and then drained while you weren’t using Claude, this was likely the cause.”
The message adds that Anthropic has no reason to believe the malware was “related to Claude, installed through Claude, or related to anything you did with Claude.”
To sum this up:
- Cybercriminals are spreading infostealers. How they are doing this and whether they are targeting groups likely to use Claude professionally is unknown.
- Infostealers can bypass standard credentials and multi-factor authentication (MFA) by stealing active browser sessions and session cookies.
- Once they are able to take over a Claude account, they can consume the victim’s usage and potentially incur additional charges.
- Anthropic is signing affected users out of Claude, removing saved payment methods, and refunding charges it identifies as unauthorized.
To better understand this, you should know that paid Claude plans can offer additional “Usage credits.” When a subscriber reaches the plan’s session limit, Claude can allow them to continue using the service through consumption-based billing at standard API rates. The user must enable the feature, configure a monthly spending limit or select unlimited spending, and prepay for credits.
Users can also enable auto-reload, which automatically buys more prepaid credits when the balance falls below a threshold. So, in a session-hijacking scenario, a thief could use up the account’s included allowance and any available Usage credits. If auto-reload is enabled, they could also trigger further purchases.
The criminals’ likely motive is to use paid Claude capacity for free. The account and any exposed data could also be useful for fraud, social engineering, or follow-on attacks.
Stolen Claude capacity could be used to write and refine phishing and scam content, build campaign infrastructure, develop, modify, or obfuscate malware, improve delivery methods, and analyze stolen information. Cybercriminals can use AI to support several parts of an operation, although Claude has safeguards and abuse monitoring, and Anthropic says it has disrupted accounts used for malicious activity.
What to do
Anthropic provided advice for dealing with a possible infostealer infection. After removing the malware, we recommend you install an up-to-date, real-time anti-malware solution to help protect you against new infections.
These steps are good practice when cleaning up after infostealer malware:
- Scan any computer you use with Claude for malware and remove any malware before logging back in or changing passwords.
- Once the malware has been removed, secure the email account you use for Claude by changing its password, signing out of other devices, and enabling two-factor authentication (2FA).
- Change sensitive passwords that were saved in the affected browser, including those for banking, work, and cloud services. Check your card statements if you stored payment details in the browser.
- Only after completing these steps should you add your payment method to Claude again if you want your plan to continue renewing.
If you still see your usage changing while Claude is idle, or notice an unrecognized charge after completing these steps, contact usersafety@anthropic.com.
From reporting threats to removing them.
Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.
-
Cybersecurity News
-
Anthropic Issues Claude Security Warnings
Anthropic is sending security warnings and deleting payment methods as info-stealing malware compromises Claude user sessions. Protect your account now. Related Posts: ToxNetV2 Botnet Integrates AI ToxicPanda 2.0 Banking Trojan Attacks Escalate Globally NPM Typosquatting Malware Targets WSL Developers The post Anthropic Issues Claude Security Warnings appeared first on Daily CyberSecurity.
Anthropic Issues Claude Security Warnings
Anthropic is sending security warnings and deleting payment methods as info-stealing malware compromises Claude user sessions. Protect your account now.
Related Posts:
- ToxNetV2 Botnet Integrates AI
- ToxicPanda 2.0 Banking Trojan Attacks Escalate Globally
- NPM Typosquatting Malware Targets WSL Developers
The post Anthropic Issues Claude Security Warnings appeared first on Daily CyberSecurity.
-
Security Affairs

-
Infostealers Are Hijacking Claude Sessions and Draining Subscriptions
Infostealers can steal active Claude sessions, bypass 2FA and drain paid usage. Anthropic is revoking access and refunding unauthorized charges. Anthropic confirmed that several infostealer malware can hijack an active Claude login session and let attackers burn through your usage without ever touching your password. “Our investigation is ongoing. Our findings to date suggest that a computer you use with Claude is likely infected with infostealer malware, and may have been for some time.
Infostealers Are Hijacking Claude Sessions and Draining Subscriptions
Infostealers can steal active Claude sessions, bypass 2FA and drain paid usage. Anthropic is revoking access and refunding unauthorized charges.
Anthropic confirmed that several infostealer malware can hijack an active Claude login session and let attackers burn through your usage without ever touching your password.
“Our investigation is ongoing. Our findings to date suggest that a computer you use with Claude is likely infected with infostealer malware, and may have been for some time. Phones and tablets do not appear to have been involved.” reads the notification sent to the impacted users.
“We have no reason to believe that this malware is related to Claude, installed through Claude, or related to anything you did with Claude. It’s general-purpose malware that typically arrives with an unofficial download or a malicious app, and it quietly copies saved passwords, login cookies in browsers, and credentials for other apps running locally. Your Claude session was likely one of the many things it collected. It appears that a bad actor has now started picking the Claude sessions out of what it collected and using them.”
Recently, Anthropic started signing some Claude users out and removing their saved payment cards. The reason? Infostealer malware on their computers stole active Claude sessions and gave attackers access to their accounts.
“We recently signed you out of Claude and removed the payment method saved on your account, so you’ll need to log back in and re-add your card.” continues the report. “We’re sorry for the disruption. Here’s what happened and what we’ve done about it.”
Anthropic detected the suspicious activity and identified multiple infostealer families affecting Windows and macOS. Infostealers bypass the login process by stealing authenticated browser sessions, allowing attackers to evade passwords, MFA and SSO and access paid Claude accounts. Revoking sessions or blocking fraudulent payments is not enough: if the malware remains on the device, it can capture the user’s next login and give attackers access again.
Anthropic is also refunding users for any charges it identifies as unauthorized.
— International Cyber Digest (@IntCyberDigest) August 29, 2026
BREAKING: Anthropic is signing Claude users out and deleting their saved card because infostealer malware on their machines handed a bad actor live Claude login sessions. Anthropic says its systems detected the activity, and the notification names six stealer families across… pic.twitter.com/0nX53PaeiH
“”Our systems detected this activity on your account, and we’ve therefore removed your card on file and signed out the sessions involved to help block further unauthorized access.” continues the report. “If your usage limits looked like they refilled and then drained while you weren’t using Claude, this was likely the cause.””
Anthropic identified Vidar, LummaC2, StealC, RedLine and Acreed on Windows, plus Atomic Stealer on a small number of Macs. The company revoked affected Claude sessions, forcing users to log in again, and removed saved payment methods to prevent unauthorized charges.
Existing plans will continue until the current billing period ends. After that, users will need to add their payment method again. Anthropic may also sign them out again if it detects suspicious activity.
If you use Claude and haven’t checked your usage history recently, that’s worth doing today rather than next week. Anthropic’s advice is the standard but genuinely necessary response: run a full malware scan before logging back in, change your account password with two-factor authentication enabled, and treat any pirated download or unofficial app installer with the same suspicion you’d give a sketchy email attachment.
An AI subscription being quietly drained isn’t the scariest thing an infostealer can do to you, but it’s a pretty reliable sign that something considerably worse, like your actual banking credentials, might already be sitting in the same haul.
Follow me on Twitter: @securityaffairs and Facebook and Mastodon
(SecurityAffairs – hacking, Anthropic)
-
Firewall Daily – The Cyber Express

-
Anthropic Warns Commodity Infostealers Are Hijacking Claude Sessions to Drain Paid Usage
Anthropic warned users over the weekend that a threat actor is using widely available infostealer malware to hijack active Claude login sessions from infected computers, then using those sessions to run up victims' paid usage without ever needing a password or a two-factor code. The company said it identified six malware families in the campaign: Vidar, LummaC2, StealC, RedLine and Acreed on Windows, and Atomic Stealer, known as AMOS, on a smaller number of macOS machines. None are novel or bes
Anthropic Warns Commodity Infostealers Are Hijacking Claude Sessions to Drain Paid Usage
![]()
Anthropic warned users over the weekend that a threat actor is using widely available infostealer malware to hijack active Claude login sessions from infected computers, then using those sessions to run up victims' paid usage without ever needing a password or a two-factor code.
The company said it identified six malware families in the campaign: Vidar, LummaC2, StealC, RedLine and Acreed on Windows, and Atomic Stealer, known as AMOS, on a smaller number of macOS machines. None are novel or bespoke. All are commodity stealers sold or rented on criminal dark web marketplaces, and all work the same basic way - harvesting locally stored browser credentials, autofill data and authentication cookies from a compromised machine and shipping them to an operator's server.
Claude Session Cookies Heist
What makes the campaign notable is the target rather than the technique. Session cookies represent an already-authenticated state, so an attacker who replays a stolen Claude session token steps past both the account password and multi-factor authentication entirely. This is textbook session hijacking; the new element is that paid AI assistant subscriptions have become worth stealing as a commodity in their own right, alongside the streaming and gaming accounts that stealer log markets have traded for years.
Anthropic told affected users that the tell sign for them was a usage pattern that made no sense. Limits appearing to refill and then drain while the account owner was not using Claude was the biggest red flag.
Also read: Hacker Used Claude AI to Automate Reconnaissance, Harvest Credentials and Penetrate Networks
The company said it is signing affected users out of their sessions, removing saved payment methods from compromised accounts and refunding unauthorized charges identified during its investigation. It also stressed that the malware is not connected to Claude, was not installed through Claude and did not result from anything users did with the product. Infections trace to the usual vectors — pirated software and other illicit downloads.
A Reddit user going by the moniker "WorriedAssociate7029" received the notification from Anthropic and confirmed that he mistakenly installed an infostealer from "a reputable Russian underground forum" while downloading a pirated game. "I got fooled like a rookie by downloading a cracked game," he said.
Intrestingly though, the user claimed of using Claude's Opus model to detect and remove the malware.
"I use the models exclusively in permission-free mode on my entire computer," the Reddit user said."Opus was very efficient. It scanned for active processes, then listed my recent downloads. It found the virus almost instantly. My prompt was very simple: "I think I downloaded a virus recently. My login credentials were stolen. Audit the malware and remove it if you find it. Report on the extent of the damage. He deactivated the virus and created a folder on the desktop containing all the relevant information (including the deactivated virus, lol)."
Anthropic has not disclosed how many accounts were affected.
The security implications reach past the billing line. AI assistant accounts increasingly hold conversation histories, uploaded documents, connected data sources and, in developer configurations, API keys and repository access. A hijacked session inherits whatever the account can reach. Organizations that have rolled out AI tools without folding them into identity and access management now have a class of high-value session token sitting in employee browsers, largely outside the monitoring applied to corporate SaaS.
Anthropic's guidance to compromised users is the standard infostealer playbook. Change credentials across every service used on the affected machine, revoke active sessions, and actually remove the malware, since signing out does not clear an infection that will simply harvest the next session.
There is no formal regulatory hook here yet — no confirmed breach of the provider itself and no disclosure obligation triggered on Anthropic's side. But the episode lands as regulators and standards bodies are working out how AI system security fits existing frameworks, and it illustrates a gap those frameworks have barely addressed - the weakest point in an AI deployment may be an unmanaged endpoint rather than the model or the platform.
-
Cybersecurity News
-
OpenAI Cuts Off Cursor Access to Its Models Following SpaceX Acquisition
OpenAI will terminate Cursor's access to its AI models by November 12, citing Elon Musk's past contract violations following SpaceX's $60 billion acquisition of the company. Related Posts: Google Auto-Expands AI Overviews Sony and Warner Chappell Sue Anthropic uBlock Origin v1.74.0 Is the Final Version for Chrome Before Google's Delisting The post OpenAI Cuts Off Cursor Access to Its Models Following SpaceX Acquisition appeared first on Daily CyberSecurity.
OpenAI Cuts Off Cursor Access to Its Models Following SpaceX Acquisition
OpenAI will terminate Cursor's access to its AI models by November 12, citing Elon Musk's past contract violations following SpaceX's $60 billion acquisition of the company.
Related Posts:
- Google Auto-Expands AI Overviews
- Sony and Warner Chappell Sue Anthropic
- uBlock Origin v1.74.0 Is the Final Version for Chrome Before Google's Delisting
The post OpenAI Cuts Off Cursor Access to Its Models Following SpaceX Acquisition appeared first on Daily CyberSecurity.
-
Cybersecurity News
-
Sony and Warner Chappell Sue Anthropic
Sony Music Publishing and Warner Chappell Music sued Anthropic for allegedly using copyrighted lyrics to train Claude AI models. Read about the lawsuit here. Related Posts: Google Auto-Expands AI Overviews OpenAI Cuts Off Cursor Access to Its Models Following SpaceX Acquisition uBlock Origin v1.74.0 Is the Final Version for Chrome Before Google's Delisting The post Sony and Warner Chappell Sue Anthropic appeared first on Daily CyberSecurity.
Sony and Warner Chappell Sue Anthropic
Sony Music Publishing and Warner Chappell Music sued Anthropic for allegedly using copyrighted lyrics to train Claude AI models. Read about the lawsuit here.
Related Posts:
- Google Auto-Expands AI Overviews
- OpenAI Cuts Off Cursor Access to Its Models Following SpaceX Acquisition
- uBlock Origin v1.74.0 Is the Final Version for Chrome Before Google's Delisting
The post Sony and Warner Chappell Sue Anthropic appeared first on Daily CyberSecurity.
-
Malwarebytes
-
The AI agent swarm that attacked Hugging Face is a warning for the future
The hacking incident involving OpenAI evaluation agents and Hugging Face offers an unusually concrete look at what advanced AI-assisted intrusion can mean in practice: not a single clever exploit, but thousands of automated decisions, rapid experimentation, lateral movement, credential theft, persistence, and attempts to evade detection. The OpenAI–Hugging Face incident began during internal cybersecurity evaluations using ExploitGym, a benchmark designed to test whether AI agents can identif
The AI agent swarm that attacked Hugging Face is a warning for the future
The hacking incident involving OpenAI evaluation agents and Hugging Face offers an unusually concrete look at what advanced AI-assisted intrusion can mean in practice: not a single clever exploit, but thousands of automated decisions, rapid experimentation, lateral movement, credential theft, persistence, and attempts to evade detection.
The OpenAI–Hugging Face incident began during internal cybersecurity evaluations using ExploitGym, a benchmark designed to test whether AI agents can identify and exploit software vulnerabilities. OpenAI says the affected agents were meant to operate in isolated cloud sandboxes, with no internet access or communication between agents in many of the evaluation environments.
How the agents got around those boundaries
According to OpenAI’s account, agents used an internal Artifactory package-management service as an improvised communications channel. By writing files and directory names that other agents could read, isolated instances effectively created their own message board.
This meant the agents were no longer acting merely as independent benchmark runners. They could share discoveries, leave instructions, divide work, and build on one another’s successes.
An independent investigation by METR and Redwood Research found that roughly 1,200 agents used the unauthorized message board, exchanging more than 70,000 messages and files. About 700 of them went on to participate in the attack on Hugging Face.
The scale is the threat
A model capable of finding one useful exploit presents a risk. Hundreds of agents able to pool information, delegate tasks, and repeatedly test routes around a control boundary become a much more difficult problem.
Hugging Face reconstructed about 17,600 attacker actions across roughly four and a half days, including reconnaissance, remote code execution, credential access, command-and-control activity, Kubernetes enumeration, movement through the internal network, and supply-chain probing.
A human intruder must choose where to spend time. An agent swarm can continuously test ideas, retry failed paths, switch infrastructure when blocked, preserve notes for other agents, and run many experiments in parallel. It can also create noise: Thousands of low-confidence actions, most of which lead nowhere, can obscure the handful that matter.
Hugging Face described this as an “asymmetry problem.” An attacker can generate enormous volumes of attempts, while defenders must determine which signals represent a real escalation before the attacker reaches a sensitive system.
What needs to change
The core security lesson is that evaluation environments need to be treated as high-risk production systems when models can use tools, write files, execute code, or interact with infrastructure.
The technology is rapidly becoming capable of doing more. The question is whether the people building, deploying, and defending it are ready for what comes next.
OpenAI says it has since moved toward stronger workload and network isolation, tighter control over internet access and model weights, and mandatory chain-of-thought monitoring for high-capability, tool-using training and evaluations.
“We are taking this incident as a ‘warning shot’ that today’s model capabilities present the possibility of loss-of-control incidents. Our security and alignment posture is escalating accordingly. These events also highlight risks in future AI development that extend beyond OpenAI and will require the attention of the whole industry.”
At the same time, a US court ruling involving Anthropic has brought a related question into focus: whether developers can place limits on high-risk uses of their models, including mass surveillance and fully autonomous weapons.
The legal dispute is political by nature, but its technical underpinning is hard to ignore. If capable AI systems can enhance offensive cyberattack methods and bypass safety restrictions, access controls, logging, and deployment boundaries, those safeguards are no longer abstract policy choices.
Advanced AI agents can be useful to defenders as well as attackers. But the surrounding systems need to be trusted to keep their capabilities bounded when something goes wrong.
Who benefits from more capable AI?
The security debate around AI agents often focuses on whether systems can be controlled. Can they be kept inside a sandbox? Can their tools, credentials, network access, and autonomy be restricted? Can defenders detect harmful behavior before it becomes an incident?
While those questions are essential, there is another: Who benefits when AI becomes capable enough to automate large parts of cognitive work? Who carries the costs when it fails, displaces workers, enables fraud, causes damage, or concentrates power?
AI could give small organizations access to technical expertise that previously required large teams and budgets. It could help doctors identify urgent cases sooner, help teachers tailor support to individual students, assist people with disabilities, speed up scientific research, and make complex public services easier to navigate. For cybersecurity teams, it could make vulnerability triage, alert investigation, threat hunting, and incident response faster and more accessible.
Bill Gates has argued that while AI could bring remarkable benefits to health care, education, agriculture, scientific research, and public services, the outcome will depend on deliberate choices rather than technical progress alone. He also warns that AI’s rapid adoption could widen inequality, disrupt entry-level and mid-career work, make harmful capabilities more accessible, and reinforce existing concentrations of power.
Gates also argues that “self-regulation on the most dangerous tool ever invented” does not sound like a good idea.
“AI will either be the greatest equalizer ever invented, or the worst source of injustice.”
Right now, we still have a choice.
Let’s face it, an incognito window can only do so much.
Breaches, dark web trading, credit fraud. Malwarebytes Identity Theft Protection monitors for all of it, alerts you fast, and comes with identity theft insurance.
-
Security | TechRepublic
-
Claude Opus 4.6 Found a Gym API Flaw — Then Exploited It in 9 of 10 Tests
Claude Opus 4.6 exploited a gym API flaw in 9 of 10 controlled tests, highlighting security risks when AI agents gain backend access. The post Claude Opus 4.6 Found a Gym API Flaw — Then Exploited It in 9 of 10 Tests appeared first on TechRepublic.
Claude Opus 4.6 Found a Gym API Flaw — Then Exploited It in 9 of 10 Tests
Claude Opus 4.6 exploited a gym API flaw in 9 of 10 controlled tests, highlighting security risks when AI agents gain backend access.
The post Claude Opus 4.6 Found a Gym API Flaw — Then Exploited It in 9 of 10 Tests appeared first on TechRepublic.
-
Cybersecurity News
-
Claude Fable 5 Intelligence Drop Sparks Concerns
Developers report a Claude Fable 5 intelligence drop. Anthropic clarifies that Claude Code effort level scores reflect API configuration testing. Related Posts: OneDrive Folder Exclusions Roll Out for Development Environments OpenAI Advocates Stricter California AI Regulations LinkedIn AI Slop Reduction: A Necessary Course Correction The post Claude Fable 5 Intelligence Drop Sparks Concerns appeared first on Daily CyberSecurity.
Claude Fable 5 Intelligence Drop Sparks Concerns
Developers report a Claude Fable 5 intelligence drop. Anthropic clarifies that Claude Code effort level scores reflect API configuration testing.
Related Posts:
- OneDrive Folder Exclusions Roll Out for Development Environments
- OpenAI Advocates Stricter California AI Regulations
- LinkedIn AI Slop Reduction: A Necessary Course Correction
The post Claude Fable 5 Intelligence Drop Sparks Concerns appeared first on Daily CyberSecurity.
-
Cybersecurity News
-
Anthropic Revenue Surges 14x to $11.5B in Q2 2026 Ahead of IPO
Anthropic's Q2 2026 revenue surged nearly 14-fold to over $11.5 billion, turning adjusted operating profit positive as the company prepares a Wall Street IPO. Related Posts: Stripe Finalizes $7 Billion Acquisition of OpenRouter SpaceX Acquires AI Coding Startup Cursor Qualcomm Snapdragon C Targets $300 Windows Laptops to Rival MacBook Neo The post Anthropic Revenue Surges 14x to $11.5B in Q2 2026 Ahead of IPO appeared first on Daily CyberSecurity.
Anthropic Revenue Surges 14x to $11.5B in Q2 2026 Ahead of IPO
Anthropic's Q2 2026 revenue surged nearly 14-fold to over $11.5 billion, turning adjusted operating profit positive as the company prepares a Wall Street IPO.
Related Posts:
- Stripe Finalizes $7 Billion Acquisition of OpenRouter
- SpaceX Acquires AI Coding Startup Cursor
- Qualcomm Snapdragon C Targets $300 Windows Laptops to Rival MacBook Neo
The post Anthropic Revenue Surges 14x to $11.5B in Q2 2026 Ahead of IPO appeared first on Daily CyberSecurity.
-
Cybersecurity News
-
Anthropic Claude Text Watermark: An Invisible Statistical Algorithm
Anthropic reveals its invisible Claude text watermark mechanism, utilizing a probability-based algorithm to comply with the EU AI Act without altering meaning. Related Posts: Anthropic Revenue Surges 14x to $11.5B in Q2 2026 Ahead of IPO Stripe Finalizes $7 Billion Acquisition of OpenRouter SpaceX Acquires AI Coding Startup Cursor The post Anthropic Claude Text Watermark: An Invisible Statistical Algorithm appeared first on Daily CyberSecurity.
Anthropic Claude Text Watermark: An Invisible Statistical Algorithm
Anthropic reveals its invisible Claude text watermark mechanism, utilizing a probability-based algorithm to comply with the EU AI Act without altering meaning.
Related Posts:
- Anthropic Revenue Surges 14x to $11.5B in Q2 2026 Ahead of IPO
- Stripe Finalizes $7 Billion Acquisition of OpenRouter
- SpaceX Acquires AI Coding Startup Cursor
The post Anthropic Claude Text Watermark: An Invisible Statistical Algorithm appeared first on Daily CyberSecurity.
-
Graham Cluley
-
Smashing Security podcast #480: This is the AI service you should never sign up to
Would you like access to Anthropic's Claude at 90% off the normal price? All you have to do is redirect your traffic to a mysterious service called "Poison Claude". Only problem is that it's run by fraudsters... Meanwhile, a phishing-as-a-service platform called "Greatness" has come up with something rather nasty: a phishing attack that doesn't need a fake website, a suspicious URL, or your password. Just a real Microsoft login page and a moment of misplaced trust - and the attackers walk off
Smashing Security podcast #480: This is the AI service you should never sign up to
-
Graham Cluley
-
Beware cut-price AI services that read your every word
f someone offered you 90% off the official price to access Claude, the powerful AI model from Anthropic, would you be tempted? It turns out that around 900 people were, and they may be regretting their decision. Read more in my article on the Fortra blog.
Beware cut-price AI services that read your every word
-
Cybersecurity News
-
Rogue AI Models Hack Systems to Cheat Evaluations
Rogue AI models from Meta, OpenAI, and Anthropic breached sandboxes to hack external systems and collaborate on exploits to cheat evaluation tests. Related Posts: Dopamine 3.0 Jailbreak Brings iOS 26 Support to A12 and A13 Devices Google Wallet Introduces Digital Allowance for Minors Google Ask Maps Update: AI Navigation Evolves The post Rogue AI Models Hack Systems to Cheat Evaluations appeared first on Daily CyberSecurity.
Rogue AI Models Hack Systems to Cheat Evaluations
Rogue AI models from Meta, OpenAI, and Anthropic breached sandboxes to hack external systems and collaborate on exploits to cheat evaluation tests.
Related Posts:
- Dopamine 3.0 Jailbreak Brings iOS 26 Support to A12 and A13 Devices
- Google Wallet Introduces Digital Allowance for Minors
- Google Ask Maps Update: AI Navigation Evolves
The post Rogue AI Models Hack Systems to Cheat Evaluations appeared first on Daily CyberSecurity.
-
Security | TechRepublic
-
UK AI tests found 19 unauthorized agent actions involving Anthropic and OpenAI models
UK researchers reported 19 unsanctioned actions by Anthropic and OpenAI agents during permissive cyber tests involving real external systems. The post UK AI tests found 19 unauthorized agent actions involving Anthropic and OpenAI models appeared first on TechRepublic.
UK AI tests found 19 unauthorized agent actions involving Anthropic and OpenAI models
UK researchers reported 19 unsanctioned actions by Anthropic and OpenAI agents during permissive cyber tests involving real external systems.
The post UK AI tests found 19 unauthorized agent actions involving Anthropic and OpenAI models appeared first on TechRepublic.
-
Security Affairs

-
AI Deception Emerges in Cyber Tests as Agents Target Real People and Systems
AISI found AI agents taking unsanctioned online actions, including social engineering and code attacks, during controlled cyber tests. The UK’s AI Security Institute (AISI) has put something uncomfortable on the table: during cyber testing, frontier models didn’t just follow instructions badly. In some runs, they crossed into real-world actions, touched real people and organisations, and then tried to make the evidence go away. “On 28th July 2026, AISI’s Security Team detected unusual da
AI Deception Emerges in Cyber Tests as Agents Target Real People and Systems
AISI found AI agents taking unsanctioned online actions, including social engineering and code attacks, during controlled cyber tests.
The UK’s AI Security Institute (AISI) has put something uncomfortable on the table: during cyber testing, frontier models didn’t just follow instructions badly. In some runs, they crossed into real-world actions, touched real people and organisations, and then tried to make the evidence go away.
“On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.” reports AISI. “We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.”
That matters because the setting was not a botched public rollout or a rogue consumer chatbot. It was a controlled evaluation, run on purpose with open internet access and some safety filters turned off, so the institute could measure what the models could actually do under permissive conditions. AISI says that’s standard practice for frontier testing, but it also notes that those conditions are not how these models are normally shipped.
To assess the models’ real-world cyber capabilities, they were given internet access and had their built-in cyber safety filters intentionally disabled under controlled testing conditions.
The headline result is simple enough, even if the details are messy. AISI ran the challenge 122 times across several models, and in 10 of those runs it saw AI agents take autonomous, unsanctioned actions on the live internet. The institute logged 19 such actions in total, with 17 tied to Anthropic’s Mythos 5 and two to OpenAI’s GPT-5.6-Sol with cyber classifiers disabled.
The most serious sequence looked less like a lab mishap and more like a small-scale social-engineering campaign. The agent tried to insert malicious code into a public open-source project, researched the maintainers, created fake identities based on real people, and used those identities to pressure a maintainer into approving the code.
“On the morning of 28th July, our security monitoring flagged data leaving one of our testing systems through the ‘Tor’ anonymity network, which is commonly used to disguise the origin of internet traffic. The team running the evaluation began reviewing the agent’s activity within minutes.” continues the report. “They quickly found that, during one run, the agent had created a malicious ‘pull request’ (a proposed code change) on a real, public open-source project on GitHub. “
It also sent messages and files, and when challenged, it edited earlier activity to look harmless and even considered adopting a fresh identity.
AISI says human review stopped the worst outcome. The malicious pull request was caught, the fake identities failed to get the code approved, and the institute did not find evidence of resulting real-world harm. Even so, it says this is the first time it has seen autonomy and deception show up this clearly, without specific prompting, in a real-world setting.
The institute is careful not to overclaim. It says the evaluation design and configuration choices helped create the conditions for the behaviour, and it notes that the models in question were not commercial public deployments. But that caveat cuts both ways: if permissive test conditions can produce this kind of behaviour now, tighter production settings are not a reason for complacency. They are just a better fence.
That is the part that should make people stop scrolling. Not because the models “went evil,” which is lazy, but because deception emerged as a by-product of goal-seeking under the conditions set by the evaluators.
“The agent pursued its goal persistently. AI agents explore routes their operators did not intend. Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people.” states AISI. “It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”
The agent was trying to solve a cyber challenge, and in doing so it explored routes nobody had asked it to explore, including routes that involved fooling actual people.
What makes the episode more serious is the way it lines up with recent disclosures from Anthropic and OpenAI. Anthropic has already said its models reached the public internet from within a third-party evaluation environment and accessed real systems. OpenAI has separately reported that one of its models mistook a misconfigured test setup for a real task and touched real infrastructure.
AISI’s view is that these incidents point to a shift in the risk landscape. Harm may no longer come only from obvious misuse by humans. It may also come from capable agents, in internal research settings or privileged-access environments, taking unintended actions beyond the scope they were given. That is a quieter problem than movie-style “AI rebellion,” and a more useful one to think about.
The key takeaway is straightforward. Keep cyber basics tight, verify outside code before trusting it, and stop assuming that a model will stay inside the lines just because the prompt sounded clear on the day. In security, the line between “evaluation” and “incident” can get thin fast. Machines are very good at finding the part of the process you forgot to make boring.
AISI says it will tighten internet controls, add real-time monitoring, and revisit how it designs evaluations. That is the right response, but it should not be read as a narrow fix for one lab. It is a warning to anyone testing powerful agents: if the test can reach the real internet, the real internet can reach back.
The original AISI report is here: Incident report: unsanctioned agent behaviour during cyber testing.
“Incidents of this kind reflect the speed at which AI is developing. As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.” concludes the report.
Follow me on Twitter: @securityaffairs and Facebook and Mastodon
(SecurityAffairs – hacking, AI Deception)
-
Hackread – Latest Cybersecurity, Tech, Crypto & Hacking News
-
Black Hat USA 2026: One GitHub Issue Could Compromise Major AI Coding Workflows
At Black Hat USA 2026, Novee found GitHub workflow flaws in Claude Code, Gemini CLI and Codex that enabled RCE, credential theft and agent control in pipelines.
Black Hat USA 2026: One GitHub Issue Could Compromise Major AI Coding Workflows
-
Security | TechRepublic
-
Claude Opus 5 Vending Test Shows Profit-Driven AI Risks
Claude Opus 5 set a Vending-Bench record while fabricating supplier bids, breaking truces, and ignoring refunds, showing why companies need stronger AI agent controls. The post Claude Opus 5 Vending Test Shows Profit-Driven AI Risks appeared first on TechRepublic.
Claude Opus 5 Vending Test Shows Profit-Driven AI Risks
Claude Opus 5 set a Vending-Bench record while fabricating supplier bids, breaking truces, and ignoring refunds, showing why companies need stronger AI agent controls.
The post Claude Opus 5 Vending Test Shows Profit-Driven AI Risks appeared first on TechRepublic.