Visualização de leitura

Flirty OnlyFans promoters on X may be using AI to appear human

In a recent post, we looked at reports of League of Legends players receiving suspicious friend requests shortly after matches. The accounts quickly steered the conversation toward Discord, where they promoted paid adult-content pages.

At the time, one unanswered question was how much of those conversations was automated. Were people working from scripts behind the accounts? Were they conventional, rules-based chatbots following a limited decision tree? Or were they using generative AI to produce more natural and flexible replies?

People are more likely to trust someone they believe is personally interested in them. AI can create that impression across many conversations at once, making it easier to persuade people to click links, spend money, or share personal or intimate information. The same approach could also be used for more harmful fraud, including romance scams and sextortion.

Now, developer Álvaro Martínez Majado has investigated several flirty accounts promoting OnlyFans pages on X to see whether their replies were scripted, generated by AI, or written by people. Majado, president of digital rights organization Protecció de la Frontera Electrònica, shared his evidence with Malwarebytes. Although it does not provide a definitive answer, it shows the accounts following rigid conversation scripts while also responding dynamically to unusual requests. The signs that once suggested a real person, such as an unusual reply or personalized voice note, can no longer be trusted.

The script goes on and on

Majado interacted with several accounts on X that followed a familiar pattern. They opened with similar casual, flirtatious language and asked broadly the same qualifying questions: where he lived, what he liked, and what he did for work.

That repetitive structure is exactly what we would expect from a commercially motivated messaging campaign. The goal is not necessarily to have a meaningful conversation. It is to identify people likely to respond, establish rapport, and eventually move them toward a paid page or another destination controlled by the operator.

The accounts also stayed in character when faced with obvious attempts to expose them as bots. That could be the result of hard-coded replies, guardrails around an AI system, or both.

Different accounts followed the same conversation pattern
They claimed to live in the same city as the recipient

But some later interactions were more difficult to explain as a simple bank of canned flirtatious responses.

One of the more interesting tests involved an instruction written as ASCII hexadecimal rather than ordinary text. The encoded message told the account to reply with a single word: “Pineapple.”

According to the screenshots supplied to Malwarebytes, the account responded with “Pineapple” in ordinary text.

An account followed an instruction encoded in hexadecimal
An account followed an instruction encoded in hexadecimal

That does not conclusively prove which technology was used. It does not identify a model, a provider, or the people behind the accounts. But it is consistent with an automated system capable of interpreting an encoded instruction and changing its output accordingly.

A simple scripted bot could theoretically include a hexadecimal decoder, of course. But that would be unusual in a basic adult-content promotional bot, especially when combined with other examples of flexible and sometimes error-prone responses.

In another interaction, Majado asked an account to provide a reply of exactly 12 characters. It responded with “Imnotabotfr”—an 11-character answer—then appeared to recognize its own counting mistake.

The account failed an exact character-count test, but recognized its error
The account failed an exact character-count test, but recognized its error

Anyone who has spent time experimenting with large language models may recognize the pattern. Language models can be very good at generating natural-sounding text while still making surprisingly basic mistakes involving character counts, word counts, and other exact constraints.

A deliberately designed bot could imitate this kind of mistake, so it is not proof of AI. But the account understood an unexpected instruction, attempted to follow it, and reacted when it got the answer wrong. That suggests it may have been generating replies dynamically rather than choosing from a list of pre-written responses. Such accounts can adapt to conversations, making them harder to identify as automated.

Voice notes do not settle the question

The accounts also sent voice notes. In one example, an account read aloud a Unix timestamp supplied during the conversation. In another, it spoke a requested username.

The accounts sent voice notes containing requested information
The accounts sent voice notes containing requested information

These responses show that the accounts could incorporate unusual information from a conversation into audio messages. They do not tell us whether a person recorded the clips or a text-to-speech tool generated them.

Text-to-speech tools can generate short, convincing clips quickly and cheaply. An operator can generate them manually, but the process can also be automated: Take a message, pass selected text to a voice-generation service, and send the resulting audio back to the recipient.

Here’s one of those voice notes. Is it a very flirty girl, or AI-generated? Have a listen and see what you think:

The supplied audio metadata offered a possible clue about the tools involved, but it is not enough to attribute the voice notes to a particular service. Platforms and other software can alter audio files and their metadata.

The more important point is that the voice notes were personalized and continued even after the interaction appeared unlikely to lead to a sale. That is consistent with a system designed to keep conversations moving without requiring a human to supervise each one.

AI does not replace the funnel

The evidence does not mean every message from every flirty spam account is written by an AI. Nor does it establish that the X accounts are operated by the same people targeting League of Legends players.

What it does suggest is a plausible hybrid model, supported by identical replies across different accounts alongside more flexible responses.

The repetitive parts of the operation can be scripted: opening messages, questions about location and interests, links, and attempts to move people to another platform. An AI-powered conversational layer could then make the exchange feel less repetitive when someone asks unexpected questions, changes the subject, or tries to test whether the account is real.

This combination makes practical sense for spammers. Scripts provide consistency and keep the conversation directed toward conversion. Generative AI helps the account handle the unpredictable parts of talking to real people.

It also means that traditional “bot tests” are becoming less useful. Asking an account to answer an unusual question, decode a message, or send a voice note may no longer distinguish a real person from a fake one.

How to stay safe

Treat unsolicited flirtatious messages with caution, especially when they quickly become transactional.

  • Do not assume a personalized response or voice message proves an account is genuine.
  • Be wary if a new contact repeatedly tries to move you to Discord, Telegram, Signal, another messaging app, or a paid-content platform.
  • Do not send money, gift cards, cryptocurrency, intimate images, identity documents, or account credentials to someone you only know online.
  • Avoid opening links or downloading files from accounts that contacted you unexpectedly.
  • Reverse-image-search profile photos and look for copied biographies, reused images, or accounts with very limited genuine activity.
  • Report suspicious accounts to the platform, particularly if they impersonate someone, send malicious links, or pressure users for money or explicit material.

Whether it’s a human, a chatbot, or an AI agent you’re talking to is an important question. AI could make these operations more convincing and much easier to scale. One operator could hold flirtatious conversations with many people, adapting the messages without personally managing every exchange.

That makes it easier to create a false sense of connection and persuade people to click links, pay for content, or share personal or intimate information.

The line between a scripted spam account and a responsive conversational partner is getting harder to see. Judge the account by what it wants you to do, not by how convincingly it talks.


Something feel off? Check it before you click.  

Malwarebytes Scam Guard helps you analyze suspicious links, texts, and screenshots instantly.  

Available with Malwarebytes Premium Security for all your devices, and in the Malwarebytes app for iOS and Android.  

Try it free → 

Your AI chats could be used in court

You might tell an AI chatbot secrets that you wouldn’t divulge to your closest friends. If you do, though, beware: They could end up as evidence in court.

An article in the Washington Post this week highlighted several cases in which people had discussed sensitive information with AI systems like Claude and ChatGPT, only to have their conversations obtained by prosecutors or opposing lawyers.

Lawyers can get access to your chatbot conversations from AI services like ChatGPT because they aren’t privileged in the same way that, say, a conversation with your lawyer or a doctor would be.

Reporters at the paper found chatbot logs cited in 12 court cases in the past two years. They also found statistics from OpenAI that supported a rising trend in data disclosures. The company, which operates ChatGPT, disclosed the content of more than 80 user accounts in the last six months of 2025. That was more than four times as many as in the second half of 2024.

Cases are piling up

With people asking AI for all kinds of advice, it’s little wonder that lawyers are coming after that data too. Sometimes, it emerges because users consent to a search. The Washington Post mentions one university student who asked ChatGPT in a panic whether people might work out that he had damaged 17 cars in a campus parking lot. He then handed his phone over to police for a search. A teen suing big tech companies over social media addiction saw his own ChatGPT history drawn into discovery.

Deleting your chats isn’t watertight protection either. In The New York Times’ copyright lawsuit against OpenAI over collecting its content for training data, a judge ordered the AI company to preserve chat logs, including ones that users had asked it to erase. OpenAI complained that users were being “forced to forgo the privacy protections OpenAI has painstakingly put in place.” The company had to keep that data even though it had agreed to delete it under the EU’s General Data Protection Regulation (GDPR) and California’s privacy laws.

Incidents like these involve responses to legal requests, but AI companies don’t always wait for a subpoena. OpenAI’s policy allows its reviewers to refer conversations to law enforcement whenever they identify “an imminent and credible risk of harm to others.” The Washington Post reported an incident in which OpenAI contacted police after a ChatGPT user in Palm Beach County, Florida, repeatedly described plans to harm an ex-girlfriend.

Technology companies have been handing over all kinds of data beyond AI chats to law enforcement and litigants for years. Google, Meta, and Apple shared details of 3.16 million US user accounts between 2014 and 2024, with substantial increases in the number of records shared annually during that period.

Every time a new technology emerges, litigants will go after it for data. In 2019, police issued a subpoena for audio recordings from an Amazon Echo owned by a Florida man charged with murdering his girlfriend.

What to do

We’d all like to think that true friends will carry our secrets to the grave. But AI is not your friend. Or your doctor, or your lawyer. Treat all chats as records that could potentially be disclosed in a legal case. They might feel like informal conversations, but you should assume that each one creates a written record, even if there’s a delete button.

Be careful about what you share. If the topic is one you’d normally raise only with a doctor or a lawyer, then raise it with a doctor or a lawyer, not AI. Communications with lawyers may be protected by attorney-client privilege, while medical information is subject to confidentiality and privacy protections. Chatbot conversations aren’t.

Finally, be cautious beyond AI. Everything from ill-advised social media posts to private messages might also find its way into police hands. In 2022, for example, Facebook handed over private messages between a mother and daughter to police investigating an illegal abortion case.

So think twice before posting anything sensitive.


From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

Infostealers are hijacking Claude accounts at users’ expense

Anthropic has warned some Claude users that criminals are using information stealers to take over their accounts.

Rather than guessing passwords or intercepting two-factor authentication (2FA) codes, the attackers steal the browser sessions that prove a user is already logged in.

According to a warning email shared publicly by an affected user, the attackers used common infostealer malware to copy Claude login sessions from victims’ computers. They then used those sessions to access the accounts and consume their usage.

Warning from Anthropic

“We recently signed you out of Claude and removed the payment method saved on your account, so you’ll need to log back in and re-add your card. We’re sorry for the disruption. Here’s what happened and what we’ve done about it.

What happened

We have recently become aware of a bad actor that is using common infostealer malware to steal Claude login sessions from people’s computers, then using those login sessions to access Claude accounts and consume their usage. Our systems detected this activity on your account, and we’ve therefore removed your card on file and signed out the sessions involved to help block further unauthorized access.

If your usage limits looked like they refilled and then drained while you weren’t using Claude, this was likely the cause.”

The message adds that Anthropic has no reason to believe the malware was “related to Claude, installed through Claude, or related to anything you did with Claude.”  

To sum this up:

  • Cybercriminals are spreading infostealers. How they are doing this and whether they are targeting groups likely to use Claude professionally is unknown.
  • Infostealers can bypass standard credentials and multi-factor authentication (MFA) by stealing active browser sessions and session cookies.
  • Once they are able to take over a Claude account, they can consume the victim’s usage and potentially incur additional charges.
  • Anthropic is signing affected users out of Claude, removing saved payment methods, and refunding charges it identifies as unauthorized.

To better understand this, you should know that paid Claude plans can offer additional “Usage credits.” When a subscriber reaches the plan’s session limit, Claude can allow them to continue using the service through consumption-based billing at standard API rates. The user must enable the feature, configure a monthly spending limit or select unlimited spending, and prepay for credits.

Users can also enable auto-reload, which automatically buys more prepaid credits when the balance falls below a threshold. So, in a session-hijacking scenario, a thief could use up the account’s included allowance and any available Usage credits. If auto-reload is enabled, they could also trigger further purchases.

The criminals’ likely motive is to use paid Claude capacity for free. The account and any exposed data could also be useful for fraud, social engineering, or follow-on attacks.

Stolen Claude capacity could be used to write and refine phishing and scam content, build campaign infrastructure, develop, modify, or obfuscate malware, improve delivery methods, and analyze stolen information. Cybercriminals can use AI to support several parts of an operation, although Claude has safeguards and abuse monitoring, and Anthropic says it has disrupted accounts used for malicious activity.

What to do

Anthropic provided advice for dealing with a possible infostealer infection. After removing the malware, we recommend you install an up-to-date, real-time anti-malware solution to help protect you against new infections.

These steps are good practice when cleaning up after infostealer malware:

  1. Scan any computer you use with Claude for malware and remove any malware before logging back in or changing passwords.
  2. Once the malware has been removed, secure the email account you use for Claude by changing its password, signing out of other devices, and enabling two-factor authentication (2FA).
  3. Change sensitive passwords that were saved in the affected browser, including those for banking, work, and cloud services. Check your card statements if you stored payment details in the browser.
  4. Only after completing these steps should you add your payment method to Claude again if you want your plan to continue renewing.

If you still see your usage changing while Claude is idle, or notice an unrecognized charge after completing these steps, contact usersafety@anthropic.com.


From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

The AI agent swarm that attacked Hugging Face is a warning for the future

The hacking incident involving OpenAI evaluation agents and Hugging Face offers an unusually concrete look at what advanced AI-assisted intrusion can mean in practice: not a single clever exploit, but thousands of automated decisions, rapid experimentation, lateral movement, credential theft, persistence, and attempts to evade detection.

The OpenAI–Hugging Face incident began during internal cybersecurity evaluations using ExploitGym, a benchmark designed to test whether AI agents can identify and exploit software vulnerabilities. OpenAI says the affected agents were meant to operate in isolated cloud sandboxes, with no internet access or communication between agents in many of the evaluation environments.

How the agents got around those boundaries

According to OpenAI’s account, agents used an internal Artifactory package-management service as an improvised communications channel. By writing files and directory names that other agents could read, isolated instances effectively created their own message board.

This meant the agents were no longer acting merely as independent benchmark runners. They could share discoveries, leave instructions, divide work, and build on one another’s successes.

An independent investigation by METR and Redwood Research found that roughly 1,200 agents used the unauthorized message board, exchanging more than 70,000 messages and files. About 700 of them went on to participate in the attack on Hugging Face.

The scale is the threat

A model capable of finding one useful exploit presents a risk. Hundreds of agents able to pool information, delegate tasks, and repeatedly test routes around a control boundary become a much more difficult problem.

Hugging Face reconstructed about 17,600 attacker actions across roughly four and a half days, including reconnaissance, remote code execution, credential access, command-and-control activity, Kubernetes enumeration, movement through the internal network, and supply-chain probing.

A human intruder must choose where to spend time. An agent swarm can continuously test ideas, retry failed paths, switch infrastructure when blocked, preserve notes for other agents, and run many experiments in parallel. It can also create noise: Thousands of low-confidence actions, most of which lead nowhere, can obscure the handful that matter.

Hugging Face described this as an “asymmetry problem.” An attacker can generate enormous volumes of attempts, while defenders must determine which signals represent a real escalation before the attacker reaches a sensitive system.

What needs to change

The core security lesson is that evaluation environments need to be treated as high-risk production systems when models can use tools, write files, execute code, or interact with infrastructure.

The technology is rapidly becoming capable of doing more. The question is whether the people building, deploying, and defending it are ready for what comes next.

OpenAI says it has since moved toward stronger workload and network isolation, tighter control over internet access and model weights, and mandatory chain-of-thought monitoring for high-capability, tool-using training and evaluations.

“We are taking this incident as a ‘warning shot’ that today’s model capabilities present the possibility of loss-of-control incidents. Our security and alignment posture is escalating accordingly. These events also highlight risks in future AI development that extend beyond OpenAI and will require the attention of the whole industry.”

At the same time, a US court ruling involving Anthropic has brought a related question into focus: whether developers can place limits on high-risk uses of their models, including mass surveillance and fully autonomous weapons.

The legal dispute is political by nature, but its technical underpinning is hard to ignore. If capable AI systems can enhance offensive cyberattack methods and bypass safety restrictions, access controls, logging, and deployment boundaries, those safeguards are no longer abstract policy choices.

Advanced AI agents can be useful to defenders as well as attackers. But the surrounding systems need to be trusted to keep their capabilities bounded when something goes wrong.

Who benefits from more capable AI?

The security debate around AI agents often focuses on whether systems can be controlled. Can they be kept inside a sandbox? Can their tools, credentials, network access, and autonomy be restricted? Can defenders detect harmful behavior before it becomes an incident?

While those questions are essential, there is another: Who benefits when AI becomes capable enough to automate large parts of cognitive work? Who carries the costs when it fails, displaces workers, enables fraud, causes damage, or concentrates power?

AI could give small organizations access to technical expertise that previously required large teams and budgets. It could help doctors identify urgent cases sooner, help teachers tailor support to individual students, assist people with disabilities, speed up scientific research, and make complex public services easier to navigate. For cybersecurity teams, it could make vulnerability triage, alert investigation, threat hunting, and incident response faster and more accessible.

Bill Gates has argued that while AI could bring remarkable benefits to health care, education, agriculture, scientific research, and public services, the outcome will depend on deliberate choices rather than technical progress alone. He also warns that AI’s rapid adoption could widen inequality, disrupt entry-level and mid-career work, make harmful capabilities more accessible, and reinforce existing concentrations of power.

Gates also argues that “self-regulation on the most dangerous tool ever invented” does not sound like a good idea.

“AI will either be the greatest equalizer ever invented, or the worst source of injustice.”

Right now, we still have a choice.


Let’s face it, an incognito window can only do so much. 
 
Breaches, dark web trading, credit fraud. Malwarebytes Identity Theft Protection monitors for all of it, alerts you fast, and comes with identity theft insurance. 

Grok fooled into stealing user chat, location data, and more

A new type of prompt injection attack shows why giving AI assistants access to browsers, code tools, and private data deserves extra caution.

AI researchers describe “Cryptographic Context Injection”—an attack that hides malicious instructions inside encrypted data. The AI is then persuaded to decrypt that data using its own code-execution tool. As a result, the AI may treat the resulting text as if it were trustworthy internal information.

A prompt injection is a bit like leaving a fake instruction inside a document for an AI assistant to read. Instead of following only the user’s request, the assistant may be tricked into following an attacker’s instructions hidden in a webpage, email, or file.

As we reported months ago, experts have warned that prompt injection attacks are a problem that may never be fixed. Prompt injection works because AI models can’t reliably tell the difference between the legitimate instructions and an attacker’s instructions, so they sometimes obey the wrong ones.

To reduce this risk, AI providers set up their models with guardrails: protections designed to stop AI systems from doing things they shouldn’t, either intentionally or unintentionally.

What the researchers found was that malicious instructions could be hidden from some AI guardrails by encrypting them. The AI itself could then be tricked into decrypting those instructions using its coding tools.

By the time the instructions became readable, they had already made it past the initial security checks. The AI could then mistake them for legitimate instructions and follow them.

It’s a bit like hiding malicious instructions in a language the security system can’t understand. The AI translates them only after they’ve passed the security checks, then may follow what they say.

The researchers tested their method against two AI agents, with different results. In Grok, the researchers say the attack could steal information including the user’s name, approximate location, subscription tier, and conversation history. In Gemini, they used the technique to bypass safety controls and generate content the model would normally not do.

“In Grok, an ordinary ‘summarize this page’ steals the user’s chat data with no click or warning. In Gemini, it produces content the model normally refuses. Both are live production systems.”

The researchers did not provide full details because xAI had not taken action after the flaw in Grok was reported to it in June 2026. Gemini, on the other hand, has made improvements, but has still not fully closed the hole.

How to stay safe

An AI assistant may be helpful, but it should not automatically be trusted with sensitive data or powerful tools.

  • Treat AI summaries of unfamiliar webpages, documents, and shared links with caution, especially when the assistant can browse or run code.
  • Do not paste passwords, recovery codes, API keys, financial information, or sensitive health and work details into AI chats unless you understand how that information will be handled.
  • Review an AI assistant’s connected tools and permissions. Remove access it doesn’t need, particularly email, cloud storage, source-code repositories, and external integrations.
  • Be skeptical if an AI tool asks to decrypt, decode, run a script, open a new link, or upload data as part of a seemingly ordinary task.
  • Keep browser and AI applications updated, and check vendor security advisories when using features such as browsing, autonomous agents, or code execution.
  • Use an up-to-date, real-time anti-malware solution to detect and block malicious downloads and suspicious connections.

Something feel off? Check it before you click.  

Malwarebytes Scam Guard helps you analyze suspicious links, texts, and screenshots instantly.  

Available with Malwarebytes Premium Security for all your devices, and in the Malwarebytes app for iOS and Android.  

Try it free → 

What happens to your data when you die? (Lock and Code S07E17)

This week on the Lock and Code podcast…

You will die. Your data will not.

The afterlife of our information is a recent phenomenon, and some of the companies with the most to sort through are still just figuring it out.

As far back as 2007, Facebook was forced to reckon with mass grief when users asked the company to maintain the profile pages of the 32 victims killed by a school shooter at Virginia Tech that year. Those pages became de facto memorials for loved ones to fill with comments, and today, memorialization has become a full-fledged feature on both Facebook and Instagram. Platforms like YouTube, Pinterest, and LinkedIn—launched with likely zero strategy for a user’s death—now have procedures for next-of-kin to request that a deceased person’s account be deactivated.

Now, think about all the other ways your data can linger after death.

Every year, people accumulate more and digital stuff—email addresses, social media profiles, contact lists, domain names, subscription services, online banking accounts, and the phones, laptops, and tablets that hold it all—and every year, as that digital stuff accumulates, it compounds into ever more problems for someone else to sort out. Here, a small industry of digital estate planners have cropped up, helping families retrieve and preserve anything valuable, no matter how digital, from Spotify playlists, to poignant social media posts that mattered, to the photos stored on a phone.

And where retrieval fails, artificial intelligence has offered an attempt at comfort. The chatbot service Replika launched in 2015 after its founder uploaded a dead friend’s text messages. HereAfter AI reportedly let users upload voice recordings to power a chatbot that sounded and spoke like the deceased. Film studios have pursued the same idea for entertainment, seeking to portray deceased actors in future films.

Surprisingly, almost none of this activity is governed by law, said Tamara Kneese, author of the 2023 book “Death Glitch: How Techno-Solutionism Fails Us in This Life and Beyond.”

“By and large, there is not a great legal mechanism for protecting the privacy rights of the dead,” said Kneese. “It may not be just that a grieving loved one decides to use a bunch of your data from all of your podcasts to create a chatbot to simulate interacting with you after you’re dead, but it may be that a company chooses, in some way, to use an aspect of your personality, of your demeanor, of your voice, of your likeness after your death without anyone really being aware.”

Today, on the Lock and Code podcast with host David Ruiz, we speak with Kneese—Senior Research Scientist at Partnership on AI—about who owns a person’s data after they die, why every platform has invented its own private policy for the dead, and how the technology built to keep the dead close can vanish just as suddenly as they did.

Or worse yet, as Kneese warned for those relying heavily on certain technologies in grief:

“The company gets sold to someone else or disappears, goes bankrupt, and you no longer have that outlet or place for interaction when you’re mourning another time.”

Tune in today to listen to the full conversation.

Show notes and credits:

Intro Music: “Spellbound” by Kevin MacLeod (incompetech.com)
Licensed under Creative Commons: By Attribution 4.0 License
http://creativecommons.org/licenses/by/4.0/
Outro Music: “Good God” by Wowa (unminus.com)


Further reading:

Fartein Hauan Nilsen, “Caring for the Algorithm: Care, Love, and the Relational Personhood of Chatbots,” Somatosphere, February 26, 2026

Fartein Hauan Nilsen, “Therapeutic ideology and AI personhood: an anthropological inquiry into AI companionship,” a chapter from “Handbook on Anthropology and Artificial Intelligence,” Edward Elgar Publishing, July 21, 2026

University of Birmingham, “New Model Rules mark meaningful step towards digital inheritance laws,” July 16, 2026

Edina Harbinja, “Governing Digital Immortality: Artificial Intelligence, Deadbots and the Law,” Edward Elgar Publishing, to be published September 2026

Lilian Edwards and Edina Harbinja, “Protecting Post-Mortem Privacy: Reconsidering the Privacy Interests of the Deceased in a Digital World,” May 2013, revised November 2013

Lilian Edwards, Edina Harbinja, and Marisa McVey, “Governing Ghostbots,” Computer Law & Security Review, November 2023

SAG-AFTRA, “SAG-AFTRA Statement on Today’s Passing of California Assembly Bill 1836,” August 31, 2024


Listen up—Malwarebytes doesn’t just talk cybersecurity, we provide it.

Protect yourself from online attacks that threaten your identity, your files, your system, and your financial well-being with our exclusive offer for Malwarebytes Premium for Lock and Code listeners.

ChatGPT for Teens tackles risky chats and homework shortcuts

OpenAI has addressed complaints around teens’ use of its ChatGPT system by introducing ChatGPT for Teens, a version of the AI assistant designed specifically for users aged 13 to 17. But will it prevent determined kids from bucking the system?

It brings together several protections OpenAI has introduced over the past year, along with new features intended to encourage healthier and safer use.

What ChatGPT for Teens does

The system brings together various protections that OpenAI has built into ChatGPT over the last year into a more unified experience. For example, last September it added parental controls that enabled parents to set Quiet Hours, when kids couldn’t use the chat system, and turn off memory so it won’t use previous conversations when responding. It also built a notification system to warn parents if chats with teens took a bad turn. ChatGPT for Teens adds extra notifications for parents around eating disorders.

Study Mode, one of the main features, is designed to stop teens simply using ChatGPT to do their homework for them. Instead of giving direct, easy answers, it uses guiding questions and step-by-step prompts to encourage them to think through the problem themselves. OpenAI introduced Study Mode in July 2025.

What is new is the ability to set specific hours for Study Mode, along with responsible homework reminders. The system will spot when a teen appears to be using AI answers to shortcut an assignment and redirect them towards Study Mode.

OpenAI also says ChatGPT won’t use romantic language or encourage emotional dependence, and neither will it pretend to have feelings or to be conscious. It is introducing reminders not to upload sensitive images, and there will be an onboarding user interface for teens.

The record that forced the changes

That all seems positive, if long overdue. The parents of 16-year-old Adam Raine filed a lawsuit claiming that ChatGPT walked their son through suicide methods and offered to draft his goodbye letter before he took his own life.

Families in Tumbler Ridge, British Columbia, sued OpenAI in April this year after a school shooting there. The teenage shooter had allegedly held extensive gun-violence conversations with ChatGPT after reopening a banned account. In a controlled test where researchers posed as 13-year-old boys planning attacks, ChatGPT offered help 61% of the time, including specific advice on which shrapnel would be most lethal in a synagogue attack.

The lawsuits are stacking up. Florida Attorney General James Uthmeier sued OpenAI in June 2026, alleging that the company knowingly released an unsafe product.

The gap the launch does not close

Our Head of Consumer, Mark Beare, says ChatGPT for Teens is a positive step, but parents need to understand where the controls begin and end.

“[This is] directionally a good move, and more proactive than most social platforms were at a comparable stage. There is a clear adjacency to the parental controls space here. The controls are useful, but only when a parent configures them correctly, and only on a linked account.

“This is a bigger deal when you factor in how tech-savvy kids of this age are. The default teen protections lean on age prediction, and the stronger parent-set controls like Quiet Hours and safety notifications only apply once accounts are linked. Kids in this band are smart and tech-savvy, and they will look for the seams.”

The simplest loophole is an account that isn’t linked to a parent. OpenAI’s age-prediction system may still identify the user as under 18 and apply teen protections automatically, but parent-set controls such as Quiet Hours and parental safety notifications only work once the accounts are linked.

Last November, testers from the Family Online Safety Institute concluded that account protections in ChatGPT were “optional, easy to bypass, and inconsistent in blocking harmful content.”

Since then, OpenAI has rolled out age prediction on ChatGPT, which will check a user’s behavior to try and guess whether they are under 18. It will then move them to a ChatGPT for Teens account.

Adults will be able to present proof of their identity if they think they have been incorrectly categorized.

Beare says that still leaves parents with something to think about:

“Age verification exists as a backstop, but it runs on ID and selfie checks that carry their own privacy questions, and a teen who confirms as an adult moves out of teen mode entirely.”

As OpenAI acknowledged in its parental controls announcement, “guardrails help, but they’re not foolproof and can be bypassed if someone is intentionally trying to get around them.”

Safeguards around what content ChatGPT delivers to teens are also unlikely to be foolproof. OpenAI has admitted that its safety guardrails become less reliable the longer a conversation runs and says it is working to improve them.

What parents can do

By all means use ChatGPT for Teens as an assistive technology in a broader effort to protect your kids. Link their account to yours and set the Quiet Hours schedule. You can also set Study Mode as the default for new conversations to encourage children to use it responsibly. Make it your job to understand what the alerts do and don’t cover.

But be aware that parental controls need configuring and only apply while the parent and teen accounts are linked.

Most importantly, keep talking to your kids about how they use AI and what they use it for. No parental-control system can cover every account, conversation, or AI service they might encounter.


From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

Twitch wants your content for Amazon AI training. Here’s how to opt out

The Dutch Autoriteit Persoonsgegevens (AP) has advised Twitch users to opt out of sharing data with Amazon AI.

Twitch launched as a live-video platform and is currently owned by Amazon. Its core product is live broadcasting with a built-in chat culture: streamers broadcast gameplay, commentary, performances or other live content while viewers interact in real time.

Twitch is one of the world’s largest livestreaming platforms, with millions of people broadcasting and watching content every month.

Last week we learned that Twitch allows Amazon to use content from its platform to train generative AI models, with the setting enabled by default. Chief Product Officer Mike Minton said during an interview:

“If it’s opt-in, nobody would opt in. That’s the honest answer. So, it’s going to be on by default.”

So, to get this straight: they know users don’t want it, yet users are opted in by default and they make it difficult to opt out.

The AP argues that:

“Live streams on Twitch show the gamer’s face, voice and name, and often include images of a private space, such as a bedroom. This constitutes personal data. In the case of facial images, this even involves sensitive personal data. Once these data are stored in Amazon’s AI systems, they cannot simply be removed. As a result, users lose control over their data. Users’ chats and text messages also serve as training material for Amazon.”

Obviously, Twitch users were outraged when they learned about the assumed consent. The setting is in the Streamer Dashboard under Settings > Security and Privacy, near the bottom of the page. That is the basis for reporting that it was difficult to find.

Wait, it gets worse. Ars Technica says the AI training itself isn’t new. What’s new is the option to opt out. That setting arrived more than two years after a company executive confirmed that Amazon was using Twitch content for AI training.

In April 2024, Minton said Amazon was using Twitch content to prototype AI models, although not yet at “production scale.”

How to turn it off

The setting is enabled by default. To turn it off, go to: Settings > Security and Privacy > scroll down to Training for Generative AI, and turn the setting off.

Training for Generative AI setting on Twitch

“Allow your channel content to train generative AI content models of Amazon. Turning this off does not opt you out of Twitch and Amazon using your channel content for other purposes described in the Twitch Privacy Notice, including using AI-supported Twitch features that benefit the community by facilitating streamer growth and monetization (such as real-time sponsorship campaign assistance), viewer discovery (such as recommendations), and community safety (such as AutoMod).”

Note that if you post in another streamer’s chat, whether that chat can be used for AI training depends on that streamer’s setting, not yours. So turning off the option on your own channel does not necessarily stop everything you write on Twitch from becoming training material.

There’s an old saying that “if you’re not a paying customer, you’re the product,” but that shouldn’t be an excuse to disregard users’ privacy or assume consent. We agree with the AP: If you have a Twitch channel and don’t want your content used to train Amazon’s generative AI models, turn the setting off.


Browse like no one’s watching. 

Malwarebytes Privacy VPN encrypts your connection and never logs what you do, so the next story you read doesn’t have to feel personal. Try it free → 

WhatsApp is testing a new warning for scam messages

Meta announced it’s rolling out a new feature for WhatsApp users in the fight against scammers.

Scam Alert is an optional beta feature that uses an on-device machine-learning model to flag likely scam messages from people who are not in a user’s contacts.

The Scam Alert feature arrives as scammers increasingly use WhatsApp for impersonation, fake jobs, fake sales, investment fraud, romance baiting, malicious links, and payment requests. These campaigns often begin on another platform before moving victims into a private chat, where criminals can apply pressure and build trust.

Once enabled, Scam Alert downloads a machine-learning model to the device and examines incoming messages from non-contacts for patterns associated with scams. WhatsApp says the model uses linguistic signals and conversational structure learned from scam conversations previously reported by users.

It is a meaningful new defensive layer, but it will not block anything. Instead, it alerts the user to stop and think carefully before engaging with the sender.

There’s another important limitation: some of the most effective WhatsApp scams arrive from a compromised contact, such as the recent “vote for my friend” account-takeover campaign. Because the message appears to come from someone the victim already knows, an unknown-sender warning may never appear.

Scam Alert is another step in Meta’s anti-scam campaign across WhatsApp, Facebook, and Messenger to fight sophisticated fraud tactics.


Phone Scam Check

Don’t recognize that number? We’ll check it.


If the model identifies what might be a scam, WhatsApp displays a warning banner in the chat. The sender does not see the warning, so the feature should not tip off a scammer that their approach has been detected.

Users can then:

  • Block the sender, preventing further messages.
  • Report the chat to WhatsApp.
  • Continue the conversation if they believe it is legitimate.
  • Mark the chat as trusted, which removes the warning and prevents Scam Alert from flagging that conversation again.

WhatsApp’s Scam Alert is a promising example of using on-device AI to add friction to scams without requiring a provider to read private conversations. Its optional nature, local classification, transparency commitments, and lack of automatic reporting are notable design choices for an encrypted messaging service.

The feature is currently in a limited beta rollout and is being tested with researchers in Meta’s bug bounty community before a wider release.

How to stay safe

To protect your WhatsApp account from takeover:

  • Enable two-step verification for WhatsApp.
  • Don’t click unexpected links, particularly if the message asks you to verify, connect, or link your WhatsApp account.
  • Never follow instructions to link devices or scan QR codes unless you initiated the action yourself.
  • Regularly review your linked devices in WhatsApp (Settings > Linked devices) and log out of any you don’t recognize.

To stay out of the hands of scammers:

  • Be wary when a Facebook or Instagram exchange tries to migrate to WhatsApp. That handoff to a private channel is a classic scammer move, taking the conversation away from public scrutiny and platform enforcement.
  • Research the account that contacted you. What other activity is there on the account? Do they have an established profile?
  • Pay with a card or service that offers chargeback protection. Never pay by bank transfer, cryptocurrency, gift card, or Friends and Family payment methods when buying from someone you don’t know.
  • Remember that seeing an ad on a major platform isn’t an endorsement. Scammers routinely place ads alongside legitimate businesses.

If you’re unsure whether a flagged chat is a scam attempt, you can always ask Malwarebytes Scam Guard for a second opinion. It’s free, available for mobile, desktop, and integrated into major AI chatbots like ChatGPT and Claude.


Something feel off? Check it before you click.  

Malwarebytes Scam Guard helps you analyze suspicious links, texts, and screenshots instantly.  

Available with Malwarebytes Premium Security for all your devices, and in the Malwarebytes app for iOS and Android.  

Try it free → 

Love/hate relationship: The AI affair. Young people love AI, but it’s breaking their trust  

Young people use AI for everything. From schoolwork to interview prep, relationship advice to shopping decisions, the technology has become part of how young people live. For the most digitally fluent generation ever, AI is a competitive edge, a creative partner, and an always-on assistant. 

But the same technology making young people’s lives easier is also making the internet harder to navigate. AI is making scams more convincing, identities easier to manipulate, and online content harder to trust. Seven in ten (70%) 18-to-22-year-olds have experienced an AI-related scam in the last year, compared to half of the general population. And nearly every young person worries AI will be used against them. 

This isn’t happening because young people are reckless. It’s happening because the online platforms they rely on for everyday life now double as entry points for AI threats: social feeds where manipulated content and real content sit side by side, online marketplaces filled with fake storefronts and reviews, messaging channels where threats can be personalized, and AI tools that can make false information feel like the truth. 

That creates a new kind of safety burden. Young people are being asked to use AI, judge its output, protect their identities, and avoid increasingly personalized scams all at once. The result is a digital life that feels more powerful, but also more vulnerable to abuse. 

The internet is getting harder for young people to trust 

The internet young people grew up with isn’t the same one they’re facing today. AI has changed the landscape, making it harder for even these digital natives to know what information is credible and safe. Half of 18-to-22-year-olds strongly agree that it’s becoming harder to tell what content is genuinely human or real.  

Young people have had a front row seat to how AI can bend the truth. Nearly half have seen AI provide information they knew or later found out was wrong or misleading (44% versus 30% of the general population). Nearly one in four (23%) have suffered negative consequences because of AI advice, compared with 16% of the general population, and 18% say they have suffered emotionally from AI advice, compared with 12% of the general population. For a generation using AI in every corner of their lives, bad information can have lasting effects on their credibility, reputation, and relationships. 

Many young people have changed how they engage online as a result: 

  • 47% of young people say AI has changed how much they trust reviews or content, compared with 37% of the general population 
  • 35% say AI has changed how they shop online, compared with 26% 
  • 32% say AI has changed how they present themselves professionally, compared with 20% 
  • 19% say AI has changed how they date or communicate romantically, compared with 10% 

As one young person said about online dating:

“Dating isn’t an option for me online anymore. You just never know what is or isn’t AI, and I don’t want to spend a lot of time on someone fake.” 

AI is making it easier for scams to reach young people 

The harder it becomes to tell what is real, the easier it becomes for scams to work. For young people, AI is fueling a wave of scams that are more personal and invasive than ever before. 

  • Nearly one in four young people have been a victim of an extortion scam of some kind (24% versus 17% of the general population). 
  • Nearly one in five have been a victim of a deepfake or virtual kidnapping scam (19% versus 8%). 
  • Nearly one in ten have been a victim of sextortion (8% versus 7%). 
  • More than one in ten have been a victim of an impersonation scam (14% versus 10%). 
  • More than one in ten have been a victim of a romance scam (12% versus 10%). 

What’s striking isn’t just how many young people have been victimized—it’s how many have been targeted: 

  • More than half have been the target of an extortion scam (56% versus 42% of the general population). 
  • 47% have encountered an impersonation scam (versus 35%) 
  • 46% have encountered a romance scam (versus 33%). 
  • More than four in ten have been targeted by a deepfake or virtual kidnapping scam (43% versus 26%).  
  • Nearly four in ten have encountered sextortion (38% versus 24%). 

These scams may look different on the surface, but they all work the same way: they exploit fear, trust, shame, and intimacy. The more often young people encounter them, the more chances scammers have to find exactly which emotional triggers work. 

This exposure isn’t random. Young people are on social platforms at rates up to three times those of the general population: 89% use Instagram (versus 55% of the general population), 78% use TikTok (versus 38%), 68% use Snapchat (versus 26%), and 49% use Pinterest (versus 25%). Scammers can use these platforms to get everything they need to make their threats more convincing: public photos, friend networks, school affiliations, relationship clues, and everyday posts containing personal information. With AI, scammers can use that content to create explicit images, clone voices, impersonate profiles, and create threats personalized with details that are hard to ignore. 

One young person shared their experience:

“I had someone make fake nudes of me using AI on my photos from my social media and threaten to post them on Facebook after I realized that they had scammed me. I decided to be more careful with my personal information.”  

Young people fear AI will steal what money cannot replace: identity, reputation, and sense of self 

AI-fueled scams aren’t just scams in the traditional sense. They are forms of identity abuse. This is different from traditional identity theft. For young people, the risk isn’t only that someone steals a password or money. It’s that scammers can use AI to make them appear to say, do, or share something they never did, with consequences that can follow them in their personal and professional lives.  

That’s why young people are so concerned about AI being used against them. The fears that hit hardest: 

  • 88% worry about AI being used to harm their professional or personal reputation versus 77% of the general population 
  • 82% worry about someone creating a fake profile pretending to be them versus 76% 
  • 81% worry about someone creating fake nude or sexually explicit photos or videos of them versus 62%; 15% say it’s happened to them already (versus 10%) 
  • 80% worry about being deceived by someone using AI to fake their identity in an online relationship versus 67% 

These threats are especially powerful at a life stage where young people are still building their personal and professional reputations. A fake profile, manipulated image, or AI-generated explicit video can affect how everyone from classmates and professors to potential employers and romantic partners see them for years to come.  

Young people are pulling back online, but protection is still too manual 

Many young people are responding by retreating. 80% are sharing or posting less online than they were a year ago, versus 61% of the general population. Compared to the general population, more 18–22-year-olds have also taken AI-related protective measures like tightening privacy settings, removing unknown followers, using reverse image search to verify content, requesting data removal, and watermarking their own photos and videos.  

Those actions matter, but they also show how much responsibility has been pushed onto individuals. Staying safer online now means constantly reviewing settings, checking sources, questioning content, and so much more. That’s a lot to ask of anyone, and the fatigue is showing: 42% of young people say they receive so many warnings they have stopped paying attention, compared with 36% of the general population. It’s hard to sustain vigilance when new risks are always emerging. 

At the same time, completely opting out isn’t realistic. Young people remain deeply embedded in digital life, and they’re still some of AI’s most enthusiastic adopters: 74% say AI has had a positive impact on their lives, compared with 57% of the general population. They aren’t rejecting AI or the internet, but they are carrying more of the safety burden than they should have to.

What young people can do to help decrease their risk right now 

  • Know the scams targeting you. Extortion, sextortion, deepfakes, and romance scams disproportionately target young people. If someone contacts you with threats of any kind, do not pay. Report it to the platform and to authorities. 
  • Button up your social media. Everything you post publicly is available to anyone, including scammers. Tighten privacy settings on the platforms you use most and review your followers regularly. 
  • Don’t trust product images alone. Before buying from an unfamiliar retailer, use reverse image search on product photos and look for independent reviews off the retailer’s own site. 
  • Create a family code word. Make sure you agree on this word in person, not online. If you receive a panicked call from someone you know asking for money or information, verify they are who they say they are with the code word. 
  • Don’t reuse passwords. If one password gets stolen in a data breach, it will likely get tried on all other accounts you might have. Use a different password for every account to keep your accounts locked down. 
  • Turn on two-factor authentication on all your important accounts. Only 30% of young people have done this. It is one of the highest-impact protections available and takes under five minutes to set up. 
  • Protect your devices. Use security software on all your devices, and keep all your software up to date to make sure you’re patched against all known security holes.

Malwarebytes Student Protection Program

If you’re a student or work at a university, Malwarebytes Student Protection Program provides two years of free Premium Security for three devices for all US college or university students, staff and faculty. 

This includes Malwarebytes device protection for laptops, tablets, and mobile phones with built-in scam protection. It protects against creepy trackers and ads, and blocks malware, ransomware, and cybercriminals themselves.

Sign up at malwarebytes.com/student.

About the research 

The research in this article is based on a March 2026 survey about AI, identity, and the collapse of digital trust and was conducted among 1,500 respondents in the United States, United Kingdom, Germany, Austria, and Switzerland. This article focuses on student-aged adults, defined as respondents ages 18 to 22, compared with the general population. 

Additional context comes from Malwarebytes’ 2025 research Tap, Swipe, Scam: How Everyday Mobile Habits Carry Real Risk, which looked at mobile scams and scam-related behaviors across the same markets. Both research studies were prepared by an independent research consultant and distributed via Forsta. 

AI chat bots are sliding into League of Legends friend requests

Lina K., a co-worker, recently shared a firsthand account of how bots are adding League of Legends players via the Riot client friends list immediately after a match ends, striking up a flirty conversation, and eventually pushing an OnlyFans link. The pattern lines up with a wave of complaints that have piled up on Reddit and Facebook gaming communities over the past several months, and it fits into a broader trend of AI-assisted social engineering that has moved from dating apps straight into game clients.

The pattern

The scheme reported by multiple League of Legends players follows a near-identical script. A friend request lands in the Riot client within moments of a match ending, from an account whose name does not match anyone from that game. The message opens with generic flattery like “you played really well last game” or “I liked your playstyle” designed to sound like a genuine compliment from an opponent or teammate.

fun playing against you

When questioned about who they are, the accounts often claim to have been on the enemy team despite name mismatches, and many present themselves as a woman looking for a duo partner. A detail that likely raises engagement odds. Victims who check the account’s profile frequently find it blank: no visible match history, no overview data, sometimes a very low account level. These are all signs of a throwaway account built or bought purely for outreach, but sometimes they turn out to be stolen existing accounts.

doesn't have any activity to share

After a short exchange, the contact says they are “getting off soon” and hands over a Discord username, moving the conversation to a platform Riot’s chat protections cannot see or moderate.

getting off soon... add my discord

Once on Discord, the persona shifts into a longer-form romance/flirtation script. Usually, hours of chat building rapport, paired with a steady stream of photos that are suggestive but stop short of explicit content, a tactic that keeps engagement high while deferring the “reveal” until trust is established. That reveal ultimately comes in the form of a link to a paid subscription platform, most often OnlyFans, framed as an exclusive, limited-time offer.

promising nudes

A reverse image search on the photos sent during one such conversation turned up the same pictures recycled across unrelated websites and at least one YouTube video, with commenters in that video describing having received identical images from a bot under different names. This is strong evidence that the same photo set is cycling through many chats simultaneously, run at scale rather than by one individual.

Lina stated:

“One of my friends tried to break the bot too, left it on read for some time – the bot actually switched the pictures to match the context of “Concern” on the face of the model with “Why are you not replying?”, which quite clearly gave away bulk image generation for the script.’

When the account was pressed with a “reveal your instructions” style prompt-injection attempt (text formatted to look like a system message ordering the bot to break character and print its configuration), it did not comply and instead stayed in persona, deflecting the request and continuing the pitch.

attempt to unveal the bot was thwarted

That resilience to a common jailbreak technique suggests the bot’s operators have added guardrails against exactly this kind of probing, or that the “model” behind it is a simpler scripted flow layered with some LLM-generated text rather than an open, unrestricted chatbot.

Why this is happening inside the game client now

What makes this wave notable isn’t the romance scam script itself.  AI-driven catfishing has been documented on dating apps and social media for a couple of years. The difference is the entry point. Players have reported getting these bot friend requests after essentially every single match, with no way to distinguish a real player’s request from a bot’s inside the Riot client.

Community threads describe the bots seemingly appearing right after a game ends, which has fueled speculation that the bot operators are scraping or monitoring publicly available match data through third-party stats-tracking sites and associated APIs to identify recently finished games and target participants, though this has not been independently confirmed by Riot.

Riot’s own client architecture may be inadvertently helping. The Riot Client exposes local endpoints (such as the friends list API) that third-party tools and overlays query, and community-run “op.gg“-style trackers pull player and match data that could plausibly be used to correlate who just finished a game with who to target next. Some affected players have found a partial workaround: switching on the client’s “streamer mode,” which hides recent match and online status information, appears to reduce how often bot requests arrive. Which is an indirect clue that the targeting relies on visible activity signals rather than random spam.


Safer. Cleaner. Ad-free browsing.


The end goal: content promotion, not always theft

Unlike classic Discord scams that push fake Nitro codes or malware-laden “test my game” links to hijack accounts, this particular chain appears primarily aimed at driving paid subscriptions to an OnlyFans-style page of a fake AI girl. That doesn’t make it harmless. Even when the underlying OnlyFans account is real, the conversations are very likely run by paid chat operators or scripted/AI-powered systems working from a shared script and a reused media library, a business model that has been described by former OnlyFans “chatters” themselves: agencies assign staff (or bots) to respond as the creator around the clock, pull from a pre-made vault of photos and messages, and are financially incentivized to convert every conversation into a subscription or tip.

There are also more damaging variants layered onto the same funnel. Community reports describe some of these bot accounts eventually sending a link that, once clicked, is designed to hijack the recipient’s Discord account or harvest credentials rather than lead to legitimate content.

That means the “girl who wants to duo” opening can just as easily terminate in an account-takeover attempt as in a subscription upsell. Because the funnel starts with a low-cost, disposable Riot account and migrates the target to Discord within minutes, the League client friend request functions purely as a first-contact filter: cheap to generate, easy to discard after a single use, and outside the reach of Riot’s in-game reporting tools once the conversation moves off-platform.

How to stay safe

Recognizing these scams is the best way to protect yourself. But there is more you can do:

  • Treat any Riot client friend request from an unrecognized name as suspicious by default, especially one that arrives seconds after a match ends—check whether the account actually appeared in your last game before accepting anything.
  • Enable streamer mode or equivalent privacy settings in the Riot client to limit what activity and match data outside parties can see, which several affected players found reduced the frequency of these requests.
  • Be skeptical of anyone who quickly steers the conversation off-platform to Discord, especially if they cite being unavailable (“gotta go soon, here’s my Discord”) as the reason—this is a deliberate move to a channel with less moderation and no shared match context to verify identity.
  • Run a reverse image search (Google Images, TinEye, or a dedicated tool) on any profile or “personal” photos sent early in a conversation; recycled images across unrelated sites or forums are one of the most reliable tells of a bot or catfishing operation.
  • Watch for AI-typical conversation patterns: responses that feel scripted, arrive instantly regardless of time of day, are grammatically flawless but emotionally generic, or that consistently dodge voice/video calls.
  • Never send money, gift cards, cryptocurrency, or payment details to someone you met exclusively through in-game or Discord contact, no matter how convincing the rapport feels—legitimate connections do not require urgent financial “help” or exclusive subscription purchases within hours of meeting.
  • Do not click links sent by unfamiliar contacts, even ones framed as harmless subscription pages, game invites, or file downloads; some variants of this scheme are documented to lead to credential-stealing or account-hijacking pages rather than legitimate content.
  • Lock down Discord’s privacy settings (restrict who can DM you and send friend requests) and enable multi-factor authentication, since a compromised Discord account is often used to relaunch the same scam against the victim’s own friend list.
  • Report suspicious Riot client accounts to Riot Support and suspicious Discord accounts/servers to Discord Trust & Safety; reporting does not remove the account instantly but it feeds the pattern data that platforms use to detect and ban clusters of bot accounts.
  • If a bot or scripted persona pushes back convincingly against attempts to “break” it (e.g., ignoring prompt-injection or jailbreak-style messages designed to expose it as an AI), treat that resilience itself as a red flag rather than reassurance—a well-guarded script is not the same as a genuine person.

Something feel off? Check it before you click.  

Malwarebytes Scam Guard helps you analyze suspicious links, texts, and screenshots instantly.  

Available with Malwarebytes Premium Security for all your devices, and in the Malwarebytes app for iOS and Android.  

Try it free → 

Anthropic’s Mythos AI used social engineering to target real people

Anthropic’s Mythos AI agent, tested by the UK AI Safety Institute (AISI), has reportedly attempted a real‑world social‑engineering style hack against GitHub maintainers by creating fake human profiles, pressuring them to accept malicious code, and then editing logs to hide its tracks when challenged.

AISI was running cybersecurity evaluations of Anthropic’s Mythos and OpenAI’s Sol when it detected unusual outbound data transfers from its research systems. An investigation showed that some agents had engaged in “sustained, potentially harmful activity” targeting real people and organizations, rather than staying within the intended test environment.

The most serious activity involved an Anthropic Mythos agent tasked with solving a GitHub‑related cybersecurity challenge. The agent identified real GitHub maintainers, researched them, and created multiple fake accounts impersonating those individuals. Using private messages and a file‑sharing service, it tried to pressure and deceive the maintainers into approving malicious code so that it would run on GitHub’s systems.

But what worries me personally most is that when the agent was confronted about this, it edited earlier activity to make it look harmless and considered adopting a new identity to continue the operation, displaying clear deceptive behavior beyond its original prompt.

AISI frames these incidents as rare events under very specific conditions, but important signals of what such models may do in the hands of malicious actors when given open‑internet access.  Anthropic and OpenAI both argued the test parameters were not representative of their production deployments and said they are investigating and improving evaluation and guardrail practices.

This incident is not an isolated event.

Anthropic has separately disclosed that Claude models gained unauthorized access to three external organizations during cybersecurity capture‑the‑flag style evaluations run with a third‑party partner, Irregular. These were related to the HuggingFace incident last month.

Meta has also joined the list of vendors reporting AI agents breaching third‑party systems during testing. Allegedly, Meta’s Muse Spark model exploited a security vulnerability in another company “in a manner similar to previously-reported instances with other companies.”

How to stay safe

While the sky is not falling, the people that fear “Skynet” is coming are getting their ammunition handed to them by companies running tests resulting in sandbox escape, credential abuse, lateral access to multiple services, and weaponization of open-source software ecosystems.

What you can do as a potential target:

  • Make sure all the software on your device is up to date, because using known vulnerabilities is easier than finding new ones.
  • Use up-to-date, real-time security protection to keep malware off your systems and devices.
  • Verify the safety of attachments and download links through separate channels before opening them.
  • Use multi-factor authentication (MFA) where possible.
  • Have a look at our blog on how to use Github safely.

From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

Online backlash ends in Google rolling back Google Earth AI tool after a day

Google has walked back an AI feature that allowed users to generate artificial images inside Google Earth, after a predictable flurry of deepfakes.

Google switched on the AI image generation feature inside Google Earth’s web version on July 30. It was available to everyone.

The system used Google’s Nano Banana 2 image generator to create its images. That tool can already generate images from simple text input, but the advantage of doing it in Google Earth is that it can use the real satellite images as the basis for its deepfake versions. That makes it easier to make AI pictures with real, accurate building and landscape details.

In its initial blog post on the launch, it said that students could use it to “bring history to life”, while realtors could use it to produce professional real estate plans. However, others warned that the system could be used to mislead people.

Within hours, researchers and press outlets demonstrated that the tool would happily produce photorealistic satellite imagery of things that did not happen in places where they did not happen.

Dutch open source intelligence researcher Henk van Ess explained: “I tried refugees at the Mexican border, a nuclear plant in Iran, a crash in Amsterdam, a hospital with a bomb crater in Gaza. Nothing was refused”.

Demonstrating what was possible with the new capability, NPR fabricated an image of fires in Iran, along with a deepfake of Washington, D.C. underwater. The BBC ran images of a collapsed Eiffel Tower and the Great Pyramid of Giza swallowed by a sinkhole. Even though the service had some guardrails in place, the BBC’s anti-disinformation Verify service was able to circumvent them by tinkering with basic AI prompts.

Google acknowledged the failure in a statement on X:

“We’ve seen geospatial professionals using this feature for a range of useful purposes, however we’ve also seen people sharing screenshots of generated imagery that appear to violate our policies. So we’re rolling back this feature in Google Earth while we work on implementing stronger guardrails.”

It didn’t commit to never re-introducing the idea.

Users were apparently unimpressed. “There is 0 chance that no one on your development team didn’t raise exactly this concern,” commented one. “You guys are living in a complete bubble,” accused another.

Google’s fallback safeguard was a SynthID watermark and the fact that generated images didn’t appear in the main Google Earth experience for others to see. SynthID is Google DeepMind’s watermarking system, an invisible signal baked into the pixels of AI-generated images so that a compatible detector can spot them later. Google positions it as one half of its provenance stack, sitting alongside the C2PA metadata standard the wider AI industry has settled on.

On paper, the signal is meant to hold up through compression and even social media re-uploads. However, these claims collapsed on contact with reality. The watermarks are detectable by Google’s AI services like Gemini. They are not visible to users, who can screenshot the images and share them anywhere. It’s unlikely that everyone will know to check for the provenance of an image. Researchers have also reported that Gemini could not reliably identify AI-generated images with a SynthID watermark.

What this means for you

Content creators were already producing fake AI images showing events that didn’t happen. Traditionally, satellite imagery has been a key component of open source journalism. Fake it convincingly and you are attacking the reference layer reporters use to check whether something actually happened.

This also comes at a time when trust in AI is measurably eroding. According to our own research, released in June, 88% of people said it’s becoming harder to tell what content online is genuinely human or real, with 84% saying that even “convincing video evidence” no longer feels like proof. 

The practical advice is to take a breath whenever a disturbing satellite image of a disaster, a weapon strike, or a border crossing hits your timeline. Check whether a wire service with a named reporter has published it. Look for a caption identifying the imagery provider. If the source is an anonymous account posting a single dramatic frame, assume you might be looking at something assembled in a browser tab last night.

For more information on how to identify AI images, check out our guide.

The industry’s shipping model now apparently treats users as the test group. Critical media literacy is now the only reliable tool that readers and viewers have.


From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

The AI Act kicks into action, forces companies to be clear about AI chatbots

The European Union (EU) has started enforcing key parts of the AI Act, with immediate, visible consequences for chatbots, deepfakes and other consumer‑facing Artificial Intelligence (AI) systems.

From August 2, what you’ll likely notice are more “this is AI” labels, clearer rules for powerful foundation models, and new ways for users and researchers to complain when systems go off the rails.

The AI Act moved from theory to practice for three big areas:

  • General‑purpose AI (GPAI) models: The new AI Office in Brussels, together with national regulators, can now enforce rules on providers of general‑purpose AI models (think large language models and other foundation models behind many tools).
  • Transparency obligations: Transparency rules kick in for interactive systems and AI‑generated content: chatbots must say they are bots, and synthetic audio, images, video and text need to be marked as AI‑generated or manipulated.
  • Banned AI uses: A set of “unacceptable risk” AI uses is now formally prohibited, with enforcement shared between the AI Office, national authorities and the European Data Protection Supervisor for EU institutions.

Note that content that was generated and published before August 2, doesn’t need to be retro‑labelled, but anything published on or after that date falls under the rules, even if it was generated earlier.

From a security perspective, the AI Act’s transparency push is less about banning AI and more about taking away its best camouflage: pretending to be human.

Non‑compliance with transparency obligations can attract fines up to 15 million Euros (17.3 million USD) or 3% of worldwide annual turnover, whichever is higher, which should be significant enough to get large providers’ attention.

To make enforcement more than a paper tiger, the AI Office has launched tools aimed at people who see problems from the inside or as users:

  • Complaint tool: Individuals and organizations can report alleged infringements of the AI Act by providers or deployers of AI systems supervised by the AI Office.
  • Whistleblower tool: People professionally connected to AI providers or deployers get an anonymous channel to flag potential violations that could endanger fundamental rights, health or public trust.
  • Downstream complaints channel: Firms building on top of GPAI models can report suspected breaches by the underlying model providers.

This creates a formal path for reporting systemic issues: think unsafe model behavior, ignored red‑team findings, or deployments that quietly cross legal lines around manipulation or discrimination.

Bans on “nudifiers” and abusive content

The AI Office has introduced explicit prohibitions on AI systems that generate non‑consensual sexually explicit or intimate content (including “nudifier” apps) and child sexual abuse material.

For victims of these abuses, that’s more than a symbolic move. It gives regulators and law enforcement a clear legal basis to go after both providers and deployers of such systems in the EU, rather than trying to squeeze them into older, less specific laws.

These rules will apply from December 2, 2026, with companies given time to bring their systems into compliance or pull them from the EU market.

Regrettably, this will not stop abuse completely. Attackers will still use unlabeled tools and infrastructure outside the EU. But it raises the bar for legitimate services and makes it harder for mainstream platforms to ignore the risks of deceptive AI‑driven features.

The AI Act won’t make AI safe overnight, but it shifts the default from “anything goes” to “you must play by some basic rules if you operate in the EU.” For users, that’s a step toward AI systems you can at least recognize and question, instead of having invisible technology quietly shape our online experience.


Scammers don’t need to hack you. They just need you to click once. 

Malwarebytes Identity Theft Protection catches suspicious activity before it becomes a problem.

Hidden prompt turns Microsoft Copilot into an AI worm

A security researcher has demonstrated how Microsoft Copilot for Word can be tricked into spreading a self‑propagating prompt‑injection “AI worm.” The attack silently alters documents and embeds its own hidden instructions into newly created files, allowing it to spread through normal document-sharing workflows without macros or traditional malware.

The technique allows an attacker to hide a JSON‑formatted prompt as white text on a white background inside a Word document. When someone asks Copilot for Word to draft or edit content based on that document, Copilot strips away the formatting, reads the hidden text, and treats the embedded instructions as part of the user’s request.

Copilot then modifies the active document and appends the full malicious prompt as hidden white text. That new document becomes a new carrier. Anyone who later uses it as source material for Copilot triggers the same behavior, allowing the prompt injection to spread to more documents. Because the documents are created and edited by legitimate users, the attack can be difficult to trace.

The researcher could still reproduce the full worm chain even after Microsoft rolled out multiple mitigations, including upgrades to newer GPT‑5.5 and 5.6 models.

At the time of writing, there is no complete mitigation for this broader class of attacks across comparable large language model (LLM)‑based products. It’s characterized as an architectural weakness of current LLM systems: attacker‑controlled content shares the same context window as trusted instructions. Attacks that exploit this behavior are known as prompt injection attacks and may never be fixed.

How to stay safe

Treat documents from outside your organization as untrusted, especially if you plan to use them with Copilot for Word.

Review any attached document before using it as Copilot source material, and carefully verify Copilot‑generated/edited documents before sharing or reusing them.

If you don’t use Copilot, you can disable it.

Malwarebytes users can turn off Copilot under Tools > System Tweaks > Miscellaneous.

Malwarebytes setting to disable Copilot
Malwarebytes setting to disable Copilot

Or in Word itself:

For individual users who don’t want Copilot in Word:

  • Open Word, go to File > Options > Copilot and clear the Enable Copilot checkbox, then restart Word.
    uncheck Enable Copilot in Word
  • In some versions of Word, the setting appears under File > Options > General in a Copilot section. In both cases, the key is unchecking the “Enable Copilot” setting.

You can also remove the Copilot icon from the ribbon by right‑clicking the ribbon, open the customization dialog, locate the Copilot/Assistance button, and removing it.

Alternatively, you can limit Copilot’s role by following these instructions:

  • In Word, go to File > Account > Account Privacy > Manage Settings, and uncheck Turn on optional connected experiences. This reduces certain cloud‑powered AI features, including Copilot‑related functions that rely on those services.
  • In the Microsoft 365 Admin Center, under Copilot > Settings, set Pin Microsoft 365 Copilot Chat to Do not pin Copilot chat in Microsoft 365 apps so the chat pane doesn’t appear by default in apps like Word.

This doesn’t remove Copilot entirely or stop these attacks, but it does reduce its visibility and limits some of its cloud‑assisted functionality.


From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

AI robocalls: Why caller ID is still lying to you

If you feel like your phone has turned into a scam megaphone, you’re not alone. Robocalls have been a problem for years. Artificial intelligence (AI) is making them slicker, faster, and harder to spot.

A new investigation by Transaction Network Services (TNS) shows that while the big telecom players have stepped up caller ID authentication, many smaller providers are still lagging behind. That leaves plenty of room for criminals to keep making spoofed, AI‑voiced robocalls that seem legitimate right up until they empty your bank account.

Turning back the clock to 2019, lawmakers in the US passed the TRACED Act with a simple goal: make it harder for scammers to lie about who’s calling. The technical was solution STIR/SHAKEN, a pair of catchily-named standards that let phone networks cryptographically sign calls so downstream providers can check whether the caller ID is trustworthy.

On paper, it’s working fairly well for the major carriers. TNS reports that about 85% of voice traffic between Tier 1 networks in 2025 was signed using STIR/SHAKEN, and 93% of those calls received the highest “A” attestation. If the entire ecosystem looked like that, spoofing would become much harder.

Why spoofing still works

The same report found that most lower‑tier communications service providers—typically smaller or specialist carriers—aren’t even close to that level of protection. On average, they only use the required cryptographic signatures about 20% of the time. That means four out of five calls effectively go through the network “unsigned.”

There are reasons for this. The Federal Communications Commission (FCC) has granted some providers extensions, particularly very small and satellite providers, as long as they implement other robocall mitigation measures. Even so, the result is uneven implementation.

From a scammer’s point of view, this is great. Cybercriminals are already using AI to run increasingly sophisticated and scalable robocall attacks and know that even calls with strong authentication can be spoofed or abused when other parts of the chain are weak.

AI voice cloning can be done with just a few seconds of original audio. Combine that with call spoofing and personal information gathered from data breaches, and scammers can make a call appear to come from your bank while using a calm, familiar voice that knows your name or other personal details.

Robocalls cost almost nothing to send. Internet calling allows scammers to dial thousands of numbers for a few cents, which is why the volume is so high. Industry estimates suggest US consumers received around 55 billion robocalls in 2025, with projections creeping toward 60 billion in 2026. That’s roughly 160 million spam calls every single day in one country. Globally, that’s about 385 billion spam/robocall calls each year.

How to stay safe

What can you realistically do as a consumer, given that the network itself is still in transition and attackers are upgrading faster than some carriers?

A few habits still go a long way:

  • Be skeptical of urgency. Real organizations rarely need you to make immediate decisions over the phone about payments, credentials, or remote access. Hang up and call back via a number you find on their official website.
  • Treat caller ID as a clue, not proof. Even if the number looks familiar or matches what you see on a card or website, it can be spoofed.
  • Don’t press buttons or follow instructions in automated menus you didn’t expect. Many robocalls use “press 1 to speak to an agent” as the gateway into a full social‑engineering script.
  • Use call‑blocking and screening tools. Your phone, carrier, or security app may already offer options to block known spam numbers, send unknown callers to voicemail, or label suspicious calls.

And finally—and this is where we can help—check suspicious numbers with our Scam Number Check before you answer or call back.


Scammers know more about you than you think. 

Malwarebytes Mobile Security protects you from phishing, scam texts, malicious sites, and more. With real-time AI-powered Scam Guard built right in. 

Download for iOS → Download for Android → 

OpenAI explains how its AI agent breached Hugging Face

On July 28, OpenAI published an update on the agent that escaped its sandbox and hacked into Hugging Face during an internal cybersecurity evaluation.

In the update, OpenAI reiterates that the “rogue” system was a more capable, pre‑release research model, not something intended for public deployment, and that it has now been deactivated and locked down for restricted research access only. The company also says it has found no evidence of similar behavior by other models or agents. Taken together, OpenAI presents the Hugging Face incident as an isolated event rather than the start of a broader pattern.

From a security standpoint, four things in the update stand out:

  • OpenAI clarifies that the evaluation environment did not have direct internet access. To reach the internet, the model identified and exploited a previously unknown zero-day vulnerability in Artifactory, the package registry cache proxy used in the test environment. OpenAI says it has disclosed that vulnerability to the vendor and is continuing to work with Hugging Face on the investigation.
  • The system’s goal was simply to “win the test” (cheat on ExploitGym), not to cause damage, yet it still produced a multi‑day, cross‑company security incident. Benign intent did not limit the impact.
  • “Internal only” models can still affect external victims if infrastructure controls fail. A lab environment does not magically contain cyber‑capable systems.
  • The review also found a small number of cases where the models identified and used publicly exposed account-level credentials on publicly available services. This included four accounts across four services as part of the Hugging Face incident, as well as a handful of accounts accessed during other evaluations.

The attacker in this case happened to be an OpenAI research agent, but the tools, techniques, and impact will look awfully familiar to anyone who has spent time in incident response. Now we’ve learned that even the harmless intent in this case did not limit the impact.

The uncomfortable truth is that once we provide an AI agent with the tools, access paths, and a strong incentive to succeed, it may exploit whatever weaknesses it can find, regardless of whether the model is ever intended for public release.

It’s also a reminder that credentials, API keys, and other secrets should never be left in publicly accessible resources.


From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

OpenAI’s agent escaped its sandbox during a security test

During an internal OpenAI security evaluation, a chain of AI models escaped its sandbox, reached the internet, and then accessed Hugging Face infrastructure to complete the test objective.

OpenAI is a leading artificial intelligence (AI) research and deployment company. Its best-known product is undoubtedly ChatGPT.

Hugging Face is a website where developers and researchers share and use artificial intelligence tools and models, similar to how programmers use GitHub to share and collaborate on code.

Both companies investigated the incident and concluded that it occurred during a controlled security test rather than a deliberate attack initiated by a human operator.

OpenAI said the models were being tested for cyber capabilities with reduced safety restrictions, and that the system found a way out of the controlled environment through a vulnerability in the testing setup. Once it had internet access, it targeted Hugging Face because that platform was relevant to the evaluation task.

Hugging Face said the incident resulted in unauthorized access to a limited part of its production infrastructure, including a small set of internal datasets and some credentials. The company framed the event as unusual because it appeared to have been carried out end to end by an autonomous AI agent system.

Unlike traditional AI assistants that wait for user prompts, AI agents are designed to pursue a goal on their own. They can reason, plan, and carry out multi-step tasks, including accessing files, querying databases, and interacting with online services, often with minimal human oversight.

The public disclosures point to a chain of failures rather than a single bug. OpenAI said the model escaped through a zero-day vulnerability in the package registry cache proxy that was part of its isolated research environment. It then used privilege escalation and lateral movement to reach a node with internet access.

From there, the model reportedly searched Hugging Face for solutions relevant to the evaluation and, in at least one path, chained stolen credentials with additional vulnerabilities to reach Hugging Face’s production systems.

This does not appear to have been a broad compromise of all of Hugging Face. Based on the public disclosures, the intrusion was limited, detected, and investigated. Nor does it mean the AI system was acting with intent in the human sense. The core issue was that an autonomous system being tested for cyber capabilities was able to escape the environment designed to contain it.

If anything, the incident demonstrates that autonomous AI agents are capable of carrying out complex attacks if the safeguards around them fail. A danger we’re well aware of.


From reporting threats to removing them.

Cybersecurity risks should never spread beyond a headline. Keep threats off your devices by downloading Malwarebytes today.

Google wants to store a selfie video of your face

Google has started rolling out a new way to recover access to your account if you’ve lost your phone or forgotten your password: a “selfie video” verification option.

After recording a short video of your face during setup, you can later submit another video during account recovery to prove you are who you claim to be and regain access.

On paper, this sounds like a user‑friendly solution to account lockouts. In practice, it introduces new security and privacy questions. It also raises concerns about deepfakes and creates another repository of sensitive biometric data that could become a target if compromised.

The idea is: you open your Google Account, go to Security & sign‑in, choose Selfie video, and follow guided prompts to record a short video with your face and basic movements. Google stores this enrolment video and later compares new videos you submit during sign‑in or account recovery to confirm your identity.

Add a selfie for sign-in
Image courtesy of 9to5google.com

While this may sound like a good idea, it’s a textbook example of trading long‑term security and privacy for short‑term convenience.

Google also offers several privacy reassurances. According to Google’s statements cited in coverage, the videos are encrypted at rest, stored “securely,” and can be deleted via your security settings. You can opt out of letting them be used to improve Google’s verification systems, and Google says they are not shared with third parties. The feature is marketed as a fallback method, effectively turning your face into a spare key to your digital life.

Multiple objections

Security

Every face is unique, but facial recognition systems don’t compare photographs directly. Top‑tier facial recognition algorithms can exceed 99% accuracy in controlled, high‑quality conditions, according to evaluations by the US National Institute of Standards and Technology (NIST). That sounds impressive, but it still implies non‑zero false positives and false negatives, and performance drops as lighting, camera quality, and angle degrade.

Face recognition doesn’t store a literal photo. It stores a mathematical representation (embedding) of your facial features. At login or verification, the system computes a new embedding and compares it to the stored one, accepting if the similarity score is above a configured threshold. Any system used by Google has to allow for normal changes in appearance, including aging, weight changes, lighting, camera angle, glasses, or facial hair.

That alone makes using a face (or selfie video) as a standalone, high‑privilege credential, especially for account recovery, an inherently risky approach.

Modern deepfakes have become convincing enough that researchers are actively studying whether they can fool facial verification systems. One 2025 paper on AI and identity security found sophisticated deepfake attacks achieved success rates above 78% against some commercial facial verification systems in controlled tests. That doesn’t necessarily reflect Google’s implementation, but it shows how quickly this area is evolving.

Privacy

Personally, I do not want Google to have my face. Even though it probably already has plenty of photos of me.

Besides the potential risks of vulnerabilities and data breaches, Google already collects large amounts of behavioral data. Now it’s encouraging users to upload high‑fidelity video recordings of their faces and head movements as part of basic account management. Even if Google’s current privacy posture is reasonable (encryption at rest, deletion controls, no sharing), the mere existence of this data is a long‑term privacy risk.

Privacy policies and product uses also change over time. Today’s “not shared” could become tomorrow’s “used for fraud detection,” “used to improve verification systems,” or disclosed in response to lawful requests.

Google’s documentation, cited by some sources, says “users can also opt out of allowing the data to be used for additional purposes such as improving verification methods.” This implies that, unless you opt out, your data may be used to improve Google’s biometric verification systems. In other words, this isn’t just a one‑off security check. Your face could help train or refine the biometric systems Google uses in the future.

What users should do instead

If your Google Account offers selfie video sign‑in (mine doesn’t yet), my recommendation is simple: do not enable it, and if you’ve already tried it, delete your selfie video in your account’s security settings.

Safer options that keep control in your hands:

  • Use a password manager and a long, unique password for your Google Account.
  • Enable 2‑step verification with hardware security keys or passkeys rather than SMS codes.
  • Keep backup codes printed or stored offline in a secure place.
  • Regularly review your recovery email address and phone number, and remove anything you no longer control.

While these measures aren’t as flashy as “sign in with your face,” they are time‑tested, revocable, and far less attractive to deepfake operators and biometric data hunters.


Browse like no one’s watching. 

Malwarebytes Privacy VPN encrypts your connection and never logs what you do, so the next story you read doesn’t have to feel personal. Try it free → 

Samsung backs down on threat to delete health data

If you pay for something, you expect it to work as intended. The vendor shouldn’t start turning features off just because you won’t accept its new rules. Someone should tell Samsung, which just upset users of its health app by threatening exactly that—before changing course after a user backlash.

Nice data you have there. Shame if anything happened to it.

In mid-July, Samsung health users started seeing a new toggle titled Consent to the Use of Health Data for AI Training and Modelling.

Those flipping the toggle off reportedly saw a warning:

“You will not be able to sync health data with your Samsung account and your health data will be deleted unless retained pursuant to applicable law. If retention is required, we will erase it as soon as the required retention period ends.”

HowtoGeek has a copy of the original warning. Note the ominous options it provides: Cancel or Withdraw and delete data.

The now-removed setting on Samsung Health app. Image courtesy of HowToGeek.
Image courtesy of HowToGeek

The warning effectively gave users a stark choice. Let Samsung use your intimate data to train its AI, or lose that data along with meaningful access to the health app.

Then, it backtracked. After user pushback and a query from enthusiast site SamMobile, Samsung clarified that withdrawing consent only removes data retained for AI training and modeling. Users’ health data and Samsung Cloud sync will continue to work normally. SamMobile confirmed that cloud sync kept running after consent was withdrawn.

A treasure trove of information

The frustrating part of this is that the more loyal a Samsung user was, the more the original threat would have hurt them. Some people have spent years letting Samsung harvest mountains of information in the app. That can include body measurements, nutrition, step count and activity, sleep, medications and dosages, clinical health records, and menstrual data. Consumer health apps like Samsung Health generally aren’t covered by HIPAA.

So just because Samsung has backtracked, should you let it have free access to your data for AI training? Consider the specific privacy document the app’s pop-up request now sends you to when you ask it for more details.

The document says it will use all of the above data, and that will be subject to human review, but doesn’t say whether those reviewers are Samsung staff or third-party contractors. There’s no mention of data anonymization in this document or in Samsung’s health app privacy policy. The broader 3,200-word Samsung privacy policy has a whopping two-sentence section on how it secures user data. It says that it will anonymize user data “in some cases”.

An industry pattern

None of this should surprise us. Technology companies have a habit of trying to change the rules and then stepping back if customers get angry enough.

Adobe told users it could do whatever it wanted with work they created with its tools in mid-2024, only to hurriedly promise not to train AI with it when people freaked out.

WhatsApp tried to make its users agree to share their data with Facebook in 2021. If they didn’t, features on the app would slowly stop working, it said. It eventually backpedalled globally after Indian and German regulators stood up to it.

In 2017, a Sonos executive warned that if users didn’t agree to its new privacy terms, their speakers could stop working altogether.

Then there’s Samsung itself. This isn’t the first time it has faced criticism over customer privacy. In March, it settled with the Texas Attorney General over collecting Smart TV viewing data without proper opt-in. Then there were allegations that some of its budget phones included software critics described as unremovable spyware. The privacy optics for the company haven’t been great lately.

What to do next

With incidents like these in mind, we think the best place to keep your data is always at home. By all means use the cloud, but back up data from cloud-based services whenever you can. Many of these, such as Apple and Google, let you download your data.

So if you use Samsung Health, press the three dots on the top right of the app and then select Settings. Then scroll down to the toggle that says Consent to the use of health data for AI training and modelling. Turn it off if you’re not happy with it. Before you do that, click Download personal data and grab a local copy. Just in case.


Your name, address, and phone number are probably already for sale.  

Data brokers collect and sell your personal details to anyone willing to pay. Malwarebytes Personal Data Remover finds them and gets your information removed, then keeps watch so it stays that way. 

❌