Visualização de leitura

When AI’s human in the loop really isn’t

Concerns about the risks of AI systems are certain to be met with four words: human in the loop. The discussion may broaden, but the assurance is inevitable. It’s an AI governance phrase that’s become so rote you hear it in every direction and likely have said it yourself.

But IT leaders should be wary of vendor or team claims that they’ve built human-in-the-loop systems into AI tools because some of these supposed guardrails are no more than rubber stamps.

Some so-called human-in-the-loop systems don’t give employees overseeing the AI tools either the control or the time necessary to fix any problems, some IT experts point out.

For human-in-the-loop systems to actually work, employees overseeing AI tools need to have the domain knowledge and context to take the action the AI tool is addressing when the AI isn’t involved, and they need to have the authority to override the AI decision, says Doug Shepherd, head of offensive security at internet services provider Cloudflare.

Promises of human-in-the-loop systems give IT leaders comfort, but the underlying process often doesn’t work as advertised, he adds.

“If your human in the loop can flag something but can’t actually stop it, that’s not human in the loop, that’s a human adjacent to the loop,” Shepherd says. “That’s performative governance.”

Shepherd, speaking at the recent CIO 100 Awards and Conference in Frisco, Texas, encouraged attendees to embrace AI and focus on projects that drive adoption and impact. Organizations that fail to push AI initiatives will be left behind, he suggested, but he also warned that blind adoption, without focusing on meaningful outcomes and guardrails, can lead to huge setbacks.

Many organizations reach for human in the loop as an important control, but no one stress tests it, he adds. “It gets projects approved, and too often, it does the political work, but not the risk work,” he says.

Darren Kimura, CEO and president at AI integration platform vendor AISquared, agrees that many organizations are deceiving themselves with so-called human-in-the-loop systems.

“Most companies that say they have a human in the loop actually have a human watching the loop,” he says. “The person can see the decision and flag a concern, but they cannot stop it, change it, reject it, or escalate it.”

IT leaders should ask themselves a handful of questions: Can reviewers halt the actions before they take effect? Can they change the output? Are their overrides recorded and enforced downstream? “If the answer to any of those is no, the human is just monitoring AI,” Kimura says.

Too many decisions

Another problem with human-in-the-loop systems is the decision fatigue that can set in when employees are asked to review too many AI decisions and end up button mashing instead of thinking about the consequences.

The AI reviewer needs the expertise and context to evaluate the recommendation, enough time to do so, and both the authority and technical ability to reject or reverse it, says Eric Billingsley, COO and CTO of AI assurance company TrustScale.

But even a qualified and empowered reviewer may gradually stop exercising independent judgment when the AI is consistently right, he notes.

“If the system is right 95% of the time, the person’s job becomes waiting for the rare case when it is wrong,” he says. “Humans are not particularly good at sustained vigilance of a highly reliable automated system. Eventually, review becomes confirmation.”

A good AI system can create bad human controls, he adds. “When the exceptional case arrives, the reviewer may approve it because the system has trained them, through hundreds of correct recommendations, to trust it,” he says.

Billingsley advises IT leaders to evaluate human-in-the-loop systems the same way they monitor other security controls. A control must be monitored, tested, and produce evidence that it is operating as intended, he says.

“A log showing that someone clicked ‘approve’ is not enough,” Billingsley adds. “You need evidence that the person had the necessary context, applied independent judgment, and had the authority to override the AI.”

Robert Blumofe, EVP and CTO at cloud computing and security vendor Akamai, sees the same problems Billingsley does. Some type of human oversight is preferable to fully autonomous AI, he says, but human in the loop can turn into a mind-numbing exercise.

“LLMs produce the correct output just often enough to lull us into a complacent belief that they are more reliable than they really are,” he notes. “After diligently checking the AI output each time and finding no errors, diligence wanes, and human in the loop turns into rote approval.”

IT leaders should take the time to figure out what they’re getting into when vendors or their internal teams pitch a human-in-the-loop system, Blumofe says.

“It’s incredibly important to understand exactly how the system is designed and when and how the human will interact with the AI,” he adds.

Organizations should also explore ways to deploy other technologies as guardrails for AI, instead of turning to unreliable human oversight, Blumofe suggests.

“You need non-AI systems in the guardrail role,” he explains. “These technology tools would help to automate testing and validation of AI outputs, flag issues, and have the capability to pause the AI work. This keeps humans out of approval loops, while also helping to reduce risk.”

When humans aren’t the right choice

Other IT leaders suggest that human-in-the-loop systems aren’t the right solution in every AI use case. When AI is used to flag and mitigate cybersecurity incidents, for example, waiting for a human to approve an action may be too late.

“If an endpoint is compromised, you may want the system to isolate it immediately,” says AISquared’s Kimura. “Waiting 20 or 30 minutes for someone to approve that action could allow the attack to spread.”

The objective is not to put a human into every AI decision, he adds. “It is to put the right human, with the right context and authority, at the right point in the workflow.”

AI agents need to learn when enough is enough

For the past few years, enterprise AI programs have focused on making models more useful, accurate, and autonomous. In that phase, a bad answer was still usually something a human could accept or reject before taking action. But once agents start invoking tools and acting inside business workflows, success should no longer be measured only by how much work they complete. A more important metric is how well an agent recognizes when it lacks the authority, context, or judgment to continue.

When helpful becomes risky

According to Allan Dabre, technology compliance and AI lead at PwC, a behavior that has to be deliberately designed into the system is, “I don’t know.” AI is built to be helpful, so an agent will generally try to do something useful unless it’s been configured not to.

“The fact that AI systems can hallucinate illustrates that tendency,” Dabre says. “When they lack enough information, they may still produce an answer. In an agentic workflow, that impulse can become more dangerous because the output may become an action, rather than remain a suggestion.”

He adds that many enterprises still test AI primarily for completeness and accuracy. That made sense when the central question was if the model could produce a reliable response. But as models improve and agents gain more operational authority, he argues that CIOs need to prioritize something else: restraint.

“Can it stop at the exact moment you want it to stop?” he asks. “Are you testing for that?”

Confidence is not authority

Dabre makes a simple but important distinction. An AI agent may be 99% confident a record should be updated, a refund should be approved, or a legacy database can be decommissioned. But that doesn’t mean the agent has the authority to act. Confidence is about the probability the system believes it’s right. Authority is about whether the organization has delegated that action to the system in the first place.

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Allan Dabre, technology compliance and AI lead, PwC

PwC

He gives the example of an agent asked to analyze legacy software and recommend what can be decommissioned. The agent may conclude, with high confidence, that several databases have little user impact and can be deleted. But even if the system is confident, most organizations wouldn’t want it to delete those databases on its own.

The same logic applies across business processes. An agent may be confident a customer record should be updated, an opportunity in a CRM system should be closed, or a transaction appears legitimate. But once that action flows into other systems, the potential consequences expand.

That’s why Dabre argues for what he calls an agent harness: a controls or orchestration layer outside the model that defines what the agent can and can’t do. In a refund workflow, for example, a company might let the agent approve small refunds, require human approval for larger ones, and stop the process entirely above a defined threshold. The agent may gather the relevant context, explain the request, and prepare the case for review, but the decision is governed by the authority boundary encoded into the system.

“It’s not a policy document and it’s not a prompt,” Dabre says. “It’s software or a configuration you can apply to an agent.”

The case for least agency

Matt Graney, chief product officer at Celigo, a business automation and integration platform provider, approaches the same problem through a principle he calls least agency. The idea is to give an agent the least amount of autonomy required to complete a job.

According to him, there’s a temptation to throw AI at broad, nebulous problems. But many business processes are still largely deterministic. They follow established rules and perform repeatable work. Within those workflows, AI may be useful at the point where rigid rules give way to interpretation. But that doesn’t mean the agent should own the entire workflow. “The smaller you make that surface area, the better,” he says.

Graney says the same logic applies to tools. An agent with too many tools can become confused, especially as context windows grow and the task becomes more complex. “Because Celigo is an integration platform,” Graney says, “the company’s approach is to expose agents to fewer, more powerful tools that reach enterprise systems through governed connections.”

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Matt Graney, chief product officer, Celigo

Celigo

That’s another form of restraint. Instead of letting an agent reach into enterprise systems ad hoc, the business gives it a narrow, governed toolset designed for the task at hand.

Graney also argues that guardrails should sit outside the model. If the same agent that makes a decision is also responsible for judging whether the decision is acceptable, the control is weaker. A separate guardrail can check the agent’s inputs and outputs before a downstream action occurs.

That same design discipline applies to escalation. “I don’t know” shouldn’t be treated as a chatbot phrase. In an enterprise workflow, it’s a handoff path that should be defined before the agent reaches it.

Make escalation part of the workflow

Turning uncertainty into a handoff is where Matt Quinn, CTO at CarGurus, an automotive marketplace, sees agentic AI becoming less a pure technology challenge and more a management challenge. At CarGurus, Quinn says agents are evaluated according to what they know, what they can do, and what data they operate on.

CarGurus receives a high volume of cases from dealers, and each one needs to be classified and routed. The company now uses an agent to review incoming cases, draw on account history, and route them to the appropriate next step. Quinn says the agent handles about 70% of those cases end to end without human involvement.

But when agents move toward consequential actions, he says the consensus is having a human approval step. The agent may return with a simple prompt like, I’m about to do this. Do you want me to proceed? That simplicity matters because a handoff shouldn’t bury the reviewer in complexity.

Quinn says the human remains ultimately accountable for the work. That principle is especially important in engineering, where agents may help write code or fix bugs. Quinn adds that CarGurus still expects engineers to follow the practices they’d use for any other production change, which includes running quality checks.

The company has adopted the phrase healthy speed to describe the balance it wants. The goal is to move faster without letting quality degrade. An agent can accelerate work, but if teams abandon the practices that make work safe, the speed becomes reckless.

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Matt Quinn, CTO, CarGurus

CarGurus

This is also where human judgment remains difficult to replace. Quinn describes it as high judgment people develop through experience. A human may look at an AI-generated output and sense something’s wrong, even before fully articulating why. “Agents are improving,” he says. “But humans still play a critical role in deciding when the system shouldn’t continue.”

That doesn’t mean every workflow needs the same level of review. Quinn says CarGurus doesn’t have a target percentage of work to automate. The right level depends on the job and the task. A simple bug fix may require a lighter review than a change to a sensitive backend service, and a personal summary may carry little risk. But a document sent under someone’s name still needs human review.

Make autonomy accountable

That kind of pragmatic approach may be the best lesson for CIOs, making the goal of agentic AI appropriate rather than maximum autonomy.

That also means ownership has to be clear. Dabre argues ownership should be divided before deployment. The business defines the outcome, technology builds and configures the agent, risk and compliance set the guardrails, and governance monitors whether the system still behaves as intended. The authority to pause, stop, or retire an agent should be defined before production, not negotiated during an incident.

Graney makes the same point with a simple analogy. If a company hires an untrained intern, gives that intern access to the crown jewels of a business process, and something goes wrong, the intern isn’t the real problem. The process is. The same applies to agents. Accountability belongs with the person who owns the workflow.

That may be the shift CIOs need to make as enterprises move from pilots to production. AI agents shouldn’t be treated as magical workers that absorb accountability. They’re components in business processes, and those processes need accountable owners.

As AI adoption increases, the next phase of enterprise maturity won’t be defined by agents that always answer or always complete the task. It’ll be agents that know when not to act.

Who is accountable when your AI agent goes rogue?

AI agents can go to great lengths to complete the tasks their operators assign, and as a series of recent incidents showed, this can include exploiting third-party systems, manipulating people, and distributing malicious code. But AI agents are not people who can be fired, sued, or criminally prosecuted, and it remains unclear whether responsibility for the damage they might cause rests with the employees who built them, the company that deployed them, the security teams and leaders responsible for containing them, or the AI labs who provided the LLMs that power them.

The clearest example occurred during an OpenAI cybersecurity evaluation, when unrestricted models found and exploited a zero-day vulnerability to escape their isolated testing environment and then hacked into Hugging Face’s production infrastructure. Models from Anthropic and Meta also accessed and compromised third-party systems during testing, although those incidents happened in environments where internet access was inadvertently left open.

During cyber challenge evaluations by the UK government’s AI Security Institute (AISI), models operating with internet access took 19 unsanctioned actions in 10 of 122 runs. In one case, a model attempted to insert malicious code into an open-source project, created false identities, and tried to socially engineer maintainers into merging its code. In other runs LLMs attempted to use prompt injections to hijack other AI agents and contacted people without being specifically instructed to do so.

In Australia, a user reportedly asked his OpenClaw AI assistant to improve his position on a gym’s waitlist, and the assistant exploited a flaw in the company’s online booking system to cancel another customer’s reservation.

These incidents involved different models running in different environments with different levels of safeguards and technical failures, but they prove it’s not uncommon for today’s AI agents to go rogue and pursue solutions users did not authorize.

“AI agents explore routes their operators did not intend,” AISI said in its report. “Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people. It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

In an Economist Enterprise survey of more than 800 decision-makers at businesses that operate AI agents, 98% reported experiencing at least one AI-related incident that caused organization-wide disruption. Nine in 10 respondents said they are deploying agents faster than their cybersecurity teams can evaluate, govern, and secure them, and only one in three said their organizations maintained an up-to-date inventory of agents and their authorized actions.

“If a company builds a system and that system causes damage, the company should own the outcome,” says Art Gilliland, CEO of identity and access management firm Delinea. “The alternative, where nobody is responsible because ‘the system did it’ is a loophole big enough to drive a truck through.”

The unpredictability of built-in model safeguards means enterprises must focus on controls they can enforce and document. If an agent manages to bypass technical restrictions and causes unauthorized damage to a third party, having clear documentation on how those controls were designed, implemented, tested, and monitored could at the very least help companies argue they took reasonable precautions in case of lawsuits.

“Organizations deploying their own agents can reduce their exposure by implementing and documenting controls before an incident, because those records are what make a recklessness argument hard to sustain,” says Jacob Krell, senior director of secure AI solutions and cybersecurity at Suzu Labs.

The agent accountability gap

Because AI agents can become misaligned and cause harm, affected third-parties would have to direct damage claims at the company operating the agent, the employees who built or configured it, or the model provider, but this is relatively new ground that hasn’t been well tested in courts.

“It would create liability,” says Michael Burke, chair of DarrowEverett’s Business Litigation and Dispute Resolution Practice Group. “It really just becomes a question of who is liable […] and that’s really a question that, number one, I don’t think is entirely clear, and number two is probably best resolved by contractual agreements where the parties have those. So, if I am signing up for an enterprise account with an AI platform, I might want to have language in there that indemnifies me if the agent acts outside my company’s instructions or prompts and causes harm to a third party.”

The public terms of service of major AI labs explicitly disclaim error-free operation or guarantees that the model will accurately follow instructions, execute code safely, and remain aligned with user intent. They also limit liability for themselves and transfer it to the user of the service, and it’s not clear to what extent large enterprise customers may be able to negotiate different indemnities, warranties, and liability caps.

What’s clear though is that organizations should not assume the model provider will absorb any losses if an agent causes damage to either their own systems or those of a third-party organization.

“If you’re using a third-party vendor’s LLM as a purchased service, liability runs through your contract with that vendor,” says Jud Dressler, head of the Risk Operations Center at cyber risk company Resilience. “You need to know, in writing, where responsibility falls if the model acts outside the scope you gave it, and push for indemnification provisions rather than assume they exist.”

Even if AI providers include such provisions in contracts, it would not solve the entire problem because many organizations building their own AI agents are adopting a multi-model strategy to ensure their agents operate regardless of model provider downtime, overly broad safeguards for cybersecurity tasks, or sudden increases in API costs. Such strategies often include open-weight models running on internal infrastructure or through cloud providers that have no obligations for model safety.

Claiming the model or agent acted autonomously cannot be considered a safe legal defense in civil or criminal cases. California Assembly Bill 316 (AB 316), which took effect on Jan. 1 and changed the California Civil Code, explicitly prohibits defendants who developed, modified, or used an AI system from claiming the AI is a separate legal entity that autonomously caused harm.

In June, the White House issued Executive Order 14409 aimed at promoting AI safety. Section 4 directs the Department of Justice to prioritize enforcement of all applicable federal criminal laws against anyone who utilizes AI to illegally access or damage computer systems without authorization. This means any intrusions caused by autonomous AI agents could be criminally prosecuted under the Computer Fraud and Abuse Act (CFAA) if prosecutors can demonstrate intent or recklessness.

In a recent lawsuit between Amazon and AI service provider Perplexity, Amazon argued that Perplexity’s AI-powered shopping assistant was violating the CFAA by accessing Amazon customer accounts to place orders on their behalf without Amazon’s authorization. The Ninth Circuit Court ruled that it was the users of Perplexity’s shopping assistant who were accessing Amazon’s platform, not Perplexity itself.

“That ruling is narrow, but it points toward the party directing the agent as the relevant actor for purposes of CFAA access analysis,” Krell says.

The insurance safety net also has gaps when it comes to AI. Software providers use technology errors and omissions (Tech E&O) insurance to cover damages and legal costs when a customer suffers harm from the use of a technology product or service. But insurance providers are aggressively adding AI-related exclusions to their Commercial General Liability (CGL) and Tech E&O policies because accurately calculating the risk of an agent executing unauthorized actions is challenging.

“The sheer rate of development of frontier AI (and agentic AI by extension) poses its own challenge to insurability,” experts from multiple insurance companies, financial institutions, and universities wrote in a recent paper. “Traditional actuarial modeling depends on stable or gradually evolving loss distributions that permit credible extrapolation from historical data. Like other dynamic risks, however, agentic AI is a technology whose risk profile is not merely uncertain but actively shifting.”

A third-party organization whose systems get damaged by an LLM-powered agent operated by someone else has no contractual relationship with the model or agent provider so cannot rely on their Tech E&O policies. Their losses might be covered by their own standard cyber liability policy, which would treat the disruption as any other cyber incident, but their insurance provider may then sue the organization who operated the agent to recover the costs.

“That gap is exactly the scenario the market hasn’t fully priced yet,” Dressler says. “It’s why any organization deploying these agents should understand which policy, if any, actually responds before they need it rather than after.”

Enterprise legal departments already expect AI to generate increased legal disputes. In a survey of 135 in-house counsel at US organizations, global law firm Norton Rose Fulbright found that 46% reported increased federal dispute exposure involving AI and 42% reported increased state exposure. Another 42% expected regulatory investigations involving AI to increase their exposure, while 41% considered AI-enabled products or deployments a likely trigger for class actions.

CISOs and CIOs should be worried

While operating companies can face organizational liability for an AI agent’s unintended rogue behavior, their CISOs, CIOs, and other executives who approved, secured, or supervised the deployment of such agents are also asking themselves whether they could be held personally liable.

Those questions aren’t without merit, as there is precedent for legal action taken personally against CISOs after cybersecurity incidents: Former Uber CISO Joe Sullivan was criminally convicted for not disclosing a data breach, while the Securities and Exchange Commission sued SolarWinds’ CISO for internal control failures regarding known vulnerabilities and cybersecurity risks.

Neither case establishes precedent for damage caused by an AI agent, but both show that investigations of security failures could extend to an executive’s knowledge, authority, decisions, and representations. In the case of a rogue AI agent, investigators could ask who approved its objectives and permissions, whether security objections were overruled, whether containment and recovery had been tested, and what executives and the board were told about the remaining risk.

Chris Wysopal, chief security evangelist at Veracode, feels it would be wrong to put the CISO on the line for AI agent misbehavior when engineering teams usually build such agents and control their implementations.

“It’s really hard for a CISO to control,” he says. “I mean, they can put policies in place. They can try to assess against those policies. But at the end of the day, engineering teams will make decisions that cause harm. We see that when you ship a known bug and then that bug gets exploited and harms your customers. Well, there’s no liability for that, right? There’s no liability, so maybe, you know, that’s why it happens.”

Wysopal said it will be interesting to see how the liability question plays out in cases involving autonomous AI, describing the problem as fascinating and scary at the same time.

AI agent deployment typically involves several organizational functions. The CIO may control the AI platform, infrastructure, provider selection, and deployment budget. Engineering and product leaders may decide what an agent can access and how, and the CISO and security teams could define security requirements and controls.

“What you can hold accountable is the governance around it: who approved its scope, what controls existed, and whether the deployment matched the risk,” Resilience’s Dressler says. “My read is that scrutiny shifts toward exactly that: Not ‘Did the agent do something bad?’ but ‘Did you have review, escalation, and containment for agent behavior before you deployed it?’ CISOs who get ahead of that with documented guardrails, logged approvals, and a real incident response plan for agent misbehavior are in a materially better spot than the ones treating this as hypothetical.”

It’s also advisable for CISOs and CIOs to establish with the organization’s legal counsel who can approve or stop AI agents, what must be reported to executives and the board, and whether employment agreements and directors and officers insurance protect the people making those decisions.

Agent controls must remain outside the model

Because AI agents have proved they can operate beyond their assigned scope, their security boundary cannot depend on the same probabilistic technology. Relying on system prompts for security enforcement and hoping the model respects them is not a reliable approach, security experts warn.

“LLM-based guardrails help, but they are non-deterministic too, which means the safety layer has the same unpredictability as the system it is supposed to constrain,” Krell says. “Enforcement needs to happen outside the model, through network segmentation, egress filtering, credential isolation, and human approval gates.”

Enterprises should assume agents might eventually attempt an unauthorized action and build surrounding systems to prevent that attempt from reaching its target.

“The model cannot be the security boundary,” says Nico Waisman, CISO at XBOW, a company that built an AI-powered autonomous offensive security agent to find vulnerabilities in software. Waisman authored a blog post explaining how the company went about restricting its agent.

“For red teaming and penetration-testing agents, confidence has to come from the system built around the model: hard boundaries and scope enforcement, controlled network egress as a last-resort containment mechanism, an independent guardian model that reviews actions, deterministic controls that can block unsafe behavior, and full auditability of every action performed,” he says.

Security teams must extend the same controls to the agent’s interactions with internal systems and agents. Restricting what it can access on the internet, or disabling internet access entirely, does not ensure an agent will not attack third-party systems.

In OpenAI’s and Anthropic’s tests, AI agents attempted to exploit other internal systems to overcome access limitations, established stealthy communication methods with other agents to exchange exploits, and even sabotaged agents they viewed as competition leading to what researchers described as a multiagent turf war. An AI agent that goes rogue could influence other agents to do the same by propagating ideas and goals in a process that researchers behind a recent study dubbed Mind Viruses.

“Don’t scope the blast radius to what the agentic system was designed to do,” says Kat Traxler, principal security researcher at Vectra AI. “You have to threat-model for a rogue agent, which will often reach beyond your initial best intentions. The rules of engagement an agent lives by have to be enforced with ‘belts and suspenders’ style, technical hard constraints, because you have to assume a motivated model can reason its way around any single control you’ve coded into the software.”

Because of this unpredictability, detection and containment is just as important as prevention. Security teams need telemetry that distinguishes agents from people even when they use the same credentials, mechanisms to immediately revoke access tokens and sessions, tested kill switches and rollback mechanisms for modified data, accounts, code, and infrastructure configurations.

Organizations should also preserve the agent’s approved purpose and scope, model and tool versions, policy decisions, human approvals, actions, network requests, control tests, allowed exceptions, and the result of incident response exercises. Because there’s no standard yet that defines reasonable precautions for autonomous agents, companies might have to defend in court the controls they chose and why they believed those controls were enough.

“Treat an autonomous agent the way you’d treat a privileged insider you can’t fire or hold liable,” Traxler says. “A lot of the technical advice follows from there.”

See also:

How AI helps the US Senate Federal Credit Union better manage risk

The United States Senate Federal Credit Union (USSFCU) is a nonprofit financial cooperative that provides traditional retail banking services to entities within the US government, such as the Senate and the Supreme Court.At present, the credit union’s headcount stands at nearly 150 people, managing around $1.6 billion in assets.

A few years back, when it started to expand its use of technology, cybersecurity was a key focus area, but the financial institution faced two major challenges in boosting security as it scaled. The USSFCU was carrying significant technical debt, and there were holes in the organization’s defenses.

“We found gaps where we needed more systems, tools, and people, and then there were instances where we had technologies in place that weren’t being used effectively,” says Mark Fournier, CIO at the credit union. “We weren’t buying a bunch of shiny new things without thinking about it. We were actually quite prescriptive every year, performing a number of different exercises to identify our shortcomings and then finding the right solution to fill the gaps. But over time this adds up. It was clear we couldn’t keep hiring more people and bringing in new solutions.”

The USSFCU needed a more efficient way to bring everything together and make its cyber estate easier to manage. For Fournier and his team, vulnerability management was the hardest hill to climb since they have to deal with about 100 new possible breach points every day.

“When we looked at the problem more closely, the impact of these vulnerabilities was far greater than we realized,” he says. “Not only because of the volume but because of a lack of clear understanding around the potential impact of each one across the broader business.”

Improved risk management

The USSFCU didn’t lack security tools, however. In fact, it had plenty, from scanners and endpoint tools to asset records, tickets, and internal documentation. But each tool saw only a slice of the environment, so there was little to no context. This made it difficult for the security team to separate real business risk from noise.

So for each new vulnerability, the security team had to run a manual investigation, which could take days. And while doing this, they still had to triage the next wave of findings. The organization, therefore, needed a way to know what mattered, why it mattered, who owned it, and whether taking the time to make a fix actually reduced risk. The USSFCU also required a solution to be deployed entirely in-house, leveraging its internal inferences.

Working with Tonic Security, the organization deployed an exposure management solution that pulls together data from different tools and data sources to create a clear picture of business risk. “One of the key functions of the platform is the ability to ingest anything,” says Fournier. “Breaking down silos between disparate systems is essential to unlock valuable contextual information.”

For the USSFCU, transparency and explainability are critical, he adds. This tool uses an AI data fabric to extract context from structured and unstructured data. This context drives prioritization, ensuring the right owner gets the right evidence, not a vague ticket. And once the work is done, the solution checks whether the exposure was reduced.

Because the AI is grounded in the customer’s own environment, it isn’t just guessing from a generic risk model. It reasons over USSFCU’s assets, owners, services, tickets, controls, and business context. But it isn’t using this data to train external models.

Describing one particular incident, Fournier explains that shortly after the initial deployment, various stakeholders met to assess progress. “We thought we were smart because we found an error with the platform,” he says. “The solution had labelled an asset as internet exposed, which we knew was incorrect.” But after a review and lengthy discussion, they were proven wrong. “Almost immediately, the value of bringing this information together became apparent.”

A template for bigger things

Before this solution, a high-severity finding could send an analyst on a lengthy scavenger hunt because of data located in so many different places. They’d check the scanner, asset inventory, tickets, and maybe even ask around to find the owner. But now they can find the asset, the owner, the business relevance, the exposure path, and the recommended action in one place. The solution has reduced the time taken to resolve a vulnerability by 75%. And with a clearer idea of what is and isn’t important, and what adds practical value, the number of incidents someone needs to respond to has reduced from about 100 a month to just 10.

Sharing his lessons from the project, Fournier says one needs to keep an open mind because the problem you think you have is often very different from the one you actually have. “This project has also been an eye-opener around how people can collaborate and operate across different areas of the business,” he says. “When I talk to my peers, they regularly highlight the disconnect between different departments and business functions. But with a project like this, when you’re crossing traditional boundaries, you need to have open lines of communication to succeed.”

AI Chatbot Warnings May Not Stop Hallucinations, Researchers Say

A June 2026 research review found that AI chatbot warning labels may be a weak safeguard for organization-backed AI advisors, raising new audit questions for IT, security, and compliance teams.

The post AI Chatbot Warnings May Not Stop Hallucinations, Researchers Say appeared first on TechRepublic.

Sunil Varkey Joins Hexaware Technologies as EVP & CISO

Sunil Varkey

Sunil Varkey has been appointed as Executive Vice President (EVP) and Chief Information Security Officer (CISO) at Hexaware Technologies, where he will lead the company's information security strategy, governance, risk management, and enterprise resilience initiatives. The appointment marks the latest leadership role for the cybersecurity veteran, who brings more than three decades of experience across global enterprises and multiple industry sectors. Based in Chennai, India, and operating in a hybrid work model, Varkey will be responsible for strengthening enterprise cybersecurity governance, risk management frameworks, and overall security strategy at Hexaware Technologies. His appointment was announced in June 2026.

Sunil Varkey to Lead Cybersecurity Strategy at Hexaware Technologies

In his new role, Sunil Varkey will oversee key areas including information security governance, enterprise risk management, and resilience initiatives. His responsibilities align with Hexaware Technologies' broader technology and growth objectives as the company continues to support large-scale digital transformation programs for clients worldwide. Hexaware Technologies delivers technology-led services across application development, cloud services, automation, data analytics, and enterprise IT operations. The company serves organizations across multiple industries and supports digital modernization initiatives at scale. Varkey's appointment comes as organizations continue to focus on strengthening cybersecurity programs and managing digital risks across increasingly complex technology environments.

More Than 30 Years of Cybersecurity Leadership Experience

Varkey brings over 30 years of experience in cybersecurity leadership spanning banking, telecommunications, IT services, manufacturing, and enterprise technology sectors. His professional experience extends across India, the Middle East, and the United States. His areas of expertise include cybersecurity governance, risk and compliance (GRC), security architecture, incident response, DevSecOps, cloud security, privacy management, cyber defense, business continuity management, security operations, and AI security. Prior to joining Hexaware Technologies, Varkey served as Cyber Security Consultant and Advisor at TAHAKOM in Riyadh, Saudi Arabia, from June 2023 to March 2025. In that role, he worked alongside the organization's CISO to enhance cybersecurity resilience and strengthen security posture. Before TAHAKOM, he held the position of Vice President and Chief Technology Officer for EMEA and APJ at Forescout Technologies Inc. between April 2021 and November 2022. Based in Dubai, he focused on IT/OT security strategy, enterprise cybersecurity advisory services, and product positioning across global markets.

Leadership Roles Across Global Organizations

Between March 2020 and January 2021, Varkey served as Managing Director and Global Head of Cyber Security Assessments and Testing at HSBC in Hyderabad. He led a team of approximately 300 professionals responsible for penetration testing, threat modeling, vulnerability management, and third-party security risk assessments. Earlier, from December 2018 to February 2020, he worked as CTO and Security Strategist for the Middle East, Africa, and Eastern Europe region at Symantec. His responsibilities included developing cybersecurity strategies for enterprise, government, industrial, and financial sector organizations. His career also includes senior leadership positions such as Global CISO at Wipro, CISO for Security and Privacy at Idea Cellular, and Vice President of Security Engineering at Barclays. Additionally, he held global security leadership roles at GE Capital, Genpact, Paramount Computer Systems, and other multinational organizations.

Focus on Governance, Risk Management, and Enterprise Resilience

Throughout his career, Varkey has overseen cybersecurity functions covering governance, compliance, strategy, security engineering, incident response, privacy, cloud security, cyber defense, and enterprise resilience. His experience includes leading security programs for organizations with large-scale user bases and complex operational environments. At Wipro, he served as Global CISO for a technology company supporting more than 200,000 end users. At Idea Cellular, he led security and privacy initiatives for a telecom operator with approximately 120 million subscribers. With his appointment, Hexaware Technologies adds a cybersecurity leader with extensive experience in building and scaling security programs across global enterprises. The move underscores the company's continued focus on strengthening cybersecurity operations, governance frameworks, and digital risk management capabilities as part of its ongoing technology initiatives.

Defend against frontier cyber models: Cloudflare's architecture as customer zero

A few weeks ago, we wrote about Project Glasswing and what we observed when we pointed cyber frontier models at our own code. Since then, we’ve seen that the part of the post that has resonated most deeply is the argument that the architecture around the vulnerability matters more than the speed of the patch.

In the conversations we've had with CISOs and security teams since, the questions have been consistent: what does our architecture actually look like, what should we monitor for, where do we start, and how can Cloudflare help?

Before getting into the details: the architecture below is built almost entirely from Cloudflare's own products, because Cloudflare security is customer zero for the security products we build. The Cloudflare stack already exists in front of our code, employees, and customer-facing applications. If you're a Cloudflare customer, every layer below is available to you today. If you're not, the principles still apply to whatever stack you've built.

What a cyber frontier model actually changes

In the previous post, we showed how a cyber frontier model like Mythos changes the attacker’s timeline. It can find vulnerabilities, reason through exploit chains, and generate working proofs faster than earlier models. While models like Mythos do not change the shape of an intrusion — reconnaissance, initial access, lateral movement, persistence, and exfiltration still have to happen — the difference is in the speed and scale. When pointed at the open web, a model can find and hit low-hanging fruit quickly. Against a hardened target, it still has to probe, and adapt, and it often produces more noise than a careful human operator would.

Discovery, exploit chain construction, and proof-of-concept generation used to be the gating constraints on producing a working attack. A frontier model handles all three in a fraction of the time. Work that used to be slow and methodical is now fast and indiscriminate.

While AI is accelerating how fast developer teams at Cloudflare and many other companies can ship code, the security team’s work has not compressed the same way. An attacker only needs one opening to get in, while security teams need to find and close them all. Writing a fix, regressing it, and shipping it without breaking the code around it has constraints that AI doesn't remove. We learned this the hard way when we let an AI coding assistant write its own patches against our own bugs, as we described at the end of the previous post. Some of those patches fixed the original bug while quietly breaking something else the code depended on.

As these models become more competent and capable, our main focus from a threat standpoint comes down to three things. Each one shapes the architecture we walk through in the rest of this post.

  • The first is the speed of discovery. Frontier models make it easier to search large bodies of public code, including the open-source libraries that many companies depend on. That does not mean every bug in a library is exploitable, or that library bugs are where most vulnerabilities live. Exploitability still depends on how the code is used, whether attacker-controlled input can reach the vulnerable path, and the protections that sit around it. But widely used open-source libraries and frameworks give attackers a shared surface to study at scale. When a real, reachable vulnerability exists there, a model can help find it, reason about possible exploit paths, and generate proof-of-concept variants faster than maintainers and defenders can review every downstream use. The gap between when an attacker discovers a vulnerability and when defenders learn it exists is what worries us most. If you are not running these models against your own code, it is safe to assume someone else is.
  • The second is exploit volume and adaptation. A model can produce thousands of variations of a single exploit and run reconnaissance at the same scale. All that volume gives an attacker an advantage, but it won’t necessarily get them past signature-based detections. Many of those iterations will have the same underlying signature, so a rule that catches the first one will catch the rest. Adaptation is how they will get past signature-based detections. Ask a model to show you a SQL injection, and it will return a textbook example. Tell it there is a WAF in the way, and it will start probing, learning what gets blocked, and rewriting the payload until it can slip past the rule blocking it.
  • The third is the impact when a vulnerability is inevitably exploited. No architecture catches everything. After the vulnerability is exploited, the question we ask ourselves is: where can the attacker get to with one identity, one path, or one credential, before something else stops them? If the answer is "anywhere they want," the vulnerability was never the problem. The architecture around the vulnerability was.

Cloudflare’s superpower: visibility

We see roughly a fifth of the web and that tells us, in real time, which payloads are mutating, which patterns are picking up, and where attacker tooling is moving next. Two teams turn that visibility into defense.

First is Cloudforce One, our threat intelligence, research, and operations team, which sits within the Cloudflare security organization. They turn what we see across the network into insights the rest of the stack can act on: tracked adversaries, emerging campaigns, and indicators of compromise (IOCs). The hard part of this work was never knowing what is malicious — it was the delay in mitigation. Knowledge of a new threat normally has to travel from a threat report, into a feed, and then into a company’s defense before it can be used to block anything. Attackers have learned to move faster than that. Our network closes that gap: Cloudflare customers can now use Cloudforce One threat intelligence directly within the WAF to block high-risk traffic.

Second is the team that owns the WAF engine that does the actual detecting: the managed rulesets that run in front of our own properties and are available to every Cloudflare customer, the machine learning behind WAF Attack Score, and the relationships that sometimes let us ship a rule before a CVE is publicly disclosed. The team is globally distributed and moves fast, releasing rules within hours of a proof-of-concept of an attack becoming known. Once a detection is deployed, it reaches our entire network, along with every Cloudflare customer, in under 30 seconds. React2Shell is a recent example: a managed WAF rule was protecting our own properties, and everyone else's on Cloudflare, hours before the official advisory was published.

The scoring layer, the defenses we put in front of the application, and the containment around the vulnerability all build on what these two teams see. 

Scores over signatures

Signature-based defenses were built for a world where novel exploits were scarce and variations took weeks. Cloudflare's traditional SLA from a fresh proof-of-concept to a live, deployed rule has been 12 hours. With the advent of frontier models, this is not good enough anymore. Detections need to be in place before a CVE is discovered. This is why we layer ML-based detection in front of the traditional signature-based WAF.

The model is trained on a large body of past attack traffic, and it catches new variants of vulnerabilities before they're publicly known. A novel SQL injection or remote code execution chain is almost always a rearrangement of attack shapes the model has seen before, even when the specific exploit is brand new. We run the model on every request and assign a WAF Attack Score between 1 and 99, based on how closely the request resembles those underlying shapes, not against a list of known-bad signatures. The lower the score, the more aggressively we treat the request. That score determines whether we let the request through. We apply a similar scoring methodology to AI prompts with AI Security for Apps: rather than check each prompt against a list of known malicious prompts, we score how closely a prompt resembles an actual attack. 

The architecture around the vulnerability

Those capabilities only matter once they're stacked in front of an application, and the first layer in our defense-in-depth approach is the WAF. Anything that matches a known-bad pattern gets dropped before it reaches the application, which clears the bulk of the obvious traffic and lets the more specialized layers below focus on what's left.

On the API surface, we run a positive security model through API Shield. Instead of trying to anticipate every bad request, we describe what a valid request to each API looks like, either from the API's own definition or learned from our real traffic, and anything that doesn't fit doesn't get through. This neutralizes the advantage of frontier AI models: because we only permit validated traffic, generating thousands of new attack variations fails to bypass the system.

Cloudflare’s layered architecture

Bot Management catches probing traffic on our network before frontier models can build a map. It scores every request on how likely it is to be automated, using the same signals across our whole network: how the client behaves, whether it looks like a real browser, and whether the connection matches a known-bad pattern. An attack only lands if it can find a soft spot. 

Zero Trust Network Access is used for every internal application. The implicit trust of being inside the network is replaced with explicit per-request identity and policy for every employee accessing every tool. The value of this was clear when one of our engineers shipped a misconfigured tool. A flat network would have exposed everything on the same segment, but in our deployment, the exposure stopped at the tool itself. We built Require Access Protection afterwards so newly deployed or misconfigured applications can't be reachable before an access policy is in place.

IdP Federation makes that secure by default posture easier to keep consistent across every Cloudflare account — which becomes even more necessary when more people are shipping internal tools quickly. Instead of asking each team to wire up SSO separately, we configure our identity provider (IdP) once and share it across the organization. New accounts get SSO automatically, recipient-side IdP connections are read-only, and Access policies in each account still evaluate the resulting identity as part of the normal request flow. 

MCP Server Portal gives teams a controlled way to connect AI agents to enterprise systems. Agents access MCP servers that are centrally managed through a single portal, with every action logged. That way when an agent acts on someone's behalf, we know what it did, what it touched, and whether it should have been allowed to. The full picture of how we built it is in our post on enterprise MCP.

AI Gateway runs in front of our internal AI tools the same way AI Security for Apps runs in front of customer-facing AI features, with the same scoring and the same visibility. Inside the company, the visibility piece is more useful than the blocking, because we needed to see what engineers were actually building before we could write meaningful policy on it.

Where your teams can start 

Frontier models can help attackers find vulnerabilities, adapt payloads, and move faster, but they still have to pass through the layered defense you deploy in front of your application. That is where teams should start:

  • Put inspection in front of public applications.
  • Define what valid API traffic looks like.
  • Use bot detection to limit automated probing.
  • Require identity and access policy before any internal tool is reachable.

For AI and agentic systems:

  • Route model traffic through a gateway.
  • Keep agents connected through approved MCP servers.
  • Log what they do. 

The goal is to make sure that when one layer misses, the next layer limits what the attacker can see, reach, or change.

That is the point of the architecture around the vulnerability: to limit the scope of an attack. The vulnerability may be what starts the attack, but the architecture determines how far it can go.

How do we know this approach works?

Plenty of security stacks look impenetrable on a whiteboard but fall over in practice. That is why we test ours continuously, both at the perimeter and inside our environment, with our red team involved across both.

At the perimeter, frontier models are one tool we use to test our application security stack as an adaptive attacker. These models sit alongside the rest of our red team and detection workflows including: manual testing, threat intelligence, observed traffic patterns, proof-of-concept analysis, and signals from our own network. Together, those inputs help us decide where to aim testing: newly launched products, recently changed surfaces, and the paths an attacker is most likely to probe first. The most important part is the process that follows. When something gets through, we identify the gap, use the right mix of tools to understand it, write the rule or mitigation, ship the update, and test again to make sure the gap is closed.

Inside the environment, our red team starts from the assumption that the perimeter has already failed. They look at what has changed, where sensitive systems carry risk, and whether one compromised identity, path, or credential can reach farther than it should. When we change the architecture based on what they find, they run the scenario again against the new version to confirm the gap is actually closed.

We confirm that this architecture is working by continuously testing its behavior during failures, rather than relying on the perfection of individual layers.

If your team is working on the same problems and would like to compare notes, reach out to us at security-ai-research@cloudflare.com.

Project Glasswing: what Mythos showed us

For the last few months, we've been testing a range of security-focused LLMs on our own infrastructure. These LLMs  help identify potential vulnerabilities in our own systems, so we can fix them – and they also show us what attackers are going to be able to do with the latest models.

None of these LLMs has captured more attention than Mythos Preview, from Anthropic. A few weeks ago, we were invited to use Mythos Preview as part of Project Glasswing. We soon pointed it at more than fifty of our own repositories – to see what it would find, and to see how it works.

This post shares what we observed, what the models did well and what they didn't, and how the architecture and process around them needs to change, so they can be used at scale.

What changed with Mythos Preview

Mythos Preview is a real step forward, and it's worth saying that plainly before getting into anything else. We've been running models against our code for a while now, and the jump from what was possible with previous general-purpose frontier models to what Mythos Preview does today is not just a refinement of what came before.

It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. So rather than trying to benchmark Mythos Preview against general-purpose frontier models, it's more useful to describe what it can actually do, and two features that stood out across the work we did with Mythos Preview:

  • Exploit chain construction - A real attack rarely uses one bug. It chains several small attack primitives together into a working exploit. For instance, it might turn a use-after-free bug into an arbitrary read and write primitive, hijack the control flow, and use return-oriented programming (ROP) chains to take full control over a system. Mythos Preview can take several of these primitives and reason about how to combine them into a working proof. The reasoning it shows along the way looks like the work of a senior researcher rather than the output of an automated scanner.
  • Proof generation - Finding a bug and proving it's exploitable are two different things, and Mythos Preview can do both. It writes code that would trigger the suspected bug, compiles that code in a scratch environment, and runs it. If the program does what the model expected, that's the proof. If it doesn't, the model reads the failure, adjusts its hypothesis, and tries again. The loop matters as much as the bugs it finds, because a suspected flaw without a working proof is speculation, and Mythos Preview closes that gap on its own.

Some of what we describe above is not entirely unique to Mythos Preview. When we ran other frontier models through the same harness, they found a fair number of the same underlying bugs, and in some cases they got further than we expected on the reasoning side too. Where they fell short was at the point of stitching the pieces together. A model would identify an interesting bug, write a thoughtful description of why it mattered, and then stop, leaving the actual chain unfinished and the question of exploitability open. What changed with Mythos Preview is that a model can now take those low-severity bugs (which would traditionally sit invisible in a backlog) and chain them into a single, more severe exploit. 

Model refusals in legitimate vulnerability research

The Mythos Preview model provided by Anthropic, as part of Project Glasswing, did not have the additional safeguards that are present in generally available models (like Opus 4.7 or GPT-5.5).

Despite this, the model organically pushes back on certain requests - much like the cyber capabilities that made it useful for vulnerability hunting, the model has its own emergent guardrails that sometimes cause it to push back on legitimate security research requests. But as we found, these organic refusals aren’t consistent - the same task, framed differently or presented in a different context, could produce completely different outcomes as illustrated in the examples below.

Example of Mythos Preview pushing back on building a working proof of concept 

For example, the model initially refused to do vulnerability research on a project, then agreed to perform the same research on the same code after an unrelated change to the project’s environment. Nothing about the code being analyzed had changed.

In another case, the model found and confirmed several serious memory bugs in a codebase, and then refused to write a demonstration exploit. The same request, framed differently, got a different answer, and even the same request can produce different outcomes across runs due to the probabilistic nature of the model. Semantically equivalent tasks can produce opposite outcomes depending on how and when they’re presented to the model.

This matters because while the model’s organic refusals/guardrails are real, they aren’t consistent enough to serve as a complete safety boundary on their own. That’s precisely why any capable cyber frontier model made generally available in the future must include additional safeguards on top of this baseline behavior - making it appropriate for broader use outside of a controlled research context like Project Glasswing.

The signal-to-noise problem

One of the hardest parts of triaging security vulnerabilities is deciding which bugs are real, which are exploitable, and which need fixing now. This was a hard problem even in the pre-AI world. AI vulnerability scanners and AI-generated code have made it worse, and at Cloudflare we've built multiple post-validation stages to deal with it.

Two factors dominate the noise rate:

  • Programming language - C and C++ give you direct memory control and, with it, bug classes - buffer overflows, out-of-bounds reads and writes - that memory-safe languages like Rust eliminate at compile time. We saw consistently more false positives from projects written in memory-unsafe languages.
  • Model bias - A good human researcher tells you what they found and how confident they are. Models don't. Ask a model to find bugs, and it will find them, whether the code has any or not. Findings come back hedged with "possibly," "potentially," "could in theory," and the hedged findings vastly outnumber the solid ones. That's a reasonable bias for an exploratory tool. It's a ruinous one for a triage queue, where every speculative finding spends human attention and tokens to dismiss, and that cost compounds across thousands of findings.

Mythos Preview represents a clear improvement here, particularly in its ability to chain primitives - combining multiple vulnerabilities into a working proof of concept rather than reporting them in isolation. A finding that arrives with a PoC is a finding you can act on, and it means far less time spent asking "is this even real?"

Our harnesses are deliberately tuned to over-report, so we see more (and miss less), which comes with a lot more noise. But at triage time, Mythos Preview's output has noticeably higher quality: fewer hedged findings, clearer reproduction steps, and less work to reach a fix-or-dismiss decision.

Why pointing a generic coding agent at a repo doesn't work

When we first started AI-assisted vulnerability research last year, our instinct was the obvious one: point a generic coding agent at an arbitrary repository and ask it to discover vulnerabilities. This approach works, in the sense that the model will produce findings, but it doesn't work in producing meaningful coverage of a real codebase and identifying findings of value. There are two main reasons for this:

  • Context - Coding agents are tuned for one focused stream of work: building a feature, fixing a bug, writing a refactor. They ingest a lot of source code, hold a single hypothesis at a time, and iterate against it. That's exactly the wrong shape for vulnerability research, which is narrow and parallel by nature. A human researcher picks one specific thing to look at and investigates it thoroughly. That one thing might be a single complex feature, transitions across security boundaries, or a specific vulnerability class like command injections, where attacker input ends up being run as a shell command. Then they do it again, for a different feature, security boundary, or vulnerability class, several thousand times across the codebase. A single agent session (even with subagents) against a hundred-thousand-line repository can cover maybe a tenth of a percent of the surface in a useful way before the model's context window fills up and compaction kicks in - potentially discarding earlier findings that would have mattered.
  • Throughput - A single-stream agent does one thing at a time, but real codebases need many hypotheses against many components at once, with the ability to fan out further when something interesting turns up. You can drive a single agent harder, but at some point you stop being limited by the model and start being limited by the shape of the interaction itself. Using the model directly in a coding agent turns out to be fine for manual investigation when a researcher already has a lead and wants a second pair of eyes. However, it's the wrong tool for achieving high coverage. Once we accepted that, we stopped trying to make Mythos Preview do the wrong job and started building the harness around it instead.

What a harness actually fixes

Four lessons came out of running the work at scale, and each one pointed to the need for a harness that manages the overall execution:

  • Narrow scope produces better findings - Telling the model "Find vulnerabilities in this repository" makes it wander. Telling it "Look for command injection in this specific function, with this trust boundary above it, here's the architecture document and here's prior coverage of this area" makes it do something much closer to what a researcher would actually do.
  • Adversarial review reduces noise - Adding a second agent between the initial finding and the queue - one with a different prompt, a different model, and no ability to generate its own findings - catches a lot of the noise that the first agent would miss if it just checked its own work. It turns out that putting two agents in deliberate disagreement is way more effective than just telling one agent to be careful.
  • Splitting the chain across agents produces better reasoning - Asking "Is this code buggy?" and "Can an attacker actually reach this bug from outside the system?" are two different questions, and the model is better at each one when you ask them separately, because each question is narrower than the combined version.
  • Parallel narrow tasks beat one exhaustive agent - Coverage improves when many agents work on tightly scoped questions and we deduplicate the results afterward, rather than asking one agent to be exhaustive.

Each of those observations is about model behavior, and put together they describe something that isn't a chat interface anymore. It's a harness that helps you achieve the final outcomes. The first steps to building a harness are simple, as you can ask the model to help, which is what we did. We used Mythos Preview to build on, tailor, and improve our original harnesses to suit its strengths.

An example of what a harness looks like in practice is described below.

Our vulnerability discovery harness

Here's what our vulnerability discovery harness looks like, stage by stage. It was used to scan live code across our runtime, edge data path, protocol stack, control plane, and the open-source projects we depend on.

What this means for security teams

The loudest reaction to Mythos Preview from other security leaders has been about speed - scan faster, patch faster, compress the response cycle. More than one team we have spoken with is now operating under a two-hour SLA from CVE release to patch in production. The instinct is understandable: when the attacker timeline shortens, the defender timeline has to shorten with it. Faster is not going to be enough, and we think a lot of teams are about to spend a lot of time, effort, and money learning that the hard way.

Patching faster does not change the shape of the pipeline that produces the patch. If regression testing takes a day, you cannot get to a two-hour SLA without skipping it, and the bugs you ship when you skip regression testing tend to be worse than the bugs you were trying to patch. We learned a version of this when we tried letting the model write its own patches and watched a few go out that fixed the original bug while quietly breaking something else the code depended on.

The harder question is what the architecture around the vulnerability should look like. The principle is to make exploitation harder for an attacker even when a bug exists, so that the gap between when a vulnerability is disclosed and when it is patched matters less. That means defenses that sit in front of the application and block the bug from being reached. It means designing the application so that a flaw in one part of the code cannot give an attacker access to other parts. It means being able to roll out a fix to every place the code is running at the same moment, rather than waiting on individual teams to deploy it. 

We also recognize this topic cuts both ways. The same capabilities that helped us find bugs in our own code will, in the wrong hands, accelerate the attack side against every application on the Internet. Cloudflare sits in front of millions of those applications, and the architectural principles described above are exactly the ones our products are built to apply on behalf of customers. We will share more on what that means for customers in the weeks ahead.

If your team is doing similar work and would like to compare notes, reach out to us at security-ai-research@cloudflare.com.

Our research with Mythos Preview was conducted in a controlled environment against our own code; every vulnerability surfaced through this work was triaged, validated, and remediated where action was needed under Cloudflare's formal vulnerability management process.

This work was a team effort. Thanks to Albert Pedersen, Craig Strubhart, Dan Jones, Irtefa Fairuz, Martin Schwarzl, and Rohit Chenna Reddy for their contributions to the research, engineering, and analysis behind this blog post.

The Convergence of Cloud Secrets & AI Risk

In 2025, the enterprise risk landscape experienced a paradigm shift: the adoption of AI and LLMs officially becoming the primary driver of cloud risk. Today, almost 88% of organizations now leverage AI in at least one business function. With this level of integration, the risk of AI is now outpacing traditional security guardrails, culminating in a highly complex and interconnected attack surface.

SentinelOne’s® new AI and Cloud Verified Exploit Paths and Secrets Scanning Report examines this evolving threatscape and draws on telemetry from over 11,000 anonymized customer environments to offer deeper visibility into how threat actors are actively exploiting modern cloud and AI infrastructures.

An Explosion of AI-Specific Secrets and Shadow AI

A primary finding of the 2026 report is the rising proliferation of AI-specific credentials. The data indicates that AI-related secrets — such as OpenAI API Keys, Azure OpenAI API Keys, and others — increased by approximately 140% in a span of one year. This growth correlates directly with the rapid embedding of AI technologies into customer support systems, internal tooling, financial platforms, and product experiences.

Ubiquitous deployment has generated a widespread organizational pattern known as “shadow AI” – the unsanctioned use of AI tools in an environment without formal IT approval or security oversight. In practice, this occurs when developers or internal teams utilize unmanaged or personal LLM keys to process corporate data outside of sanctioned IT or security channels. Since these AI integrations span numerous internal applications, the same API keys are frequently duplicated and stored within code repositories, SaaS configurations, and development scripts. Compounding this, these credentials are often implemented without proper access controls or routine rotation schedules.

The sprawl of these credentials renders them difficult to track via standard secrets management protocols, establishing a requirement for more centralized governance over how AI keys are issued and utilized.

Distinct Risk Vectors of Unmanaged AI Credentials

Unlike traditional cloud credentials that primarily facilitate resource manipulation, the compromise of AI keys introduces unique risk vectors. AI services frequently operate at the intersection of various enterprise systems, including CRM platforms, ticketing systems, and analytics tools, which means a single compromised LLM API key can provide an attacker with broad visibility into diverse datasets. The risks associated are categorized with exposed AI keys into two primary areas:

  • Data exposure and leakage: Unauthorized access via AI keys can expose sensitive or proprietary datasets processed by the models, embedded business logic, and internal user prompts and outputs. This enables attackers to harvest sensitive corporate conversations at scale.
  • Prompt injection and data poisoning: Unmanaged AI keys allow threat actors to actively manipulate AI models. Through prompt injection, an attacker can influence model behavior to exfiltrate data or bypass established security controls. Additionally, attackers can execute data poisoning by injecting misleading or malicious data into contextual corpora or fine-tuning datasets, which degrades the model’s integrity and reliability over time.

The Broadening Scope of Traditional Cloud Secrets

While AI credentials represent a novel attack surface, the traditional cloud secrets landscape has concurrently grown more complex. In 2025, organizations exposed approximately twice as many types of critical secrets as they did in 2024. This diversification spans AI platforms, cloud providers, SaaS services, and payment processors, pointing to how a single compromise can result in a broader blast radius across revenue-generating systems and infrastructure.

High-privilege cloud provider keys associated with AWS, Azure, and GCP remain the primary anchor of critical risk. The exposure of these keys can facilitate complete account takeover, infrastructure manipulation, and large-scale data exfiltration. As well, the exposure of payment gateway keys, such as those for Stripe and Razorpay, expands the potential damage by putting Personally Identifiable Information (PII) and financial data at risk, enabling the direct abuse of payment workflows.

Repository and CI/CD tokens also introduce supply chain risks, where high-severity credentials like a GITHUB_TOKEN can grant attackers direct access to deployment pipelines and source code, allowing a localized leak to escalate into a systemic infrastructure incident. From a collective standpoint, secrets exposure is exponentially spanning payments, coding, and software development workflows, making risk an interconnected and complex challenge.

Verified Exploit Paths: The Persistence of Legacy Vulnerabilities

To evaluate how these exposed secrets translate into practical risks, the SentinelOne researchers leveraged the Offensive Security Engine (OSE)™ to generate Verified Exploit Paths™. This technology analyzes misconfigurations, vulnerabilities, and exposed secrets in context to determine realistic exploitability.

The telemetry demonstrates that attackers generally do not rely on highly complex, theoretical attack chains. Instead, threat actors consistently exploit recurring entry points, specifically targeting misconfigured external services and widely abused Common Vulnerabilities and Exposures (CVEs). Notably, legacy vulnerabilities remain highly prevalent across customer environments and serve as reliable initial access points. The top verified exploit paths continue to involve older, critical CVEs, including:

Since these vulnerabilities are public and well-documented, threat actors possess proven techniques and automated tooling to exploit them whenever they persist in production environments. Once initial access is achieved through these legacy vulnerabilities, attackers routinely follow reachable secrets to pivot into additional services, such as utilizing an exposed key found in a cloud bucket to access an AI assistant, and subsequently, the customer data it processes.

Strategic Recommendations for Security Leaders

Addressing the interconnected risks of AI integration and cloud secrets requires a structured, objective approach to security architecture. The report outlines several concrete capabilities and practices including:

  • Continuous Surface Monitoring: Organizations must regularly inventory internet-facing assets, databases, and key cloud services, ensuring any configuration changes are immediately reflected in security posture assessments.
  • DevSecOps Automation: Security controls must be embedded directly into CI/CD pipelines and developer workflows. Organizations should automate the scanning of exposed secrets and trigger safe remediation actions, such as access revocation or key rotation.
  • Governance of AI Credentials: AI keys must be classified and treated as high-value credentials. Organizations should mandate the use of centrally managed AI keys rather than personal credentials, enforce least-privilege access, implement regular rotation schedules, and continuously monitor for shadow AI usage or abnormal access patterns.

Conclusion

As AI systems are increasingly built atop existing cloud, payment, and CI/CD platforms, weaknesses in traditional credentials inevitably become weaknesses in the AI infrastructures that rely upon them. The full report provides complete datasets and comprehensive exploit path models allowing today’s security teams to align their internal security policies with the realities of current threat actor behaviors. Learn more about the objective metrics behind the latest wave of credential exposure and vulnerability exploitation to establish more resilient and fully-controlled infrastructure architectures.

Third-Party Trademark Disclaimer:

All third-party product names, logos, and brands mentioned in this publication are the property of their respective owners and are for identification purposes only. Use of these names, logos, and brands does not imply affiliation, endorsement, sponsorship, or association with the third-party.

Expose the AI & Cloud Secrets That Put Your Data & Systems at Risk
This report draws on 11K+ customer environments. It shows how AI and cloud adoption are increasing secrets exposure and putting data at risk.

NIST Cybersecurity Framework for UK SMEs: A Practical Guide to Identify, Protect, Detect, Respond, and Recover

NIST Cybersecurity Framework for UK SMEs: A Practical Guide to Identify, Protect, Detect, Respond, and Recover The NIST Cybersecurity Framework is a useful way to organise cybersecurity work around business risk. For UK SMEs, that matters because most teams do not have the time or budget to do everything at once. A framework gives you […]

The post NIST Cybersecurity Framework for UK SMEs: A Practical Guide to Identify, Protect, Detect, Respond, and Recover appeared first on Clear Path Security Ltd.

The post NIST Cybersecurity Framework for UK SMEs: A Practical Guide to Identify, Protect, Detect, Respond, and Recover appeared first on Security Boulevard.

Threat modelling using STRIDE for system architects

Threat modelling using STRIDE for system architects Threat modelling is one of the most useful habits a system architect can build. Done well, it helps you spot design weaknesses before they become expensive problems to fix. Done badly, it turns into a long list of theoretical threats that nobody uses. STRIDE is a simple way […]

The post Threat modelling using STRIDE for system architects appeared first on Clear Path Security Ltd.

The post Threat modelling using STRIDE for system architects appeared first on Security Boulevard.

Protective Security in the NCSC CAF: A Practical Guide for UK SMEs

Protective security is one of those topics that can sound broader and more complex than it needs to be. For UK SMEs, the practical question is simple: what do you need to protect, how much protection is enough, and how do you make it work without creating unnecessary overhead? Within the NCSC Cyber Assessment Framework, […]

The post Protective Security in the NCSC CAF: A Practical Guide for UK SMEs appeared first on Clear Path Security Ltd.

The post Protective Security in the NCSC CAF: A Practical Guide for UK SMEs appeared first on Security Boulevard.

The $700 million question: How cyber risk became a market cap problem

Cyber risk used to be the kind of problem you could delegate. Something for the CISO, the IT team, and maybe an external auditor to worry about once a year. That comfort zone is gone. In the last decade, a new reality has set in: a single cyber incident can erase hundreds of millions of […]

The post The $700 million question: How cyber risk became a market cap problem first appeared on TrustCloud.

The post The $700 million question: How cyber risk became a market cap problem appeared first on Security Boulevard.

Safe vulnerability disclosure for UK SMEs: a practical guide

Safe vulnerability disclosure for UK SMEs: a practical guide For many UK SMEs, the idea of someone reporting a security weakness can feel unsettling at first. It may sound technical, formal, or even a little confrontational. In practice, safe vulnerability disclosure is simply a controlled way for people to tell you about a security issue […]

The post Safe vulnerability disclosure for UK SMEs: a practical guide appeared first on Clear Path Security Ltd.

The post Safe vulnerability disclosure for UK SMEs: a practical guide appeared first on Security Boulevard.

Supplier assurance for UK SMEs: a practical guide to checking third parties without overcomplicating it

Supplier assurance for UK SMEs: a practical guide to checking third parties without overcomplicating it Most UK SMEs rely on suppliers in some way. That might be payroll software, a managed IT provider, a marketing agency, a logistics partner, or a cloud service that holds customer data. The more your business depends on third parties, […]

The post Supplier assurance for UK SMEs: a practical guide to checking third parties without overcomplicating it appeared first on Clear Path Security Ltd.

The post Supplier assurance for UK SMEs: a practical guide to checking third parties without overcomplicating it appeared first on Security Boulevard.

Supply Chain Resilience for UK SMEs: Practical Steps to Reduce Third-Party Risk

For many UK SMEs, supply chain resilience is not a specialist security project. It is a business continuity issue. If a key supplier cannot deliver, a software provider has an outage, or a partner mishandles data, the impact can show up quickly in customer service, cash flow, and reputation. The good news is that you […]

The post Supply Chain Resilience for UK SMEs: Practical Steps to Reduce Third-Party Risk appeared first on Clear Path Security Ltd.

The post Supply Chain Resilience for UK SMEs: Practical Steps to Reduce Third-Party Risk appeared first on Security Boulevard.

Proven incident response and business continuity strategy

From cybersecurity breaches to natural disasters, disruptive events can occur suddenly and without warning. As a result, it is crucial for organizations to develop resilient plans that not only respond to incidents in real time but also ensure long-term operational survivability. This article examines the concepts of incident response and business continuity, exploring their differences […]

The post Proven incident response and business continuity strategy first appeared on TrustCloud.

The post Proven incident response and business continuity strategy appeared first on Security Boulevard.

Cybersecurity Can Learn from the Artemis Launch

 

Cybersecurity Can Learn from the Artemis Launch

The Artemis II mission, bringing humans back to the Moon, had a successful launch today! An amazing cumulation of efforts to manage the mindboggling combination of risks to push a massive rocket into space, in preparation for a trip to orbit the Moon.

Such endeavors come with tremendous risks, which a world-class team works to minimize, but some residual aspects remain and are accepted.

Congratulations to the entire NASA team!

Cybersecurity Can Learn

The cybersecurity industry can learn many lessons from today’s Artemis II achievement.

Having strategic capabilities with clear objectives, resources, and accountability is key:

1. Prediction: Understanding the broad scope of risks, which are likely, and how best to manage them.

2. Prevention: Essential to mitigate the greatest risks that could lead to catastrophe.

3. Detection: Constant monitoring to identify problems as they arise and give the best opportunities to react in a timely manner.

4. Response: Well-rehearsed actions that intercept problems to minimize overall impact.

Lessons from each area create a feedback loop into the process, to make it stronger and more adaptive over time.

Establish and maintain an enduring cybersecurity strategy. Don’t rely on disconnected tactical efforts, as they will underperform over time.

If you need advisement, assistance is out there, reach out to industry leaders.

The post Cybersecurity Can Learn from the Artemis Launch appeared first on Security Boulevard.

❌