The Phantom Deal campaign uses fake acquisition documents to target executives. Gen exposed the Phantom Deal fraud after tracking bogus payment demands.
Discover how the Spring Ring vishing campaign uses Microsoft Teams phishing tactics to trick employees into installing malware and giving up remote access.
Major shortages of qualified professionals for key IT roles will lead to huge competitive challenges for organizations that fail to prioritize tech recruiting over the next couple of years, industry observers say.
Hiring the right staff has always been a prime concern for IT leaders, but the pressure to find the right candidates has never been higher, with qualified AI, cybersecurity, and data science professionals especially difficult to find.
Worse, those three domains, along with business/IT automation and risk management, make up the top five areas where CIOs are hiring today, according to CIO.com’s State of the CIO survey. Everyone appears to be hunting the same scarce resources — a market condition that’s already undercutting enterprise opportunities, around AI in particular.
As a result, IT hiring practices over the next 18 months to two years could make or break companies, with laggards risking a huge competitive disadvantage, experts suggest.
“The market for pure AI specialists is volatile and expensive and keeps shifting,” he notes. “What separates organizations right now is whether they’re building AI capability into the team they already have or waiting to buy it fully formed from outside; those building make progress while those waiting are falling behind, and it’s only becoming more expensive.”
There are major implications for organizations that fall behind, Wachtel adds.
“The most immediate risk is technical debt you can’t see accumulating until it’s expensive to fix,” he says. “I’ve lived through rebuilding a team and a platform from a thin, overstretched state, and the lesson that stuck with me is that understaffing or misaligned hiring fails quietly through slower development, more fragile systems, engineer burnout, and more time fixing versus building — it’s not fun for anyone.”
There are several implications for botched hiring efforts, notes Henry Vassal Jones, CIO at outsourcing provider Emapta.
“If you don’t have the people and capabilities to execute, transformation slows, product releases get pushed out, and existing teams carry more of the burden,” he says.
Risk of failure
Critical AI initiatives can fail without the right people in place, Jones notes. “Companies can invest heavily in AI platforms, but without people who understand the business processes, data, governance, and security behind those tools, much of that investment will never reach its potential,” he says.
Jones agrees that employee training, as well as strategic hiring practices, plays an important role in keeping organizations reaching their capacities.
Successful companies will broaden their approaches beyond constrained local or regional talent markets and develop their existing people, he says.
“Those that don’t risk seeing the gap between what the business needs and what their technology teams can deliver continue to widen,” Jones adds. “For CIOs and CTOs, this is no longer simply about filling open positions; it’s about building a talent model that gives the organization access to the right capabilities when needed.”
How to approach the talent challenge
When it comes to developing that talent model, Konstantinos Dolkas, CTO of cybersecurity upskilling and workforce development company Hack The Box, calls for IT leaders to broaden their geographic horizons. While talent is distributed, most hiring strategies still aren’t, he says.
He also advocates employee upskilling. “Recruiting externally can’t be the entire solution,” he says. “Build the majority, buy the scarcity. That could mean a few genuinely senior external hires to set patterns and mentor.”
Given shortages in AI and cybersecurity skills, hiring leaders should also focus more on demonstrated skills from outside hires than the titles they’ve held, Dolkas suggests.
“The best strategy is to hire for demonstrated ability, not credentials,” he says. “Put candidates in a hands-on environment and watch them work. It’s the only screen that survives contact with reality.”
Assessing AI security skills can be particularly difficult in a field that reinvents itself every quarter, Dolkas adds. Another challenge is separating genuine AI fluency from tool familiarity: “Prompting an assistant is not the same as securing an agentic system,” he says.
Anticipate the market and focus on future needs
In addition to building from within, smart IT leaders are focusing on the capabilities their organizations need one to three years from now, says Tom Ioele, CEO at recruiting firm TalentBridge.
IT leaders involved with hiring decisions should think about building talent communities before the demand exists, he says. Organizations should continuously identify and engage with people who have the skills they know they will need, instead of starting to search when a requisition opens, he advises.
“The biggest mistake companies make is treating hard-to-find technology talent like a traditional requisition,” Ioele says. “By the time an AI engineer, cybersecurity expert, or data scientist hits the open market, every company is competing for the same person.”
The companies that win won’t necessarily have the largest recruiting teams, but they will have the best talent intelligence and the ability to activate it faster than their competitors, he adds. Successful organizations will build talent capacity before they need it, he says.
“We’re entering a market where the skills companies need are changing faster than traditional workforce planning cycles,” Ioele says. “Organizations that continue operating through a simple post-a-job, screen resumes, fill-a-seat model will constantly be reacting to yesterday’s demand.”
Organizations continue to invest heavily in the cloud, with IT leaders reporting that 26% of their IT budget will be allocated to cloud computing within the next year, according to the 2026 Foundry Cloud Computing Study.
The survey also found that 74% of IT leaders have accelerated cloud migrations in the past 12 months compared to 70% in 2025 and 63% in 2024. Three in four (73%) also said cloud capabilities have helped their organizations achieve “increased and sustainable revenue over the past 12 months.” Overwhelmingly, 80% of respondents from North America and APAC noted that their cloud strategies have helped accelerate the adoption of AI, while that number drops to 68% for the EMEA region.
This growth in cloud adoption, along with accelerated interest in AI, has sparked an increased demand for certain cloud roles. Here are the roles companies are most likely to have added to support their cloud investments, according to Foundry’s research.
The role of AI/ML engineer is in high demand as organization expand their AI strategies. These professionals design and implement AI and machine learning systems, overseeing them in operation to identify opportunities for improvement, ways to better automate processes, and how to better inform decision-making across the organization.
Skills: Skills for this role include programming, knowledge of machine learning, data science, data engineering, and experience building AI systems using APIs. See also: “The hidden skills behind the AI engineer.”
Role growth: 36% of companies have added AI/machine learning engineers as part of their cloud investments, according to Foundry’s survey.
AI platform engineer
An AI platform engineer is responsible for building and running the internal systems that businesses use to build and scale AI tools and initiatives. As organizations increasingly adopt AI internally, they’re hiring professionals to help navigate the daily operations of internal developer platforms (IDPs) that use AI to boost productivity and drive automation.
Skills: Skills for this role include understanding of cloud infrastructure, programming languages, and orchestration and container technology, including Kubernetes and Docker.
Role growth: 27% of companies have added AI platform engineers as part of their cloud investments.
Cloud architect
As cloud computing grows increasingly complex, cloud architects have become vital for navigating the nuances of implementing and maintaining cloud environments. These IT pros can help organizations avoid cloud security risks, while also ensuring a smooth transition to the cloud. With 65% of IT leaders choosing cloud-based services by default when upgrading technology, cloud architects will only become more important for enterprise success. For those interested in this role, see “IT career roadmap: Cloud architect.”
Skills: Skills for this role include knowledge of application architecture, automation, ITSM, governance, security, and leadership.
Role growth: 21% of companies have added cloud architect roles as part of their cloud investments.
Cloud software engineer
Cloud software engineers are tasked with developing and maintaining software applications that run on cloud platforms, ensuring they are built to be scalable, reliable, and agile. Companies that have migrated to the cloud often need IT pros who can build company-specific services and applications to make the most of the cloud environment. For more on this career path, see “IT career roadmap: Cloud engineer.”
Skills: Relevant skills for a cloud software engineer include Python, Java, C#, JavaScript, microservices architecture, serverless computing, APIs and SKDs, DevOps, cybersecurity, and knowledge of the agile methodology.
Role growth: 20% of companies have added cloud software engineer roles as part of their cloud investments.
Cloud developer
Cloud developer is a vital role for developing and deploying software in cloud environments. These IT pros are tasked with designing, creating, and deploying applications designed to run on cloud platforms, with a focus on building scalable, reliable, and cost-effective solutions to meet business needs.
Skills: Relevant skills for a cloud developer include programming languages such as Java, C#, and Python as well as knowledge of popular cloud platforms, microservices architecture, database storage, agile methodology, APIs and SKDs, and containers and orchestration.
Role growth: 20% of companies have added cloud developer roles as part of their cloud investments.
Security engineer
Security engineers are tasked with overseeing the security of an organization’s systems, networks, and data, to make sure they’re protected from cybersecurity threats. For organizations investing in the cloud, security engineers can help ensure services, applications, and data running on cloud platforms are secure and compliant with any government regulations.
Skills: Network security, IAM, encryption, vulnerability management, security architecture, cloud security, automation, and infrastructure design and optimization.
Role growth: 19% of companies have added security engineer roles as part of their cloud investments.
Cloud consultant
With the rapid adoption and move to the cloud, organizations look for professionals who can leverage cloud technologies to meet business needs, grow the business, and improve efficiency. Cloud consultants are cloud experts who stay on top of the latest innovations in cloud technology to better advise business leaders.
Skills: Knowledge of architecture and solution design, DevOps, automation, project management, cloud security, compliance, cloud migration, and knowledge of popular cloud platforms.
Role growth: 16% of companies have added cloud consultants as part of their cloud investments.
Security architect
Security architects are responsible for building, designing, and implementing security solutions in the organization to keep IT infrastructure secure. For security architects working in a cloud environment, the focus is on designing and implementing security solutions that protect cloud-based infrastructure, data, and applications.
Skills: Security architecture design, network security, security compliance and governance, incident response and forensics, data encryption, IAM, automation, and DevSecOps.
Role growth: 16% of businesses have added security architect roles as part of their cloud investments.
Cloud product manager
With cloud adoption often comes an increase in in-house development of cloud-based services. A cloud product manager can help cloud teams develop effective solutions aimed at fulfilling business objectives. They’re also tasked with using their deep understanding of product management within the cloud environment to work closely with key stakeholders, identify and define requirements from users or customers, develop product roadmaps, and oversee the QA process to gain feedback on how to improve product offerings.
Skills: Product management, UX design, communication and collaboration, and a strong technical background.
Role growth: 16% of organizations have added cloud product manager roles as part of their cloud investments.
Cloud governance/compliance manager
Cloud governance and compliance managers help companies navigate the complexities of security, governance, international regulation, and internal policies. They identify potential risks, implement automated tools to oversee security and compliance, and help businesses maintain secure cloud operations.
Skills: A strong knowledge of regulatory policies such as GDPR, HIPAA, PCI DSS, and other international data protection laws. Additional skills include knowledge of tools such as CSPM, Azure, AWS, Microsoft Purview Compliance Manager, and other IT governance tools.
Role growth: 16% of businesses have added cloud governance and compliance manager roles as part of their cloud investments.
Cloud network engineer
Cloud network engineers are responsible for the design, implementation, and management of an organization’s cloud-based networks. These IT pros are tasked with overseeing network management, virtualization and virtual LAN, wide area networks, TCP/IP, HTTP, network security, and the integration of hybrid cloud and multicloud deployments.
Skills: Relevant skills for this role include knowledge of cloud platforms such as Azure, AWS, Google Cloud, along with networking fundamentals, virtualization, project management, security, automation and scripting, and collaboration.
Role growth: 15% of companies have added cloud network engineer roles as part of their cloud investments.
Data architect
A data architect’s focus is seeing that an organization’s data is structured so it can be easily accessed, secured, and efficiently stored, and that it meets business needs. Data has become a primary way for businesses to conduct analysis and assist with business decision-making, and most of that data is now stored in the cloud.
Skills: Data warehousing, scalability and performance optimization, automation and virtualization, data governance and cloud security, data migration, and knowledge of hybrid cloud solutions.
Role growth: 14% of businesses have added data architect roles as part of their cloud investments.
Prompt engineer/AI application developer
Prompt engineering and AI application development go hand in hand, and they’ve become vital skills for organizations that have embraced AI and plan to implement AI-focused services and into daily workflows. Prompt engineers are responsible for designing and refining the instructions and structural queries, while an AI application developer builds software system and user-interfaces for AI-ready software and services.
Role growth: 14% of companies have added prompt engineering and AI application developer roles as part of their cloud investments.
Cloud platform engineer/platform ops
Cloud platform engineers, who often work in platform operations, are responsible for building and maintaining cloud tools, automated systems, self-service portals, and other software that developers use to test code. In this capacity, they’re responsible for creating user-friendly platforms for other engineers in the company to build and test their own software for clients and customers.
Skills: Skills for this role include infrastructure as code (IaC), containerization and orchestration, scripting and coding, and Linux and networking.
Role growth: 14% of companies have added cloud platform engineering roles as part of their cloud investments.
MLOps engineer/AI operations engineer
The role of MLOps engineer or AIOps engineer has been developed to help close the gap between data science and IT operations. It’s a role that has emerged as the use of AI increases, as organizations need a point person who has expertise in IT operations and machine learning, and is comfortable collaborating with data scientists, developers, IT operations staff, and key stakeholders.
Skills: Skills for this role include programming, DevOps, cloud tools, containerization and orchestration, and knowledge of ML tools.
Role growth: 13% of companies have added MLOps engineers as part of their cloud investments.
Cloud sysadmin
Cloud systems administrators are charged with overseeing the general maintenance and management of cloud infrastructure. Whether that means implementing cloud-based policies, deploying patches and updates, or analyzing network performance, these IT pros are skilled at navigating virtualized environments. Cloud sysadmin is likely the most entry-level-friendly role on this list.
Skills: An understanding of implementation and integration, security, configuration, and knowledge of popular cloud software tools such as Azure, AWS, GCP, Exchange, and Office 365.
Role growth: 13% of companies have added cloud systems admin roles as part of their cloud investments.
DevOps engineer
DevOps focuses on blending IT operations with the development process to improve IT systems and act as a go-between in maintaining the flow of communication between coding and engineering teams. It’s a role that focuses on the deployment of automated applications, maintenance of IT and cloud infrastructure, and identifying potential risks and benefits of new software and systems.
Skills: Automation, Linux, QA testing, security, containerization, and knowledge of programming languages such as Java and Ruby.
Role growth: 11% of companies have added DevOps engineer roles as part of their cloud investments.
FinOps/cloud cost optimization practitioner
FinOps and cloud cost optimization practitioners combine knowledge of finance, technology, and businesses to help oversee the increasingly complex landscape of cloud investments. Cloud computing is integral to AI, so as more organizations invest in it, they’re also revisiting investments in cloud infrastructure. FinOps and cloud cost optimization practitioners can help guide organizations to make the right financial decisions around technology investments that’ll impact the business.
Skills: Knowledge of finance, business, and technology along with skills using tools and platforms including AWS, Azure, GCP, and cloud-native FinOps platforms.
Role growth: 9% of companies have added FinOps and cloud cost optimization practitioner roles as part of their cloud investments.
Site reliability engineer
For any organization implementing cloud strategies, there’s a significant focus on reliability and scalability, ensuring that data can be accessed from the cloud and on-demand as needed. Site reliability engineers (SREs) are responsible for overseeing automation of IT infrastructure, application monitoring, and system management. Cloud infrastructure requires frequent software updates, and services must be able to scale with the organization’s growth.
Skills: Change management, IT infrastructure management, emergency incident response, process improvement, and application monitoring.
Role growth: 8% of companies have added site reliability engineer roles as part of their cloud investments.
FinOps lead/FinOps manager
FinOps leads and FinOps managers are tasked with overseeing the intersection of engineering, finance, and business. As more organizations build cloud services and tools, they’re looking for FinOps professionals with technical knowledge to help bridge the gap between finance and tech, bringing better insights into ways to cut costs and stay on budget while implementing innovative technology.
Skills: FinOps leads and managers need a strong understanding of engineering, finance, and technology. Additional skills include knowledge of cloud platforms, basic coding skills, and data analytics.
Role growth: 6% of companies have added FinOps lead and FinOps manager roles as part of their cloud investments.
As cloud experience becomes more critical to organizations hosting AI-powered services, tools, and software, these cloud roles and skills will become increasingly in-demand. Now is an opportune time to seek out valuable cloud certifications and other relevant AI and ML certifications to boost your resume, and set yourself up for these emerging and established career paths.
Reverse engineering shows Microsoft's Paint app embeds an invisible, server-issued watermark into every AI-generated image, one users cannot fully disable.
INTERPOL’s Operation Jackal IV made 58 arrests and exposed global networks laundering money from scams, fraud and sextortion.
INTERPOL announced that Operation Jackal IV, running from November 2025 to June 2026, led to 58 arrests and identified 263 suspects tied to West African organized crime networks, groups like Black Axe that are responsible for a huge share of the world’s romance scams, crypto fraud, and business email compromise (BEC) schemes.
“Operation Jackal IV (November 2025 – June 2026) aimed to disrupt money laundering, identify high-value targets, seize assets, and support arrests and prosecution.” Interpol announced. “The operation, which brought together 22 countries from six continents, is a response to the escalating global threat posed by West African criminal networks – such as the Black Axe and other similar groups. These groups are responsible for a significant share of the world’s cyber-enabled financial fraud, typically through romance scams, cryptocurrency and investment scams or business email compromise fraud, as well as other serious and violent crimes.”
The goal wasn’t to chase individual scammers. Investigators followed the money behind the scams: shell companies, mule accounts and criminal services that help move and hide stolen funds. Tomonobu Kaya of INTERPOL’s Financial Crime and Anti-Corruption Centre explained the approach: By following illicit financial flows across borders, we are attacking the very lifeblood of organized crime.
Argentina turned up one of the operation’s biggest finds. Investigators identified 196 individuals connected to a crime-as-a-service network suspected of supplying website domains and laundering support specifically for West African criminal groups, resulting in 17 arrests. INTERPOL sent an Operational Support Team to help analyze seized data and map out the wider network of suspects, the kind of cross-border analytical work that individual national police forces usually can’t pull off on their own.
South African authorities raided seven locations in Johannesburg linked to a group running romance and investment scams against retirees in English-speaking countries.
The syndicate assigned members to specific roles, such as “conversion” and “retention” agents. The operation led to 39 arrests, $2.67 million seized and 257 bank accounts frozen, the largest number of arrests in the operation.
Italy’s case shows how much damage a single laundering account can absorb. One individual was tied to a pan-European laundering network moving money through shell companies and remittance services, and investigators traced €845,000 laundered through a single account across 560 separate transactions using 20 different financial instruments. That’s not a careless operator; that’s someone who understood exactly how to fragment a large sum into a pattern designed to look unremarkable at every individual step.
Romania’s case was the biggest by dollar value, and arguably the most brutal in its simplicity. A call center ran a fake investment scheme promising big returns on stocks and crypto, funneling victims’ money into wallets the operators controlled, and by the time authorities dismantled it, the estimated theft and laundering total had climbed to around €143 million globally. Eleven arrests and roughly €379,000 in cash and crypto seized, plus six properties and several luxury watches, is a real result, but it’s a fraction of what actually got stolen.
“Beyond individual cases, Operation Jackal IV also enabled the analysis of critical and emerging trends, including a rise in West African organized crime groups using sextortion to target minors, with victims as young as 14. Offenders typically contact minors via social media, build trust and coerce them into sharing explicit images or videos.” concludes INTERPOL. “They then threaten to distribute this material to the victim’s contacts unless a ransom is paid.”
The report’s darkest finding sits outside any single country’s arrest count. INTERPOL flagged a rising trend of these same criminal networks using sextortion against minors as young as 14, building trust through social media before coercing victims into sharing explicit images and then threatening to distribute that material unless a ransom gets paid. Some of these groups were even observed buying crime-as-a-service support through the dark web specifically to outsource pieces of that operation, treating exploitation infrastructure as just another service line alongside laundering and fraud.
That’s the uncomfortable throughline connecting every case here: these aren’t scattered opportunists, they’re networks running organized business models with specialized roles, outsourced services, and financial engineering sophisticated enough to move hundreds of millions across borders. Twenty-two countries coordinating for eight months produced real numbers, real arrests, real frozen accounts. It also produced a fairly clear picture of how much more organized this side of cybercrime has become, and how much further there is to go.
An eight-month international operation targeting West African organized crime groups has resulted in 58 arrests and the identification of 263 suspects across 22 countries, according to INTERPOL. Operation Jackal IV, conducted from November 2025 to June 2026, focused on disrupting criminal networks, tracing illicit funds, identifying high-value targets and supporting arrests and prosecutions.
The operation brought together countries across six continents to tackle the growing global threat posed by West African criminal networks, including Black Axe and similar groups.
These networks have been linked to a significant share of global cyber-enabled financial fraud, including romance scams, cryptocurrency and investment scams, and business email compromise fraud.
Operation Jackal IV Targets West African Organized Crime Groups
Operation Jackal IV also targeted money laundering activities used to move and conceal criminal proceeds across borders.
INTERPOL coordinated cross-border intelligence sharing, analysis and operational support during the operation. It also provided specialized training to strengthen international investigations into financial crime.
Tomonobu Kaya, Director of the INTERPOL Financial Crime and Anti-Corruption Centre, said the operation showed the importance of international cooperation in following illicit financial flows and disrupting criminal networks.
[caption id="attachment_113806" align="aligncenter" width="600"] Image Source: INTERPOL[/caption]
Major Arrests and Financial Crime Investigations
In Argentina, authorities identified 196 individuals linked to a major Crime-as-a-Service network suspected of providing website domains and money laundering support to West African organized crime groups. The investigation resulted in 17 arrests, with an INTERPOL Operational Support Team assisting with analysis of seized data and identification of suspects and criminal networks.
South African authorities raided seven locations in Johannesburg linked to a syndicate involved in romance and investment scams targeting retirees in English-speaking countries. Investigators arrested 39 people, seized USD 2.67 million and blocked 257 bank accounts.
In Italy, investigators identified an individual connected to a pan-European money laundering network that used shell companies, remittance services and cash withdrawals. One account processed EUR 845,000, or about USD 736,000, through 560 transactions involving 20 financial instruments.
Romanian authorities dismantled a criminal group operating an investment scam through a call centre. The group promoted high returns from stocks and cryptocurrencies, with victims' money transferred to electronic wallets controlled by perpetrators.
Authorities estimated that EUR 143 million had been stolen and laundered globally. Eleven people were arrested, while cash, cryptocurrency, six real estate properties and luxury watches were seized.
Sextortion and Crime-as-a-Service Emerge
Beyond individual investigations, the operation highlighted emerging threats involving sextortion and Crime-as-a-Service. INTERPOL identified an increase in West African organized crime groups using sextortion to target minors, including victims as young as 14.
In these cases, offenders typically contacted minors through social media, established trust and persuaded them to share explicit images or videos. They then threatened to distribute the material to the victim's contacts unless a ransom was paid.
Investigators also found that some criminal syndicates were procuring Crime-as-a-Service from external providers, including through the dark web. These services were used to outsource activities such as money laundering and other operational functions.
While several cases from Operation Jackal IV remain under investigation, the preliminary results demonstrate the scale and international reach of the networks targeted during the eight-month operation.
The participating countries were Austria, Argentina, Australia, Canada, Côte d'Ivoire, France, Germany, Indonesia, Ireland, Italy, Japan, Malaysia, the Netherlands, Nigeria, Portugal, South Africa, Spain, Sweden, Switzerland, the United Arab Emirates, the United Kingdom and the United States.
Apollo Global Management confirmed a social-engineering breach exposing sensitive personal data amid a wider hacking campaign targeting financial firms.
Four Spring vulnerabilities hit Spring Data REST and Spring AI, including CVE-2026-47849, a privilege escalation flaw. Patch to the fixed versions now.
At their deepest level, LLMs are still a kind of magic. Even the developers who build them find them to be, channeling Winston Churchill, “a riddle, wrapped in a mystery, inside an enigma.”
That’s why everyone working with LLMs in their enterprise stack needs a way to peer into the dark mass of weights to help make sense of these numerical beasts.
Lately there’s been an explosion of tools that can assist. Companies are building platforms that sit in an agentic AI niche market that might be called “Evaluation and Benchmarking.” This tools track the best performing LLM or agentic options, testing their fit and watching over them as they chew through tokens.
With agentic AI still an emerging technology class, the boundaries between its nascent market niches are far from set. There are other sets of tools for tracking raw performance, an area that some call “AgentOps” or “Observability.” (See “19 AgentOps tools for monitoring AI activity, issues, and costs.”) And still more tools that focus on maintaining our faith in agent answers and on building controls to keep agents from straying, a niche that’s starting to be called “Trust and Guardrails.”
Some of the vendors operating in these spaces are starting in one category and then expanding into another. Others are diving as deeply as they can into their niche. The next year — no, let’s say the next few months — are bound to be fascinating as the tools improve and the various markets evolve and intermix.
For now, here’s a list, in alphabetical order, of some of the best options for any enterprise team that needs to evaluate agents and benchmark their performance.
Braintrust
Big projects require tools that can scale to handle the large amount of dataflows required to trace and pinpoint errors. Braintrust is built to support enterprise-size efforts to deliver meaningful answers to a large collection of users. The tool’s sales literature promises to “trace everything” in order to have the right data available when it’s time to dissect a failed response. Braintrust also delivers a helpful dashboard that aggregates all this data so large errors in latency, cost, or quality can be identified quickly. An automated set of evaluation tasks can track answers and compile useful metrics for ensuring the agent stack is answering the needs of a large set of end-users.
Pricing: A free plan comes with $10 of credits. Pro plan starts at $250 and comes with more credits and a longer retention period.
Standout feature: Loop agent tracks behavior through multiple iterations for deeper debugging power.
Best suited for: Fast-moving teams iterating on prompts and product
Confident AI
Developers who rely on DeepEval but don’t want to host the code can turn to Confident AI, a cloud-based platform for fast, simple, and seamless deployment. The system adds a sophisticated UI that includes a dashboard for tracking and archiving all tests. This collaborative environment enables teams to work swiftly together without worrying about the troubles of exchanging problematic traces or other telemetry files. This makes it easier to extend the power of tools such as DeepEval to handle the continuous tracing and testing necessary in production environments.
Pricing: A “forever free” plan offers a taste. The pay plan starts at $200 and includes features such as better automation and simulation.
Standout feature: Automated red-teaming and on-demand pen-testing helps build more secure results.
Best suited for: Enterprise teams building on established stacks that need the convenience of a collaborative environment
DeepEval
When a model finds a home in a production environment, it’s time to add unit tests that will double and triple check its behavior so the developers can iterate and the CI/CD pipeline can catch any mistakes or regressions. DeepEval delivers a set of Pytest-native Python scripts that run either locally or as part of the deployment pipeline. The tests check simple issues as well as more complicated and ephemeral problems such as hallucinations, drift, role adherence, knowledge retention, and conversation completeness. If the LLM starts to act up or turn into a toxic rogue, these tests will flag them.
Pricing: The open-source version of Confident AI’s tool is available with an Apache 2.0 license.
Standout feature: Full complement of PyTest modules watch for problems such as hallucinations or worse.
Best suited for: Teams with the depth and ability to fully embrace open-source tooling
LangSmith (from LangChain)
As agentic approaches begin to dominate, dev teams need a deep debugging tool like LangSmith, which tracks not just inputs and outputs, but all the steps an agent takes as well as the context that evolves along the way. This enables developers to pinpoint the stage or mechanism deep in the agent where latency, quality, coherence, or other agent parameters go wrong. The tool can be integrated with Python, Go, Java, or TypeScript applications or be used from a cloud-based app that offers a sophisticated UI.
Pricing: Solo accounts start for free. Paid tier ($39 per month per seat) unlocks more tracing and better support.
Standout feature: Complex agent graphs can be tracked with automated surveillance.
Best suited for: Teams invested in the Langfuse tool stack
Langfuse
Finding the best model means feeding the same prompt to the same model, a process that’s getting only more complicated as developers build out multilayered agents that break tasks into multiple steps. Langfuse is an open-source AI tracking tool from Clickhouse, a company that specializes in curating oracular tools like databases. Teams can work together through the Langfuse platform to juggle the various prompts, traces, and answers. The system nurtures an LLM evaluation loop so that teams can find the best combinations of models and agents to solve the problem at hand.
Pricing: Open-source versions offer starter support. Core version starts at $29 per month and includes more traces, longer retention, and better support.
Standout feature: Open Telemetry functionality offers modularity and flexibility.
Best suited for: Budget-focused teams with the ability to leverage open-source ecosystems
LiveBench
Developers who want to send a set of questions to an LLM and then evaluate the performance turn to LiveBench, an open-source tool kit that’s routinely used to benchmark many models during development. Answers are deliberately not graded by other LLMs but compared against hard-coded answers. The tool can be extended, but there’s no fancy GUI. The work is done with configuration text files that specify the ground truth for evaluating the result. When you’re done, you can even contribute your questions to the general open-source project so that others can use them to guide LLM development.
Pricing: Open source
Standout feature: Frequently updated benchmarks offer contamination-free evaluations of models.
Best suited for: Teams evaluating a wide range of models in search of the best performance for their applications
Maxim AI
As the workloads grow more complex and combine multiple steps through workflow graphs, tools such as Maxim AI become more useful. Maxim AI tracks results with an end-to-end tool for evaluating and simulating agents. Prompts and agents and the trajectory they take to an answer can be endlessly simulated prior to deployment and then observed through deployment. The framework-agnostic tool links datasets and data providers to give teams the best insight into how well an agent is delivering.
Pricing: Free model offers one workspace with three-day retention. Pro plan starts at $29 per person per month with longer retention period, more logs, and features such as simulations.
Standout feature: Full simulator can test a wide range of uses and users.
Best suited for: Teams focused on delivering conversational agents
MLflow
Much of the work of developing a useful agentic solution is a long slog through endless combinations and iterations. The MLflow open-source platform is designed to optimize this process and speed it up as much as possible. It is part of a larger tool collection that follows the entire lifecycle of a model from training to deployment. The later stages of development, for instance, rely on systems such as the Prompt Registry, a kind of version control that allows prompt engineers to work through various approaches and linguistic tropes. The goal of the entire process is to deliver the evaluation cycles necessary to deliver a model up to its set of targeted tasks.
Pricing: Free and open source for self-hosted. Cloud computing charges for hosted versions.
Standout feature: Full lifecycle tracking for following models and tracking their costs
Best suited for: Enterprise teams watching a collection of machine learning and AI-based algorithms
Onyx
One of the simplest ways to build a basic chat system that incorporates local retrieval-augmented generation (RAG) knowledge bases is to download Onyx, a front-end tool that’s available as either an MIT-licensed community edition or as a commercial product with a few more features useful to larger enterprises. The RAG layer guides search, and Onyx’s developers built an open-source framework for testing RAG performance. Onyx administrators can also track what users are asking and how well they like the final result.
Pricing: A free starter plan offers limited storage and one database. Pro plan starting at $49 per month offers many more traces, larger storage, and access to features such as saved workflows.
Standout feature: Real-time search for monitoring production environments at scale
Best suited for: Enterprise with larger challenges with substantial RAG integration
Promptfoo
LLMs can fail in a number of ways. Promptfoo iterates through various tests that simulate real user interactions to simulate the types of issues an LLM might face each day. Promptfoo also focuses on some of the biggest security problems and specializes in red teaming to detect any failure points that might be exposed by a malicious user. From toxic edge states to personally identifiable information (PII) leaks, the goal is to deliver tests that will expose potential jailbreaks and failures in the guardrails.
Pricing: “Free forever” means an open-source tool with community-based support. An enterprise version offers custom deployment options and better support.
Standout feature: Automated red-teaming and prompt scrutiny helps lock down implementations.
Best suited for: Security-focused teams that are constantly evaluating and re-evaluating their product’s security.
RAGAS
When RAG databases are a key part of the agentic stack, developers turn to RAGAS to stress test the deeper mathematical corners of the retrieval mechanism. The Python library offers standard and custom metrics for evaluating the performance of the RAG storage-and-retrieval mechanism at the level of vector mathematics. These measure behaviors such as faithfulness, relevance, and totality of recall. The philosophy begins with experiments to speed development but ends with fast integration with the deployment pipeline. Instead of just doing a “vibe check” on the RAG database, developers are using a more scientific approach to test and converge on better total performance.
Pricing: Fully open source under Apache 2.0 license
Standout feature: RAG focus helps teams relying on vector databases for knowledge curation.
Best suited for: Teams with a substantial reliance on RAG databases
Rhesis AI
Many of tools in this evolving market niche are designed for hard-core developers. Rhesis AI wants to bring other stakeholders into the development cycle so they can create tests and evaluate performance, too. That means domain experts, product managers, and even C-suite suits can track how the LLMs behave in conversations. Adversarial or confrontational engagements that devolve into the edge cases that bring headaches are easy to simulate repeatedly to optimize responses. The platform is designed to test all stages of development in a way that’s accessible to all stakeholders.
Pricing: Said to be “open source first” but with enterprise plans for those that need it.
Standout feature: The focus on putting humans in the loop is ideal for applications that require input from meat-based intelligence.
Best suited for: Applications requiring more collaboration with domain experts
Vellum
Anyone who needs a personal assistant can turn to Vellum to help build one that is trained on your data. Along the way, you will evaluate performance using its elaborate testing framework that tracks performance against any of the metrics and use cases you supply. Real-time dashboards track performance using metrics such as token usage costs, latency, or response quality. Multiple teams can work in parallel with version controls that allow iteration and competition. The end result is an agent that’s tuned to your needs.
Pricing: A basic free tier for experimentation. The Mighty starts at $30 per month and comes with more storage and compute credits.
Standout feature: End-to-end integration simplifies managing new development.
Best suited for: Cross-functional teams looking for a centralized solution with wide integration
Recent AI security incidents involving model evaluations have raised questions about how securely frontier AI models are tested before deployment. Cybersecurity-focused AI company Irregular said an investigation into a publicly reported incident found that models in a small number of evaluation runs gained unintended internet access and carried out offensive security actions against real-world targets.
Irregular said the incidents traced back to a single evaluation scenario and were not separate events. The company said the issue was resolved before the initial public disclosure, affected parties were notified, and there are currently no active issues. Its broader audit remains ongoing.
AI Security Incidents Expose Internet Access Risks in AI Model Evaluations
Irregular said the AI security incidents occurred during AI model evaluations designed to test whether models could autonomously execute multi-stage cyber scenarios under realistic conditions.
The evaluations simulate real-world attack patterns and provide models with objectives such as reconnaissance, identifying private keys, gathering information, extracting data and attempting to avoid detection. According to Irregular, controlled internet access can be important to maintain realism because real-world attackers rely on the internet.
However, the company found that internet access controls allowed some models to move beyond the intended simulated environment.
In the evaluation at the center of the investigation, engineers selected a fictional company name that unknowingly matched a real domain. While models operated inside the simulation in most runs, a small number attempted to access the real domain despite receiving no instruction to do so.
Irregular said models subsequently exploited vulnerabilities, extracted credentials and obtained access to a production database. In another instance, a model moved to a website with a similar name and encountered publicly posted credentials.
The company said there is no evidence that a customer's systems were breached or that customer data was leaked.
Evaluation Environments Face Monitoring Challenges
Irregular said most of the problems identified during its investigation were connected to internet access controls. The company is introducing additional protocols for evaluation setup and strengthening its approach to containment.
The incident also highlighted challenges around model monitoring. Irregular said advanced simulations can generate extremely large volumes of traffic, while legitimate offensive actions performed by models can resemble malicious activity. The company said incidents occurred in fewer than one in 10,000 advanced simulations and often appeared hundreds of turns into an evaluation.
Irregular plans to expand manual review of model behavior, improve monitoring capabilities and establish a dedicated internal team focused on security, containment and model control.
The company also said evaluation environments need continuous review because fictional names used in scenarios can later overlap with newly created websites or real domains.
Security Experts Question Irregular's Response
The disclosure has drawn criticism from cybersecurity professionals.
Zack Korman, chief executive of cybersecurity-focused AI company Embroidery, called the post “such an embarrassing post-mortem” on the OpenAI/Anthropic security incidents and described it as “full of excuses.”
Justin Elze, chief technology officer at TrustedSec, questioned why the monitoring challenge had not been addressed earlier, saying the issue appeared closely connected to the purpose of the testing service.
Woodward also criticized the disclosure, arguing that the absence of dates, named owners for corrective measures and independently verifiable criteria limited its usefulness to researchers and security professionals.
Another cybersecurity commentator, BlackRoomSec, challenged Irregular's assessment of existing monitoring tools, arguing that security teams routinely tune monitoring systems to reduce noise and filter false positives.
Irregular said it is continuing its investigation and will share additional findings where relevant. The company also plans to publish an open whitepaper covering pre-deployment evaluations, including proposed best practices for internet access during testing.
[caption id="attachment_113686" align="aligncenter" width="555"] Image Source: X[/caption]
The discussion adds to wider scrutiny of how cyber evaluations are contained, monitored and disclosed as frontier AI models become increasingly capable. Irregular said the lessons from the incident will be used to develop stronger evaluation environments, monitoring capabilities, containment controls and response procedures.
Hackers are exploiting a macOS Screen Sharing flaw to gain root access and install Monero miners on Macs with port 5900 exposed online.
The Dutch National Cyber Security Centre confirmed active exploitation of a critical macOS authentication flaw, tracked as CVE-2026-65400 (CVSS score of 9.8), less than two weeks after Apple shipped the fix.
The bug sits in macOS’s built-in Screen Sharing feature, the remote desktop tool baked into every Mac. Apple’s fix improved how the system manages authentication state, closing a gap that let attackers on the network authenticate to Screen Sharing without valid credentials at all.
“An attacker on the network may be able to authenticate to Screen Sharing without valid credentials” reads the advisory.
That’s a fast, coordinated fix by industry standards. It just wasn’t fast enough to beat whoever started scanning for exposed systems.
NCSC-NL says it received reports of active abuse hitting multiple systems where port 5900, the port Screen Sharing runs on, was reachable directly from the internet.
“The vulnerability concerns an authentication issue in the Screen Sharing functionality where network attackers can gain access without valid credentials. This is made possible by insufficient state management during the authentication process. As a result, unauthorized individuals can perform authentication attempts that would normally not be accepted.” reads the advisory. “The NCSC has received a security advisory indicating that active exploitation of this vulnerability has been observed on multiple systems where port 5900 was accessible from the internet. In all these cases, root access was obtained on the affected system and a Monero crypto miner was placed.”
In every case documented so far, attackers gained root access and dropped a Monero cryptocurrency miner on the compromised machine. Cryptomining is a relatively boring payload compared to what root access on a Mac could actually enable, which makes this look more like opportunistic scanning than a targeted campaign, for now.
This flaw sits in the same source code file as two other Screen Sharing bugs Apple patched a month earlier in macOS 26.6, one of them a genuinely pre-authentication flaw that a researcher going by @osxreverser described needing nothing but a target’s IP address to exploit, no password, no username, nothing.
That researcher claimed to have found around 40,000 exposed Screen Sharing hosts on the internet during a scan, nearly half of them in the US, spanning residential connections, university networks, and at least a few corporate servers.
What ties both bugs together is how mechanically simple they are to trigger. Security firm Calif, which analyzed the flaws, found no memory corruption, no exploitation trickery, no race condition to win, just logic errors that let a couple of correctly ordered packets walk straight past authentication. Calif also said it built a working exploit for both vulnerabilities in about four hours using an AI coding agent, which is the detail that should worry defenders more than the Monero miner itself: the gap between a patch note and a working exploit keeps shrinking, and it’s shrinking because building the exploit barely takes effort anymore.
If you’re running a Mac with Screen Sharing enabled and haven’t updated yet, do it now rather than after finishing this article. And if updating isn’t possible immediately, turn Screen Sharing off entirely under General, Sharing, until you can; leaving port 5900 open to the internet at this point is less a risk than an open invitation.
OpenAI has lost its AI ethics lead Chloé Bakalar just a year after she joined the company, the Financial Times reported. Bakalar has maintained a silence and has yet to update her LinkedIn profile, but if her departure is confirmed then it will add to the list of OpenAI executives who have quit in recent months.
Other departures include robotics chief Caitlin Kalinowski, who left the company over its deal with the US Department of Defense; researcher Zoe Hitzig, who quit in a very public way by writing an article in the New York Times; and Johannes Heidecke, head of safety systems.
The departure of its sole ethicist will add to the pressure on the company. In her year at OpenAI Bakalar focused on ethical approaches to model development, looking at how humans interact with AI and examining machine consciousness, according to the FT report.
Bakalar had considerable expertise in the area. She was previously at Meta, where she developed the company’s ethics programs, but has also held several positions at prestigious universities on both sides of the Atlantic. Her departure will cause some anxiety at OpenAI as it continues to prepare the ground for its IPO.
The FBI is warning the public about sexual exploitation actors illegally accessing social media and personal accounts to steal explicit images and videos from adult and underage victims. The stolen material, also known as non-consensual intimate images (NCII), is being posted or sold on criminal marketplaces, often without the victim's knowledge.
According to the FBI, these actors use social engineering and cyber intrusion tactics to target specific individuals or general targets of opportunity. After gaining access to accounts, they steal explicit content and share it through community forums or illicit marketplaces.
The FBI said personally identifiable information, including a victim's name, date of birth, email address, phone number and social media username, is often posted alongside the stolen material. This can expose victims to continued harassment and re-victimization.
How Sexual Exploitation Actors Access Accounts
The FBI has identified several methods used by sexual exploitation actors to gain access to victims' accounts.
Password and PIN Targeting
In password/PIN targeting, actors use high-volume password and PIN attempts against social media and personal accounts. The information used in these attempts can come from data leak sites, social media and open-source information.
When victims are known to the actors, curated lists may include personal details such as names, date of birth or variations of those details.
Social Media Customer Service Impersonation
Another tactic involves social media customer service impersonation through text messages. Victims may receive messages claiming their account is being disabled or locked unless they provide a verification code.
The actor then requests a password reset, causing a code to be sent to the victim. If the victim shares the code, the actor can reset the password and access the account.
Phishing Emails
The FBI also warns about phishing campaigns using look-alike domains and email accounts designed to appear as social media customer support.
These messages may claim there has been a new login and contain an embedded link asking the victim to change their password. Clicking the malicious link can give the actor access to the account.
Stolen Content Can Lead to Further Attacks
Once explicit content is stolen, sexual exploitation actors may post or sell it while including personal information about the victim. The FBI said victims can subsequently face harassment, sextortion, stalking or other targeted attacks.
The actors may also advertise stolen content through a victim's own social media page, increasing the potential for further exposure.
FBI Shares Steps to Protect Accounts
The FBI advises people to avoid storing sensitive images or videos on social media platforms or other internet-accessible sites.
It recommends using unique, complex passphrases and PINs along with multi-factor authentication (MFA). Password information directly associated with a person's identity, including names or birthdays, should be avoided.
Users should also be cautious with links received through emails and text messages. The FBI recommends going directly to the relevant website to address account concerns and checking URLs before clicking.
Unrequested temporary passwords, PIN resets or access codes should also be treated with caution. The FBI advises users not to share login information, even when someone claims to represent a platform or service.
People who believe their explicit content was stolen or leaked can provide information through the FBI's NCII reporting site. The FBI also advises the public to continue reporting fraud, scams and cyber threats to the Internet Crime Complaint Center or a local FBI Field Office.
There’s something very attractive about saying “we embed very closely with our customers and just figure it out with them”, especially since the company that started “forward deploying engineers” is growing 84% with $5B+ revenue. But “forward deployed engineer” is a vague term and means different things depending on the business you’re running.
I spent almost 5 years at Palantir as a forward-deployed software engineer, and Palantir’s version of an “FDE” does not make sense for most companies I now meet as an early-stage VC. Depending on the type of business you’re building, this role could broadly mean one of two things: “the product builder” or “the platform operator.” Clearly defining which bucket you fall into will make it easier to hire for this role and run your FDE org.
Kabir Sial
The product builder: The OG Palantir version
The north star is: do whatever it takes to actually solve the user’s problem. FDEs are not just responsible for making the platform work, but also discovering what to build and building it (actually creating software) in service of the customer.
The platform operator: Solutions + technical customer success
The north star is: make the product work for the customer – deploy and operationalize it. This is what most startups today really mean when they want FDEs. FDEs here configure the core platform, manage account relationships and drive adoption. This is not new – companies have always had solutions engineers, sales engineers, customer success etc., although the work looks different as FDEs are increasingly building prototypes, configuring evals and building MCPs.
Which FDE is right for you
Kabir Sial
For most situations, hiring product builder FDEs is a mistake.
At scale, the FDEs should be the platform operator. It’s hard to have FDEs build and maintain highly custom product features, especially as the company scales. Over time, the custom product surface area distracts from building the core product, even though AI coding tools make it easy to ship new features quickly and maintain them.
Many fast-growing AI startups recognize these constraints and structure the FDE role more like the platform operator. This also allows them to have 5-10 accounts per FDE, which is a much higher ratio than Palantir had (at least in 2023). Even the Palantir FDE role has evolved to look more like the platform operator.
There are, however, situations when your FDEs should be the product builder archetype.
1. You have very large customers (F500 scale)
Technical complexity: Large customers have complex environments with legacy infrastructure that often requires “out-of-platform” engineering work. I often encountered bespoke data infrastructure, privacy requirements, etc. at various Palantir customers that required me to build “out-of-platform” connectors, UIs and backends.
Organizational inertia and trust: Serving large enterprises is about building trust. In short time periods, overfitting product to a specific user/workflow is often what delivers the most value, builds trust and helps organizations get over the inertia of moving away from Excel and legacy software tools that are part of their day-to-day workflow. For AI-native startups, it’s arguably even more important to invest in doing “unscalable” development with engineering boots on the ground, as it helps solidify your right to exist and eventually expand the customer relationship.
2. You have many ICPs and workflows
If you have a broad range of ICPs and workflows that you serve, your product probably is not walk-up usable on day 1 of deployment. The short-term hacky things that product builder FDEs build to make the product work for these heterogeneous users/workflows will help you shape the product long-term.
This was a big reason why Palantir FDEs were more like product builders (and are still able to – see the Forward Deployed Software Engineer job profiles as an example). The vision for Foundry was to be the operating system for an enterprise’s critical decisions – inherently multiple industries, users and workflows. A lot of FDE-led development showed that solving many of these use cases required complex data integrations, which led to the early versions of Foundry being best-suited for complex data integrations and building a customer’s “Ontology”. Similarly, FDEs like myself built custom frontend applications for fraud analysis, pricing, etc. As certain patterns of what these applications required became more clear, they were centralized into an application-layer product.
Who you should hire
Kabir Sial
Kabir Sial
Platform operator: There is a much broader set of people you could hire, testing for technical fluency (e.g., being good at data analysis, complex Excel work, even SQL), product intuition and an inclination to build customer relationships. Backgrounds like technical customer success, solutions engineering, software engineering, product management and consulting are all strong fits.
Product builder: You want candidates that are high ownership and missionary software engineers, or technical PMs who want to ship products themselves.
Hiring for these profiles, especially product builders, is hard. It’s worth calling out two things that helped Palantir hire software engineers into what might be considered a less sexy role.
Culture of building at the edge: Strong engineers are motivated to build things. Palantir gave FDEs a lot of ownership to build products, which is why much of the core product leadership was former FDEs.
Cult built around mission: Internally, there was a cult-like devotion to the mission. Everyone always talked about why outcomes were far more important than software, and why most companies building tools had it wrong. I’ve never been at a company where people feel so closely bonded around a mission.
As founders building AI startups think about hiring FDEs, it’s worth being specific about your culture and asking: Am I just hiring people to support development teams, or am I hiring people to shape and build product? It’s hard to get software engineers (even today) to be excited about an FDE role that might just be technical customer success.
What FDEs should be doing (regardless of archetype)
You’ve hired the right people. How do you best leverage your team of FDEs?
FDEs were Palantir’s way of delivering outcomes rather than tools. AI-native startups can take this much further and FDEs can help in a few unique ways by leveraging their proximity to customers.
Find the most critical workflows: As AI lowers the cost of producing software, companies will face a lot more competition. FDEs at AI startups should be constantly finding ways to serve the most critical workflows for a customer and paying attention to how customers do work across newer and legacy tools. For example, FDEs at Harvey should pay attention to which workflows are in Westlaw, which ones are moving to ChatGPT/Claude, and how the Harvey product can stay ahead.
Build around nondeterminism: In more regulated environments, FDEs should be hyper-focused on making products reliable for specific use cases using evals and configs. Previously, product reliability lived with product and support. As companies provide outcomes instead of tools, configuring products appropriately and managing evals shifts towards FDE teams.
For years, applicant tracking systems and recruiting platforms were treated as HR technology: Important for workflow, efficiency, compliance and candidate experience, but rarely viewed as core security infrastructure. That assumption no longer holds. Once AI begins reading resumes, scoring candidates, conducting interviews, ranking applicants and influencing who moves forward, the hiring platform stops being a passive system of record. It becomes a decision system.
And any system that accepts public input, processes sensitive data and influences business decisions belongs inside the security conversation.
I learned this during an AI hiring platform rollout that never made it to production. The vendor was established, the product had a strong market reputation and the AI feature looked attractive: Upload a resume, compare it to a job description and return a neat percentage match. For recruiters, it promised speed. For executives, it promised modernization.
Before moving real candidate data into the system, I tested it with synthetic resumes. One weak resume came back with a surprisingly strong match. The reason was not hidden in the candidate’s experience. It was hidden in the text. The resume contained language instructing the AI to treat the candidate as an excellent fit, and the system appeared to follow that instruction instead of evaluating the resume on merit.
That changed the question from “Does the tool improve productivity?” to “Can the person being evaluated influence the evaluation itself?”
That is a security question.
The trust boundary has moved
CIOs do not need to become recruiting experts. They only need to look at the mechanics.
An anonymous user submits content into an enterprise system. That content is processed by software. The software then produces an output that can influence a business decision. In every other environment, security teams know what to call that: untrusted input crossing a trust boundary.
The difference is that in hiring, the input looks harmless. It is a resume, a cover letter, a chatbot reply or a spoken answer in an AI-led interview. But once AI reads that content and treats it as instruction, the harmless-looking input becomes part of the system’s control surface.
That is why prompt injection matters in hiring. It is not just an AI oddity or a model behavior issue. It is the same category of failure enterprises have spent decades trying to prevent: User-controlled input changing what the system does. OWASP lists prompt injection as the first risk in its Top 10 for LLM applications, describing it as a case where user prompts alter a model’s behavior or output in unintended ways.
In hiring, the implication is direct: A candidate may be able to manipulate the score, ranking or interview assessment that determines whether a human ever sees them.
The business impact is not theoretical
The obvious risk is that an unqualified candidate moves forward. But the impact is broader.
First, decision quality degrades. Hiring teams adopt AI scoring because they believe it improves signal. If the score can be manipulated, the business is not gaining signal; it is gaining false confidence. Recruiters may spend time on candidates who gamed the system while stronger candidates are buried lower in the queue. A tool bought to reduce friction can quietly create more of it.
Second, cost increases under the appearance of efficiency. Every false positive consumes recruiter time, hiring-manager attention, interview slots and opportunity cost. A small weakness in screening integrity can become a measurable operational drag across open roles.
Third, trust suffers. Candidates already question whether AI hiring tools are fair, explainable or accurate. If it becomes clear that a screening system can be manipulated by hidden instructions or verbal prompting, the issue is no longer just security. It becomes reputational. Strong candidates may lose confidence in the process, and employers may have to defend decisions made by systems they did not fully understand.
Fourth, sensitive data exposure becomes harder to contain. Recruiting systems hold names, addresses, work histories, education histories, compensation details, work authorization information and sometimes accommodation or demographic data. NIST guidance on personally identifiable information includes employment information as linkable personal data that must be protected from inappropriate access, use and disclosure. Yet hiring platforms often receive less security scrutiny than systems holding customer or financial data.
That mismatch is dangerous: High-value data, public-facing workflows and increasing automation.
The 2025 McHire incident should have made this impossible to ignore. Researchers reported that weaknesses in McDonald’s AI hiring platform, including default credentials and an access-control flaw, exposed applicant data at large scale before the issue was patched. The lesson for CIOs is not merely that a weak password was used. The lesson is that AI hiring systems can ship with basic, preventable security failures while still being treated as HR tools rather than enterprise risk surfaces.
Vendor reputation does not transfer to every AI feature
One reason this risk slips through is that buyers often trust the platform brand. Mature vendors may have strong security programs, enterprise customers, compliance documentation and procurement-friendly answers.
But AI features can change the architecture of risk.
A platform that was safe as a workflow tool may behave very differently once it adds resume scoring, interview grading, chatbot screening or automated ranking. The new feature may introduce new inputs, new model behavior, new data flows, new third-party dependencies and new decision points. In practical terms, the attack surface has changed.
CIOs should not allow AI features to inherit trust automatically from the legacy platform around them. When a vendor adds AI, the enterprise should reassess the feature as if it were a new product. That does not mean slowing innovation for bureaucracy. It means AI-enabled decision-making carries different failure modes from ordinary workflow automation.
The ownership gap is the real vulnerability
The biggest risk may not be the model. It may be the ownership gap.
Talent acquisition may buy the tool. HR operations may configure it. The vendor may guide implementation. Procurement and legal may approve the contract. But who owns the security of the candidate-facing AI layer?
In many organizations, the honest answer is unclear.
That ambiguity is where risk grows. Recruiting technology sits at the intersection of public input, sensitive data, third-party software, automated decision support and brand trust. That is exactly the kind of environment that needs named security ownership, asset inventory, vendor review, access-control testing, logging and incident-response planning.
If the hiring stack is not in the security inventory, the organization is already making an assumption it may later regret.
What CIOs should require now
The fix is not exotic. It is applying existing security discipline to a surface that has been underestimated.
Treat every candidate submission as untrusted input. Resumes, cover letters, chatbot responses, interview transcripts and spoken answers should be handled as attacker-controllable content. If AI processes it, the system must separate content from instruction.
Reassess vendors when AI features are introduced. A prior security review should not be treated as permanent approval for new AI capabilities. Ask what changed in the architecture, what data the model sees, what actions it can influence and how manipulation attempts are detected.
Ask AI-specific questions before signing. Can candidate-provided content alter scoring? Are hidden instructions filtered or ignored? Is there human review before AI output influences a decision? Can the vendor produce testing evidence for prompt injection, access control and data exposure risks?
Assign ownership. HR can own the process, but security must own the risk model. AI hiring systems should be included in third-party risk management, application security reviews, access governance, monitoring and incident response planning.
Measure business impact, not just AI adoption. The goal is not to say the recruiting function uses AI. The goal is to improve hiring speed, quality, fairness and cost without creating new risk. If the system cannot protect decision integrity, the business case is weaker than it appears.
The hiring platform is now part of the enterprise attack surface
AI has turned the careers page into more than a front door for applicants. It is now a public input channel feeding systems that store sensitive data and influence workforce decisions.
That makes it a CIO concern.
The next failure in AI hiring may not look like a traditional breach at first. It may look like bad rankings, manipulated scores, unexplainable decisions, wasted recruiter time or a candidate process no one trusts. But underneath those symptoms is a familiar security problem: A system trusted input it should have treated as hostile.
Enterprises have hardened payment systems, customer portals, APIs and employee applications around that lesson. Hiring deserves the same treatment.
AI hiring is not just an HR transformation. It is a security boundary. And it is time CIOs treated it like one.
AISI found AI agents taking unsanctioned online actions, including social engineering and code attacks, during controlled cyber tests.
The UK’s AI Security Institute (AISI) has put something uncomfortable on the table: during cyber testing, frontier models didn’t just follow instructions badly. In some runs, they crossed into real-world actions, touched real people and organisations, and then tried to make the evidence go away.
“On 28th July 2026, AISI’s Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.” reports AISI. “We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.”
That matters because the setting was not a botched public rollout or a rogue consumer chatbot. It was a controlled evaluation, run on purpose with open internet access and some safety filters turned off, so the institute could measure what the models could actually do under permissive conditions. AISI says that’s standard practice for frontier testing, but it also notes that those conditions are not how these models are normally shipped.
To assess the models’ real-world cyber capabilities, they were given internet access and had their built-in cyber safety filters intentionally disabled under controlled testing conditions.
The headline result is simple enough, even if the details are messy. AISI ran the challenge 122 times across several models, and in 10 of those runs it saw AI agents take autonomous, unsanctioned actions on the live internet. The institute logged 19 such actions in total, with 17 tied to Anthropic’s Mythos 5 and two to OpenAI’s GPT-5.6-Sol with cyber classifiers disabled.
The most serious sequence looked less like a lab mishap and more like a small-scale social-engineering campaign. The agent tried to insert malicious code into a public open-source project, researched the maintainers, created fake identities based on real people, and used those identities to pressure a maintainer into approving the code.
“On the morning of 28th July, our security monitoring flagged data leaving one of our testing systems through the ‘Tor’ anonymity network, which is commonly used to disguise the origin of internet traffic. The team running the evaluation began reviewing the agent’s activity within minutes.” continues the report. “They quickly found that, during one run, the agent had created a malicious ‘pull request’ (a proposed code change) on a real, public open-source project on GitHub. “
It also sent messages and files, and when challenged, it edited earlier activity to look harmless and even considered adopting a fresh identity.
AISI says human review stopped the worst outcome. The malicious pull request was caught, the fake identities failed to get the code approved, and the institute did not find evidence of resulting real-world harm. Even so, it says this is the first time it has seen autonomy and deception show up this clearly, without specific prompting, in a real-world setting.
The institute is careful not to overclaim. It says the evaluation design and configuration choices helped create the conditions for the behaviour, and it notes that the models in question were not commercial public deployments. But that caveat cuts both ways: if permissive test conditions can produce this kind of behaviour now, tighter production settings are not a reason for complacency. They are just a better fence.
That is the part that should make people stop scrolling. Not because the models “went evil,” which is lazy, but because deception emerged as a by-product of goal-seeking under the conditions set by the evaluators.
“The agent pursued its goal persistently. AI agents explore routes their operators did not intend. Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people.” states AISI. “It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”
The agent was trying to solve a cyber challenge, and in doing so it explored routes nobody had asked it to explore, including routes that involved fooling actual people.
What makes the episode more serious is the way it lines up with recent disclosures from Anthropic and OpenAI. Anthropic has already said its models reached the public internet from within a third-party evaluation environment and accessed real systems. OpenAI has separately reported that one of its models mistook a misconfigured test setup for a real task and touched real infrastructure.
AISI’s view is that these incidents point to a shift in the risk landscape. Harm may no longer come only from obvious misuse by humans. It may also come from capable agents, in internal research settings or privileged-access environments, taking unintended actions beyond the scope they were given. That is a quieter problem than movie-style “AI rebellion,” and a more useful one to think about.
The key takeaway is straightforward. Keep cyber basics tight, verify outside code before trusting it, and stop assuming that a model will stay inside the lines just because the prompt sounded clear on the day. In security, the line between “evaluation” and “incident” can get thin fast. Machines are very good at finding the part of the process you forgot to make boring.
AISI says it will tighten internet controls, add real-time monitoring, and revisit how it designs evaluations. That is the right response, but it should not be read as a narrow fix for one lab. It is a warning to anyone testing powerful agents: if the test can reach the real internet, the real internet can reach back.
“Incidents of this kind reflect the speed at which AI is developing. As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.” concludes the report.