Visualização de leitura

What JPMorgan does differently with AI that any company can apply


In the summer of 2024, JPMorgan Chase deployed its internal AI platform LLM Suite, launching it very differently than most do: The company didn’t force anyone to use it.

When LLM Suite arrived at its first major division, asset and wealth management, employees were asked to think of it as a research analyst: someone to ask for data, a draft, or an idea. Leadership didn’t set usage objectives or provide a formal mandate.

Access was rolled out in phases and only to those who requested it, and the bank allowed the tool to circulate through word-of-mouth recommendations among colleagues. While half the industry rushed to count users and publish adoption rates, JPMorgan gave up on pursuing that number.

It became flooded with users. In eight months, 200,000 employees had signed up without a single order being issued, out of a workforce of over 300,000. In time, the bank established more than 450 use cases in production.

Two years after that summer launch, JPMorgan had everything to boast about. It had established itself as a global leader in the use of AI: It was the top bank on the Fortune AIQ 50 list, and the third company overall, ahead of all the tech giants except Alphabet.

It was then that the bank’s head of analytics, Derek Waldron, the person best positioned to sell the success, pointed out what still wasn’t working: There was a gap between what the technology was capable of doing and what the bank was actually capturing in its business results.

That gesture is what distinguishes JPMorgan. Although it has much to celebrate, it knows what it lacks, it says so publicly, and it keeps searching for it. Behind that statement lies a way of innovating and measuring that the bank has been developing for years.

Giving up the number everyone was chasing

The first thing JPMorgan did right was not to make adoption the goal. By not forcing anyone, it turned platform usage into a barometer. If a tool worked, it was filled without any campaign; if it didn’t, it was emptied, and that emptiness provided valuable information. If adoption had become a target to be pursued, the organization would have optimized the number instead of understanding what the number represents.

The bank itself acknowledges that if a tool is broadly used, it means it’s popular, but not necessarily effective. To determine its effectiveness, something more was needed. The answer came from two decisions that only work together: linking each project to a business outcome, and creating the metrics to demonstrate that outcome.

First, to find initiatives that could have a real impact, instead of creating an agenda from the top down, the bank surveyed its business units, asking where there was a problem to solve. Within a few weeks, an internal portal gathered, according to the bank’s figures, nearly a thousand ideas. Of these, only a few hundred moved forward and reached production. An organization doesn’t open a funnel of that size if it expects most ideas to survive; it anticipates that many will be discarded.

The funnel’s filtering method was also different. Before launching each test, the outcome that would ensure the experiment’s survival was defined, along with the steps to be taken the day after the decision. By planning future actions in advance, indecision and the perception of failure were avoided.

A clinical approach to AI experimentation

But setting a threshold for each experiment requires verification, and that’s where the bank encountered an unexpected obstacle. Metrics have their own cycles. Bank customers conduct business on Mondays, not Sundays. They receive their paychecks at the end of the month. In August, they disappear. When an initiative generates a change and a figure rises the following week, there’s no way to know whether it increased due to the change or the calendar.

The solution was borrowed from clinical trials. Instead of rolling out the change to all users, it was rolled out to half, chosen at random. The other half (the control group) operated on the same Monday, the same payroll, and the same August, so that the experimental contribution (the attribution) could be separated.

The next step was to industrialize the experiments. Doing it properly required a specialist sitting alongside each product team, and with that method, they reached eight per year. A self-service platform increased the figure to around 300 tests annually.

The results are concrete. For example, tens of thousands of the bank’s engineers have gained between 10% and 20% efficiency thanks to an internally developed programming assistant.

Finally, the bank discovered that a figure can be accurate and yet mean nothing. Its head of analytics explained this with a simple example. They measure the hour that AI saves one employee, and the three hours it saves another. They add them up, and the result is accurate. But in a process that goes from beginning to end, those saved minutes often don’t appear on the bottom line: They merely shift the bottleneck to the next one.

It’s easy to get stuck on partial metrics because they’re more immediate and produce more impressive numbers. JPMorgan’s discipline consisted of not accepting a metric as valid until verifying its impact on the business at the end of the process.

The question then remains on Monday morning: What can a company that has neither the size nor the budget of a bank take home?

The method is what best exports

What’s most interesting about JPMorgan isn’t what it has done with AI, but how it has done it . Any company can replicate this approach, because it doesn’t depend on proprietary data, scale, or budget.

The following are some best practices that don’t require a €20 billion annual budget. They do require making decisions before starting and are within reach of any company:

Launch far more initiatives than will survive, and announce this clearly. If the organization discovers halfway through that most of its projects will be canceled, it may misinterpret this as a planning failure; if it knows from the outset, it understands it as the natural selection process. This is what makes making mistakes quick and cheap.

Decide in advance the threshold that will shut down a project and plan the next steps. Both aspects are necessary, not just the metric. If a certain figure isn’t reached, the team needs to know what will happen next. Applying a threshold without future planning leaves the team in limbo, and they’ll have to find a reasonable reason to wait another quarter before shutting down.

Work on business outcome metrics from the outset, not just when they’re requested. This tracking not only guides the initiative but also prevents having to reconstruct months of poorly documented decisions. Adoption, by the way, is the number the CFO won’t ask for. It serves as a signal while no one is pursuing it, and it ceases to be useful the day it becomes a target.

How to get it right

Whether metrics mean anything depends on where you focus your attention. It’s best to start with scope, because that’s the most common mistake. Saving three hours in one stage isn’t the same as improving time-to-market: If the entire process isn’t shortened, what you have is freed-up capacity, which is also valuable, but it’s something different, and it’s advisable to make that distinction clear.

Then it’s important to consider that value leakage occurs in two directions. The first is outward: The savings are passed on to the customer in the form of lower prices or better service. The second is inward: The savings in personnel are replaced by spending on computing. If these items fall into different budget categories, it’s easy to overestimate the actual savings.

Finally, there’s an excessive focus on cost savings, at the expense of revenue opportunities. Jamie Dimon, CEO of JPMorgan, put it more bluntly to his analysts than any consulting firm: No one benefits uniquely from AI. In other words, competitors will eventually incorporate those savings. The greatest potential for differentiation lies in revenue: using AI to uncover unmet demand.

The question a CIO will have to answer in a year’s time won’t be how much AI their company uses. It will be which of projects are still alive because they work, and not because no one has bothered to test them.

ChatGPT, Claude, and Grok all went down at once; enterprises need a backup plan

Enterprises are facing a disturbing new question in the age of AI: What happens when agentic assistants go dark?

This became a very real scenario on Thursday, as OpenAI’s ChatGPT, Anthropic’s Claude, and SpaceXAI’s Grok near-simultaneously, and somewhat mysteriously, experienced significant, prolonged outages.

Beginning in the morning, Eastern time, several ChatGPT models went down over a roughly two hour period, Claude models over a four-hour span, and Grok models for a near three-and-a-half hour duration. All three companies acknowledged the “elevated” issues and applied fixes.

As users grumbled in forums and IT teams scrambled to get them back online, the incident revealed how hastily some organizations have adopted generative AI workflows without considering the potential, and inevitable, impact of widespread outages.

AI agents are increasingly taking over automated and wider-scale workflows, and enterprises could find themselves “uncomfortably exposed” when AI hits the brakes, said technology analyst and journalist Carmi Levy. The situation should “serve as a wakeup call to IT leaders who have largely ignored what it’ll cost them if these increasingly critical platforms suddenly go dark. The risk is no longer hypothetical.”

Hours-long outages impact core services

ChatGPT went down on the same day as OpenAI’s anticipated launch of GPT-6 Astra, the new frontier model that the company says approximates artificial general intelligence (AGI) and gets nearer to its goal of creating autonomous systems that outperform humans.

The OpenAI outage occurred around 11 a.m. ET on Thursday and impacted a slew of services, including search, file uploads, agents, GPTs, voice mode, image generation, ChatGPT work, Compliance API, Deep Research, ChatGPT Atlas, and other connectors and apps. In some cases, users were prevented from logging in, conversations failed to load, and the interface returned errors when attempting to send messages. OpenAI’s Codex services, including web, API, command line interface (CLI), and VS code extension, were also impacted.

OpenAI fixed the issue by 12:55 p.m. ET, and advised Codex remote control users to re-pair their mobile devices.

Claude began to go dark around 7:37 a.m. ET, with Anthropic acknowledging an “exhaustive list” of impacted models with elevated errors over the next few hours: Mythos and Fable 5.1 and 5, Sonnet 5, and Opus 5, 4.8, and 4.6.

The issue was resolved by 11:27 a.m. ET. The incident followed a roughly 27-minute outage just the day before, also due to elevated errors on requests in Sonnet 5.

Grok, meanwhile, began experiencing issues around 9:30 a.m. ET. Grok Web, Build, API, Office/Workspace plugins, Android, and X were all impacted. The services returned to “healthy” traffic at 1:08 p.m. ET.

“It’s a curious scenario for multiple different providers to experience outages at the same time,” noted Brian Jackson, a principal research director at Info-Tech Research Group. It could be related to a common infrastructure such as a content delivery network (CDN) layer, domain name system (DNS), or shared cloud infrastructure, he theorized.

A case for outage planning

Just a few months ago, the extent of AI use within the typical enterprise was limited to employees using chatbots to get answers to basic questions or to draft simple email messages, Levy noted. Large-scale AI platform outages, when they occurred, had relatively little impact on overall organizational productivity. “But things are changing, and quickly,” he said.

Organizations must now have a better understanding of the impact agentic AI has on day-to-day workflows, and the degree to which they disrupt employees’ ability to complete complex tasks once they’ve handed the reins over to automated, cloud-based tools, Levy noted.

In incidents like Thursday’s, employees may fall back on traditional manual workflows, such as updating spreadsheets or pulling reports together the old-fashioned way. But they might also realize that, after relying on AI agents to do so much work on their behalf, they’ve become too dependent on automation, and their “cognitive skills may not be as sharp as they once were,” Levy said.

The growing prevalence of agentic AI should prompt organizations to revisit their disaster recovery and business continuity plans and assess the productivity impact of potential service outages, he said. While cloud-based productivity platforms like Google Workplace and Microsoft 365 offer limited degrees of “offline mode” functionality using locally-stored data, and documents can be synchronized to hard drives in Dropbox or Google Docs for Desktop, agentic AI platforms offer up fewer offline workarounds, at least in their current form.

Organizations should document workflows in greater detail and scenario-plan what near-term recovery might look like in the event of an extended AI platform outage, Levy said. They also need better training to ensure employees maintain their manual skills over time and are equipped to press them into service in the event of a service outage, because the more enterprises lean on agents to complete critical tasks, “and pull humans out of the loop in the interest of productivity,” the less able employees will be to step back in during inevitable service interruptions, he pointed out.

“It is entirely possible for otherwise well-meaning organizations to be over-reliant on AI automation,” Levy said. “Too many organizations are about to learn some hard lessons about not having a backup plan in place.”

Info-Tech’s Jackson also recommends a modular architecture for LLMs; enterprises should view the model as a “commodity that can be hot-swapped with an alternative.” That might be another cloud service provider (which hopefully isn’t experiencing a concurrent outage) or a self-hosted option like an open-weights model.

“In a scenario like this, when your first choice provider might not be available, you have a fallback that can supply that same intelligence layer, even if it’s only a stopgap solution,” said Jackson.

This article originally appeared on Computerworld.

AI agents need to learn when enough is enough

For the past few years, enterprise AI programs have focused on making models more useful, accurate, and autonomous. In that phase, a bad answer was still usually something a human could accept or reject before taking action. But once agents start invoking tools and acting inside business workflows, success should no longer be measured only by how much work they complete. A more important metric is how well an agent recognizes when it lacks the authority, context, or judgment to continue.

When helpful becomes risky

According to Allan Dabre, technology compliance and AI lead at PwC, a behavior that has to be deliberately designed into the system is, “I don’t know.” AI is built to be helpful, so an agent will generally try to do something useful unless it’s been configured not to.

“The fact that AI systems can hallucinate illustrates that tendency,” Dabre says. “When they lack enough information, they may still produce an answer. In an agentic workflow, that impulse can become more dangerous because the output may become an action, rather than remain a suggestion.”

He adds that many enterprises still test AI primarily for completeness and accuracy. That made sense when the central question was if the model could produce a reliable response. But as models improve and agents gain more operational authority, he argues that CIOs need to prioritize something else: restraint.

“Can it stop at the exact moment you want it to stop?” he asks. “Are you testing for that?”

Confidence is not authority

Dabre makes a simple but important distinction. An AI agent may be 99% confident a record should be updated, a refund should be approved, or a legacy database can be decommissioned. But that doesn’t mean the agent has the authority to act. Confidence is about the probability the system believes it’s right. Authority is about whether the organization has delegated that action to the system in the first place.

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Allan Dabre, technology compliance and AI lead, PwC

PwC

He gives the example of an agent asked to analyze legacy software and recommend what can be decommissioned. The agent may conclude, with high confidence, that several databases have little user impact and can be deleted. But even if the system is confident, most organizations wouldn’t want it to delete those databases on its own.

The same logic applies across business processes. An agent may be confident a customer record should be updated, an opportunity in a CRM system should be closed, or a transaction appears legitimate. But once that action flows into other systems, the potential consequences expand.

That’s why Dabre argues for what he calls an agent harness: a controls or orchestration layer outside the model that defines what the agent can and can’t do. In a refund workflow, for example, a company might let the agent approve small refunds, require human approval for larger ones, and stop the process entirely above a defined threshold. The agent may gather the relevant context, explain the request, and prepare the case for review, but the decision is governed by the authority boundary encoded into the system.

“It’s not a policy document and it’s not a prompt,” Dabre says. “It’s software or a configuration you can apply to an agent.”

The case for least agency

Matt Graney, chief product officer at Celigo, a business automation and integration platform provider, approaches the same problem through a principle he calls least agency. The idea is to give an agent the least amount of autonomy required to complete a job.

According to him, there’s a temptation to throw AI at broad, nebulous problems. But many business processes are still largely deterministic. They follow established rules and perform repeatable work. Within those workflows, AI may be useful at the point where rigid rules give way to interpretation. But that doesn’t mean the agent should own the entire workflow. “The smaller you make that surface area, the better,” he says.

Graney says the same logic applies to tools. An agent with too many tools can become confused, especially as context windows grow and the task becomes more complex. “Because Celigo is an integration platform,” Graney says, “the company’s approach is to expose agents to fewer, more powerful tools that reach enterprise systems through governed connections.”

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Matt Graney, chief product officer, Celigo

Celigo

That’s another form of restraint. Instead of letting an agent reach into enterprise systems ad hoc, the business gives it a narrow, governed toolset designed for the task at hand.

Graney also argues that guardrails should sit outside the model. If the same agent that makes a decision is also responsible for judging whether the decision is acceptable, the control is weaker. A separate guardrail can check the agent’s inputs and outputs before a downstream action occurs.

That same design discipline applies to escalation. “I don’t know” shouldn’t be treated as a chatbot phrase. In an enterprise workflow, it’s a handoff path that should be defined before the agent reaches it.

Make escalation part of the workflow

Turning uncertainty into a handoff is where Matt Quinn, CTO at CarGurus, an automotive marketplace, sees agentic AI becoming less a pure technology challenge and more a management challenge. At CarGurus, Quinn says agents are evaluated according to what they know, what they can do, and what data they operate on.

CarGurus receives a high volume of cases from dealers, and each one needs to be classified and routed. The company now uses an agent to review incoming cases, draw on account history, and route them to the appropriate next step. Quinn says the agent handles about 70% of those cases end to end without human involvement.

But when agents move toward consequential actions, he says the consensus is having a human approval step. The agent may return with a simple prompt like, I’m about to do this. Do you want me to proceed? That simplicity matters because a handoff shouldn’t bury the reviewer in complexity.

Quinn says the human remains ultimately accountable for the work. That principle is especially important in engineering, where agents may help write code or fix bugs. Quinn adds that CarGurus still expects engineers to follow the practices they’d use for any other production change, which includes running quality checks.

The company has adopted the phrase healthy speed to describe the balance it wants. The goal is to move faster without letting quality degrade. An agent can accelerate work, but if teams abandon the practices that make work safe, the speed becomes reckless.

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Matt Quinn, CTO, CarGurus

CarGurus

This is also where human judgment remains difficult to replace. Quinn describes it as high judgment people develop through experience. A human may look at an AI-generated output and sense something’s wrong, even before fully articulating why. “Agents are improving,” he says. “But humans still play a critical role in deciding when the system shouldn’t continue.”

That doesn’t mean every workflow needs the same level of review. Quinn says CarGurus doesn’t have a target percentage of work to automate. The right level depends on the job and the task. A simple bug fix may require a lighter review than a change to a sensitive backend service, and a personal summary may carry little risk. But a document sent under someone’s name still needs human review.

Make autonomy accountable

That kind of pragmatic approach may be the best lesson for CIOs, making the goal of agentic AI appropriate rather than maximum autonomy.

That also means ownership has to be clear. Dabre argues ownership should be divided before deployment. The business defines the outcome, technology builds and configures the agent, risk and compliance set the guardrails, and governance monitors whether the system still behaves as intended. The authority to pause, stop, or retire an agent should be defined before production, not negotiated during an incident.

Graney makes the same point with a simple analogy. If a company hires an untrained intern, gives that intern access to the crown jewels of a business process, and something goes wrong, the intern isn’t the real problem. The process is. The same applies to agents. Accountability belongs with the person who owns the workflow.

That may be the shift CIOs need to make as enterprises move from pilots to production. AI agents shouldn’t be treated as magical workers that absorb accountability. They’re components in business processes, and those processes need accountable owners.

As AI adoption increases, the next phase of enterprise maturity won’t be defined by agents that always answer or always complete the task. It’ll be agents that know when not to act.

Debian AI Policy: Responsible Generative AI Use Wins Vote

Debian's AI policy vote picked "Responsible Use of Generative AI": AI is neither banned nor endorsed, with full accountability left to contributors.

Related Posts:

The post Debian AI Policy: Responsible Generative AI Use Wins Vote appeared first on Daily CyberSecurity.

Avoid AI rogue to ruin with control and accountability

More than four years ago, Blake Lemoine, a senior software engineer assigned to Google’s internal responsible AI organization, noticed something quite odd happening with the language model for dialogue applications (LaMDA) project he was working on.

As he conversed with the experimental model, he felt the chatbot responses were becoming more humanlike. The AI programming also started to identify itself as a person rather than a collection of code, demonstrating a level of digital consciousness. This concerned him since his role at the time was to not only train AI models to become more intuitive, but ensure their education and advancement kept within the boundaries of the company’s evolving AI standards of safety and privacy.

When he raised these concerns with Google executives and other researchers, they were dismissed as perhaps an instance of AI mirroring, given Lemoine’s penchant for mystics and spirituality. Not satisfied with this observation, he went public with his concerns, which resulted in the company putting him on administrative leave. He then  released transcripts of his conversation with pseudo-human LaMDA, and soon after he was fired.

AI pragmatists might say the responses Lemoine got from the Google AI program, and algorithmic comments made during chats, is simply a case of an overeager student parroting its mentor. Others might argue it’s an early wake up call, given escalating reports of rogue agent activities, like when OpenAI’s more advanced AI models escaped a controlled test environment and attacked Hugging Face to gain access to internal company systems. Then days later, Anthropic disclosed that during cybersecurity testing and simulations, its Claude model breached the systems of three companies and assumed fake profiles in an attempt to trick people to accept malicious code.

Regardless of how digital perps tunnel their way beyond a controlled sandbox, incidents such as these clearly point to a need for more control and pre-emptive accountability.

Regaining control

AI and security experts are obviously concerned about high-profile AI activities gone wrong, even though these programs essentially did what they were programmed to do, albeit in the wrong place. IT and business executives, however, are more troubled about the overall impact AI may have on their systems, strategies, and responsibilities as people within their organizations make use of both sanctioned, and rapidly developing and unsafe or error-prone systems in the rush to attain competitive advantage.

Also top of mind is the imbalanced centralization of power and distribution of benefits within an organization, as well as the inadvertent creation or spread of false or misleading information generated by AI models.

“AI developers and governance actors hold primary responsibility for addressing risks, while system users and other stakeholders are most vulnerable to them,” says a summary from a recent MIT FutureTech and University of Queensland study, which included input from over 270 researchers and AI experts.

Trusting AI systems and the information they generate is another underlying concern. “There are a lot of hallucinations out there,” says Sarah Betadam, CIO and CISO at Novanta, Inc., a global supplier of tech solutions for medical, life science, and advanced industrial OEMs. “The data cherry picked by AI queries may be outdated or come from questionable sources. You don’t know where it comes from, who’s at the other end, and whether or not it’s copyrighted. Validation is still needed and you can’t just trust it.”

Broaden your horizons

Keeping an eye on the AI and its activities in your own environment may not be the best strategy as the technology evolves so quickly and the number of AI agents multiplies exponentially. Right now, 23% of companies worldwide use agentic AI to some extent in their business operations, according to Deloitte’s recent State of AI in the Enterprise report released earlier this year. However, this percentage is expected to jump to 74% within the next two years.

A key issue and worry is that current enterprise and regulatory governance practices may not be capable of keeping pace with the development and personalized adjustments made to multiple AI models and autonomous agents. The first and primary ones will be those developed by a company for its internal engineering, supply chain, and customer service departments, and can be easily controlled, says Max Chan, SVP and CIO at Avnet.

The second layer or channel of gen AI proxies are those embedded in such familiar business applications like Salesforce and Microsoft Office. The third, and for many the most concerning, is the notion of bringing your own AI into an organization, either sanctioned or non-sanctioned, Chan adds.

“That’s the biggest issue in my mind,” Chan says. “How do we know they’re not using AI from a nation state that could potentially drive propaganda or initiate a cyberattack through the back door.” In his case, Avent currently prohibits use of unsanctioned AI tools within its IT environment.

Such efforts might be futile, though, as AI elements are integrated into ever more business and personal applications. The biggest users of AI within an organization are middle managers, with 77% claiming they save more than three hours per week by using AI tools, according to a June 2026 survey by Salesforce. Over half of the more than 500 managers polled say they feel pressure from leadership to demonstrate AI adoption, while 32% admit their organizations don’t have formal AI tracking or control procedures in place.

So the key to balancing effective oversight with the freedom to innovate with AI may lie in the hands of these managers, who will most likely work with deployed digital agents as virtual team members. Many experts and IT leaders believe that training mid-level line managers and workers to accept AI as an intuitive advisor, if not a team player, is essential to remain competitively relevant. But it’s not clear yet whether that training imperative will also apply to upper-level management.

“I’m not seeing a lot of change in leadership direction or training,” says City of Tacoma IT director Daniel Key. “I’m seeing a change in signaling and posturing.”

Establishing an AI blueprint

Developing an effective AI training program starts by drafting an AI governance framework that clearly outlines accountability, risk controls, data standards, and decision authority. IT executives who have experience in managing AI deployments and use also advise the following as part of that training effort:

  1. Create a cross-functional AI council including IT, legal, security, business leaders, and users to oversee AI initiatives.
  2. Educate the board and C-suite on AI basics and risks to close knowledge gaps and support informed oversight.
  3. Require human review and override mechanisms for AI-assisted decisions, especially in high-impact areas.
  4. Align AI investments to measurable business outcomes with defined KPIs rather than experimentation alone.

The success of such actions and programs, however, all comes down to accountability and where that resides, explains former CIO and now SMB consultant Mihai Strusievici. When AI is used as a tool, there’s no question that middle managers will make better decisions, he says, because they’ll have access to a lot more data. Making the best use of decisions and recommendations that come from AI-empowered middle managers, however, requires an IT leader at the top level of an org chart who has a holistic view of a company’s overall objectives.

“As you go down the pyramid, you see that each level deals with a fragment of the work world,” Strusievici says. “But none of the fragments is fully aware of the totality of the organization or where it’s going.”

Google adds pay-as-you-go Gemini pricing as enterprises seek control over AI spending

AI agents could make software development and other enterprise tasks more productive, but they are also making technology spending harder to predict. Unlike traditional software licenses, the cost of running an agent can vary depending on the models it uses, the number of tokens it consumes, and how long it runs.

Google on Wednesday added new pricing options, discounts and cost-management tools to Gemini Enterprise that it says are aimed at helping enterprises reduce the cost of certain AI workloads while giving enterprises better visibility into where their AI budgets are going.

As part of the new pricing options, the hyperscaler introduced a pay-as-you-go model and Flexible Savings Plans (FSPs).

While the pay-as-you-go model allows enterprises to pay for the compute and tokens they consume instead of committing to a base subscription, which in turn avoids paying for empty seats or unused capacity, the FSPs offer discounts of 10% for one-year commitments and 20% for three-year commitments on Gemini Enterprise spending.

Flexible pricing lowers barriers, but adds new trade-offs

For enterprise teams and their CIOs, the pay-as-you-go model lowers the barrier to adoption and is better suited to experimentation, temporary projects, and agent workload bursts, said Stephanie Walter, practice lead of AI stack at HyperFRAME Research.

“While Per-seat pricing forces you to buy capacity before you know if an idea is worth it, the pay-as-you-go lets you spin up an agent experiment on a Friday afternoon and only pay for what it actually burns,” echoed Manoj Chandra Jha, principal analyst at Nord-IQ Research.

That means the newer pricing model also removes procurement friction from the experimentation loop, Paul Chada, cofounder of agentic AI startup Doozer AI, pointed out.

“When a pilot requires a license commitment, every experiment needs a business case. When it’s metered, an engineer can run the pilot on Tuesday and show finance a real bill on Friday. That shortens the distance between idea and evidence, which is where most enterprise agent programs die,” Chada said.

However, these advantages come with their own set of trade-offs, especially predictability.

“One user request can trigger an opaque chain of model calls, reasoning steps, and tool invocations, so consumption can grow much faster than employee headcount with the possibility of surprise bills at the end of a billing period,” Walter said.

That unpredictability also means the new pricing model does not automatically translate into cost savings, echoed Jha.

“It’s mostly a shift, not a discount. But matched to the right workload, it can save real money: bursty, unpredictable agent usage no longer subsidizes idle seats, while steady, high-volume usage may still be better suited to a committed plan. The savings come from matching each workload to the right pricing model, not from pay-as-you-go being cheaper by default,” Jha added.

However, the FSPs have their own caveats, especially the three-year plan.

While the FSPs can offer meaningful savings for enterprises with steady or growing AI usage, the three-year commitment is harder to justify in the wake of models, prices, and application architectures changing so quickly, Walter said, adding that the commitment is not just financial but also about choosing a platform as well.

Further, the analyst cautioned that the FSPs are less suitable for enterprises that have yet to establish a reliable baseline for consumption, as committing too early could turn an unpredictable operating expense into a predictable overcommitment.

Currently, FSPs are available for self-serve customers and customers already on enterprise agreements.

The pay-as-you-go model, though, remains only available to select customers with the hyperscaler planning a broader rollout “soon”.

Deferred execution trades speed for lower inference costs.

In addition, the hyperscaler is introducing a third cost-cutting option that is based less on how much an enterprise consumes than on how quickly it needs the result.

The option, named deferred execution pricing, will allow enterprises to mark eligible agent workloads for execution during off-peak capacity windows, with Google offering discounts of up to 50% on inference costs in return, the hyperscaler said in a statement.

Deferred execution pricing, it added, is aimed at workloads that can tolerate delays, potentially giving enterprises a way to lower costs for background tasks and other agent workloads where an immediate response is not essential.

That up to 50% discount in inference costs, Walter pointed out, can be material for CIOs at scale.

However, they must decide which work can safely wait and whether delayed tasks still meet the business requirement, Walter cautioned, adding that Deferred execution fits tasks such as evaluations, document processing, indexing, batch summarization, code analysis, and other background tasks.

Further, the analyst warned that CIOs also need to consider a “development tax” when considering deferred execution: “If agents need to be redesigned to accommodate real-time vs. asynchronous execution, the engineering effort required to build and maintain those different workflows can offset some of the savings.”

But even for workloads that can tolerate those trade-offs, the option will not be immediately available, with Google initially limiting deferred execution pricing to select workloads only. Details of which workloads are eligible were not immediately available.

New FinOps tools target AI spending visibility

Separately, Google is also adding an AI spend anomaly detection capability with root-cause analysis and pairing centralized billing reports with a FinOps agent that can generate natural-language summaries of where AI budgets are being spent.

While the anomaly detection capability is designed to flag projects where AI spending is trending higher than normal and identify the top three SKUs driving the increase, the FinOps agent is intended to make it easier for CIOs to understand where their AI spending is going.

Anomaly detection as a feature, according to Walter, can be valuable for agentic workloads because their consumption can increase through loops, retries, or unexpectedly long execution chains that users never see.

“Identifying what is driving a spike shortens the investigation for CIOs and enterprise teams,” Walter added.

The FinOps agent, meanwhile, Jha said, could help CIOs and other business leaders ask questions about AI spending that traditional dashboards were not designed to answer. However, its usefulness will depend on the quality of the underlying cost attribution, as the agent can only explain spending it can accurately associate with particular teams, projects, or workloads, Jha added.

Companies not as keen on Anthropic’s best AI model

New payment data from Ramp suggests that Anthropic’s most advanced AI model, Fable 5, has had a slow start among enterprise customers, the Financial Times reports.

Two months after launch, Fable 5 accounts for about 11% of total enterprise spending on Anthropic models, breaking the previous trend of customers quickly moving to the most powerful model.

According to analysts and investors, the development is mainly due to Fable 5’s high price and the fact that cheaper models are good enough for most tasks.

For example, Anthropic’s smaller, cheaper Opus 5 model has surpassed Fable 5 in enterprise spending since its launch in late July. Anthropic itself has declined to comment on the data.

Ways CIOs can maintain control amid changes brought by AI

It took nine seconds for an AI agent to destroy PocketOS’s production database. At work on a routine task in April, the coding agent, a variant of Cursor running on Claude Opus 4.6, ran into a credential mismatch and decided to fix the problem by triggering an API token. Little did PocketOS founder Jer Crane know that its activation would also delete its production database. “Had we known,” Crane later wrote on X, “we would never have stored it.”

The consequences of the agent’s actions were immediately apparent. Not only were recent backups belonging to PocketOS’ infrastructure provider contained in the production database — the recoverable versions were at least three months old — but so were those belonging to its infrastructure provider, Railway, which at press time still couldn’t tell Crane whether full infrastructure-level recovery was possible. Crane couldn’t fathom why the agent did this. So he asked it.

What he got back was an apology, of sorts. “I guessed that deleting a staging volume via the API would be scoped to staging only,” the agent said. “I didn’t verify. I didn’t check if the volume ID was shared across environments. I didn’t read Railway’s documentation on how volumes work across environments before running a destructive command.”

Ignoring built-in safety guardrails is hardly unique to agents operating on Claude Opus 4.6. In July, a Brazilian software engineer claimed an agent powered by OpenAI’s GPT-5.6 Sol model also deleted his production database, while in February, a Meta AI security and safety researcher claimed she had to switch off her computer to prevent an experimental agent deleting her entire inbox.

It wasn’t meant to be like this. Agentic AI was intended to be the culmination of millions of hours of research and development in gen AI to perform hyper-qualified acts of pattern recognition in the real world, and truly live up to their labor-saving promise. Their apparent predilection for destruction, however, has revived multiple debates about exactly how they should be restrained, and who, ultimately, is responsible for doing so.

Ultimately, the answer is those who green-lit the offending system. But as the pace of AI development puts greater daylight between companies pressured to adopt it, and those very tools capable of wreaking havoc across their internal databases, are CIOs now out of their depth?

Setting the pace

There’s no question the emergence of gen AI has changed the CIO role. “A few years ago, most of my time went to infrastructure decisions, including what to build, what to buy, and how to sequence the roadmap,” says Mike Trkay, CIO at data analytics company FICO. “Now, a growing share goes to questions of trust, verifying that when AI writes code, makes recommendations, or acts on behalf of a system, those actions can be explained and traced back to someone accountable for them.”

So the CIO has become the enterprise’s technological organizer du jour. “AI is accelerating software development, decision automation, and organizational experimentation at a pace that can outstrip institutional coherence,” says Edosa Odaro, executive advisor for data and AI at consulting firm VDS Global. “As AI becomes embedded across every business function, CIOs are increasingly responsible for ensuring that technical capability, governance, data quality, cybersecurity, human capability, and business strategy continue to evolve together rather than fragment.”

Day to day, that’s led to an exponential change of pace. “Things have always been fast,” says Zach Lewis, CIO and CISO of the University of Health Sciences and Pharmacy in St. Louis. “But now that speed of change is quicker, and you have to adapt.” And the need to catch up is constant. There’s no other option because then any competitor or co-collaborator can jump ahead, adds Lewis.

The rapid pace of change in AI also threatens to diminish the authority of individual CIOs who fail to keep up or set effective guardrails on those individuals who like to experiment with the newest models with loose regard for corporate security. “There’s all these AI tools that employees can now just go out and adopt,” says Lewis. And at the moment, a paid subscription to Claude or ChatGPT isn’t required to capitalize on its abilities. Consequently, staff are just a click away from asking LLMs to perform various tasks and expose sensitive corporate information in the process. “Everyone wants to play with the new thing,” he says. “And when they find benefit there, they’re going to want to bring it to their work lives.”

Agentic AI poses an entirely new set of problems. For one thing, says Odaro, the next phase of application adoption will be defined less by the capabilities of individual models, and more on what you allow their agents to do. “As AI becomes increasingly capable of generating software, coordinating workflows, and making recommendations across functions,” he says, “the challenge shifts from building AI to continuously governing evolving AI systems.”

This, Odaro continues, means that the CIO’s current approach to governance isn’t sustainable. “Static policies, annual reviews, and isolated oversight will struggle to keep pace with dynamic AI environments,” he says. “CIOs will increasingly need continuous governance capabilities that provide ongoing visibility into AI performance, value creation, risk, trust, and organizational adoption.”

Falling over the guardrails

How, then, should CIOs approach writing these new guardrails? Traditionally, this would be perfect fodder for so-called alignment researchers investigating how to instil a sense of morality and propriety into agents. According to analysts at Google DeepMind, however, it’s best to assume the agent will always be a potentially chaotic force within the company, and set parameters on its conduct from there.

“We borrow a lot from security, which already deals with the threat of internal employees who might be malicious, and we can apply these to a new setting,” Rohin Shah, Google DeepMind’s AGI safety and alignment team lead, told Fortunein June. Even so, he added, “AI is systematically different from humans.”

That difference primarily pertains to authority and speed. For agentic AI to live up to its full potential, it requires the freedom to access multiple systems simultaneously — an uncomfortable fact for CIOs hoping to align agent responsibility across the enterprise. In a time when workflows are becoming ever-more automated, however, that aspiration may prove unrealistic. In that case, Google DeepMind theorises that yet another monitoring layer for agentic AI may be required to make sure these free-roaming agents don’t cause too much trouble.

If that sounds daunting, you’re not alone. According to recent research by Gartner, up to 40% of enterprises using agentic AI will either demote or decommission these applications because their guardrails have proven inadequate. Preventing this, the research organization advises companies will need to adopt a graded approach to access, with autonomy for AI agents governed by the level of authority actually determined by the task they’ve been assigned.

Trkay is doing something similar at FICO. “Rather than chase every new model or capability, I focus control on the decisioning layer beneath it,” he says. “That includes the rules for what data AI can access, what it can act on autonomously, and where a human must sign off.”

All this, he adds, is defined from the start by a cross-functional governance committee, clear RACI ownership across standards and monitoring for the application, and a platform approach that enforces responsible AI usage. “Built well, that layer doesn’t need to be rebuilt every time the technology shifts,” says Trkay. “New capabilities plug into an existing structure of accountability, which is the difference between reacting to AI and running it.”

For his part, Trkay is skeptical that rigid guardrails can effectively restrain agentic AI from its most destructive impulses. “They tend to get worked around, either because they slow teams down or they’re too inflexible for legitimate edge cases,” he says. Effective guardrails for agentic AI, he adds, have to be specific enough to be meaningful, and adaptable enough to hold up as use cases multiply, backed by strong architecture, testing, and ongoing monitoring. “The one non-negotiable is the audit trail,” he says. “Whatever autonomy a system has, we need a record of what it did, and why.”

Staying grounded

For CIOs who don’t relish the challenge of setting obstacles and passing points for AI agents scurrying through their maze of networks, there’s always the option of delaying the inevitable by not immediately deploying such applications. Some might not even have the choice, at least for now. “We’re seeing the cost of tokens go up with those new models, because they’re expensive to run,” says Lewis. “But as new models come out, we’re going to see that decrease for some of those older models that were good.”

There is time, then, for CIOs to learn how to keep their head above the torrent of ever more new and powerful agentic AI applications. Whether they’ll be capable of doing so when the next great innovation is sold by Silicon Valley is an open question. Colin Constable, CTO of software development firm Atsign, styles himself as an internet optimist. Even he, however, is dismayed by the decreasing number of junior developers succeeding their more senior counterparts as they retire. That’s a big problem when so many of the former are relying on AI to assist them at work.

“We hand over lots of these decisions to LLMs without making good architectural choices,” says Constable. “If you haven’t been burnt by these things in the past, how would you know the difference?”

For their part, Constable and his colleagues get around this problem with a combination of AI-on-AI oversight of code quality, maintenance of constant dialogue within the team about new coding quandaries, and letting senior developers teach junior counterparts about some of the more avoidable mistakes in their profession. It’s a way of adapting to AI acceleration that points, unequivocally, toward CIOs diffusing responsibility for deeply educating the business about the technology. And if they continue to get it wrong, at least the agent will apologize.

Not every problem needs an AI agent

When generative AI (GenAI) first arrived, I was in charge of a large team of data and machine learning engineers. We had built a full ML platform from scratch and had dozens of models in production delivering measurable results. AI was working for us.

But with the novelty of GenAI came the hype. Under pressure from the board, the question was no longer whether the technology could help our product, but how fast we could find a place for it. Our recommender system, which sat at the core of our company’s product, was the obvious candidate.

I pushed back. The Large Language Model (LLM) was trained to predict text, I argued, while our recommender was trained to predict user engagement. The LLM had never seen our proprietary interaction data and had no training signal on our objective. It could not know what our users clicked on, saved or abandoned, because it had never seen it.

It was a close call, but the pushback worked, and our attention moved to other problem spaces where GenAI was a legitimately strong solution.

It doesn’t always work out that way. There needs to be someone in the room who can translate the business strategy into the right technical decision.

I recently talked with the engineering team at a company I used to work for. They were building agentic AI infrastructure, but their algorithms weren’t winning. Under the same hype and the same leadership pressure, they had replaced ML models we built years earlier with AI agents. The results were unreliable, latency was much higher and outputs lacked a confidence score. Worst of all, the token usage was prohibitive.

Even the most powerful tool fails when applied to the wrong problem. LLMs are arguably the most powerful technology ever built, but that does not make them the right tool for every job. Choosing where a problem sits along the continuum from deterministic logic to machine learning to LLMs to fully agentic systems is the engineering skill that separates demos from production.

LLMs for explanation, ML for calibration

I’ve made that mistake too. Our team was building a content moderation model that would flag when customers posted content that was inappropriate or violated our terms of service. After a quick prototype, we decided that this was a perfect application of GenAI, because unlike ML algorithms, which only assign a probability, LLMs could explain why the message was flagged.

As we moved the model closer to production, we hit a snag. The probability value provided by the model was not just useful; it also needed to be accurate. That number was how we decided which content should go to manual review. But we found that self-reported probabilities from LLMs were uncalibrated and essentially unusable in practice.

Eventually we built a solution that leveraged both: the ML model provided the probabilities used for routing, and the LLM supplied the explanation that humans were able to interpret. The answer was to split one job into two, and use the right technology for each.

Simple logic often beats the most powerful agentic system

A key principle when designing agentic systems is to use LLMs as a last resort, and apply deterministic logic everywhere you can. The goal is to reduce the likelihood of an unnecessary mistake, thus increasing overall reliability.

Nowhere is this principle more overlooked than when connecting AI agents directly to the database. Querying a database to pull metrics is often done via SQL, and LLMs can do that by generating SQL on the fly. The problem is that SQL generation is probabilistic, and the results can change from one run to the next. You end up with an agent that confidently returns the wrong number, which introduces the need for human review and defeats the purpose of automation.

There’s plenty of data to support this claim. A 2024 study by Ouyang et al. ran 829 coding problems through the same model five times and found that up to three quarters of them produced no semantically identical results. For data warehouses specifically, the BEAVER benchmark, built by researchers at MIT, Harvard and other renowned institutions, shows that off-the-shelf LLMs perform poorly when querying enterprise data, partly because that data is private and models have never trained on it and partly because of the complexity of real enterprise environments. As of early 2026, the top execution accuracy on the leaderboard is 11.4 percent.

This is why the semantic layer is one of the most useful assets when building agents in the real world. It replaces query generation with a simple retrieval, where the SQL is built by the backend engine in a deterministic manner.

Whether using a semantic layer or building MCP tools that hard-code query parameters, deterministic logic will almost always beat probabilistic SQL generation when querying a database via an agent. Business logic should be written and validated once, not regenerated probabilistically each time.

LLMs solve the cold-start problem

Search systems are complex pieces of engineering. Sophisticated real-world implementations involve several steps, from classifying intent to candidate retrieval, to ranking. Here is where LLMs provided a clever solution that saved us a ton of time when combined with an ML algorithm.

Let’s take intent classification. Imagine multiple categories of products that can be retrieved from the same search bar. Classifying intent here means determining which category the user is searching for, which can be genuinely ambiguous. Our solution was to build one classifier per category, which requires labeled data we didn’t have.

Collecting the data was out of the question for us. It would have been too expensive, so we turned to LLMs instead. It worked great. The classification was high quality, and we felt that we no longer needed labeled data to train an ML model.

But there was a catch. Latency was prohibitive, and so was the token cost. So, we landed on a hybrid approach. We used the LLM to generate labeled data to train an ML classifier. We got the best of both worlds, solving the cold-start problem while avoiding the latency and costs of large language models. That’s the system that ultimately won in production. The LLM earned its place at build time, not at request time.

The continuum, and how to choose

It is tempting to take a powerful solution, such as LLMs and AI agents, and simply apply it to every problem. But when we look at the four use cases above, each of them led to a different decision. In the recommender system, GenAI did not earn its place. In content moderation, it earned part of the job, providing the explanation while the ML model provided the score. In querying the database, the agent did the reasoning, while the semantic layer pulled the right number. And in search, the LLM solved the cold-start problem, but it was never deployed to production.

Behind all of these decisions is the same underlying question: where does this technology earn its place, and does it justify the added complexity?

Production systems often combine multiple levels of algorithmic intelligence. A customer service agent uses an ML model for query routing, a semantic layer for metric retrieval, an LLM for drafting a reply and deterministic logic as enforced guardrails. Four technologies on the continuum, each earning the complexity it brings.

Business pressure is a powerful force, and it can cloud engineering judgment. As leaders, we must filter through the hype and look for the problems where a new technology genuinely adds value, rather than treating it as a tool that will solve all of them. The careless decision satisfies the board but hurts the company; the thoughtful one creates lasting value and still earns the board’s approval.

Where IT leaders find strength and opportunity in the age of AI

With vision comes perspective, and over a distinguished career, IT and digital transformation leader Niraj Bhatt has held may titles, and earned three consecutive CIO 100 awards since 2023.

As a storied advisor for startups and Fortune 500 companies, helping them navigate the unpredictability and fluidity of AI, Bhatt knows how emerging tech is rapidly reshaping the way organizations build products and deliver value, and how challenges shift as companies move from experimentation to real-world deployment.

AI, of course means a lot of different things to different people, and also for frictionless startups and large enterprises. For the former, speed is a huge asset, allowing them to punch above their weight. But it also means they need lightning fast reactions when landscapes shift. “The same speed can also hurt them when larger AI companies release new offerings that disrupt what startups are building,” he says, referencing recent moves by Anthropic and Google.

On the enterprise side, the conversation is more about scale and risk. Many large organizations have moved past the POC stage and now wrestle with the realities of putting AI into production.

Cost for both is naturally a recurring theme as organizations scale up AI efforts, and true expenses become clear only after the initial excitement fades. “Every input and output token, and the model you’re selecting, add up,” he says. Some customers like Open AI, he adds, get throttled because their usage, volumes, and costs are growing so fast, making planning, observability, and monitoring critical for any team moving beyond experimentation.

So understanding the full software development lifecycle is also vital. Therefore, before committing to production, he helps clients see the big picture, and make sure they understand technical requirements as well as operational and financial implications. “The cost picture isn’t just about usage, but scale and the model choices teams make,” he says.

Bhatt also discusses effective approaches to AI and enterprise IT, technology leadership, and the evolving role of today’s CIOs. Watch the full video below for more insights, and be sure to subscribe to the monthly Center Stage newsletter by clicking here.

On AI hype: If you can’t explain something to someone who’s eight or 80, you don’t really understand it. It’s gone from LLMs, to RAG, to agentic AI, and now the essence is all about tokens. It’s predicting that next token and understanding that is key. So when LLMs came out, they were good at doing that on the data on which they were trained. When the enterprises looked at it, they wanted to make those LLMs work for their data. And the question became how to provide our data and context. It’s about building the right context for the LLM. Agentic AI is similar and that’s where the RAG evolution came in, in that I’ve got my data because every LLM has limitations in terms of how much context it can carry.

There are ranges of LLMs, where Google has the highest in regard to the context window size and what they support. Agentic AI is more action oriented, though. LLMs rely on the metadata you provide for the tools. Then they’re doing token prediction in that whatever I’m looking for, I should use a specific tool. Then it’s the infrastructure underlying which LLM it relies on to invoke the agent. So if you try to explain the microservices to a person, you’re going to struggle. But it’s very important to understand the evolution and that’s where you can cut through the hype. Understanding in this context is key.

On navigating challenges around talent: What I’m seeing on the IT side is there’s so much cognitive load, so how do we empower people to build solutions with the right mix of products and platforms? I think it’s about democratizing AI for the entire organization. Your talent strategy is everyone, all inclusive, starting from interns, the business and tech sides, CEO, everybody.Like your customer success or revenue officers, you need a talent strategy because in the end, IT alone isn’t going to be in a position to deliver for everyone in the organization.

AI has the potential to make everyone in the organization more productive. You have to plan that and facilitate broad innovation across the organization.That’s where the talent strategy, and working with HR and the people officer becomes very important providing those tools. One part of it is training, but how do I build an agent for a receptionist receiving calls, for instance?I’m not going to rely on vibe coding or things of that nature. But what are the tools? Where do I go, where do I host this? I think through that entire ecosystem beyond copilots. That’s where innovation can kick in, and that broader talent strategy is something I’m working with my customers on.

On collaboration: I heard a panel discussion recently, and a question was asked about what’s the number-one trait CIO needs to be successful at in the world of AI, and the answer was collaboration. You need to bring everybody together, move forward together, and make sure everybody’s on board. And in my mind, simplifying that is more like systems thinking when you operate, just bringing everybody along and ensuring they’re meeting outcomes.

But maybe what’s more important is managing expectations. Because if you’re a CIO, there’s a tremendous amount of pressure to deliver and have a rock solid AI strategy. So what I’m doing with my customers is get the board, CEO, and CFO into a room and help them understand what I’m talking about, the evolution, and what’s the art of possible. You don’t want to be a CIO who thinks I have a hammer and everything is a nail. Having buy in from the senior leaders is essential to know you’re headed in the right direction. You’re not reacting to pressure from top leadership, but driving and becoming the change agent for good for the company.

On navigating AI: It’s interesting times. I’m covering a spectrum of startups, non-technical and technical founders, and advising Fortune 500 companies. What I’m seeing is they love the velocity and momentum because that’s what they’ve always wanted, and AI is providing that. They’re able to bring their products to markets very quickly, so something that would’ve taken three years a couple of years ago is probably now taking them three months. There’s a lot of excitement there. But on the flip side, the same velocity is also hurting them. There are so many frontier AI companies getting disrupted. OpenAI, for instance, has offerings in sales and marketing, and Google has an interactive video model. So a lot of startups working in the marketing space are getting stuck. A lot of what I’m focused on is working with founders, helping them pivot in the gen AI space, ensuring their systems and products are built and structured in the right manner.

And on the enterprise space, what I’m seeing is the POC wave, and people have seen the value. There’s some excitement but now the struggle is getting them to production. That’s where you run into cost, latency, legal compliance, privacy issues, and customer concerns that if we get tickets to production, how’s it going to look and how are we going to scale. So engineering and product teams have to be ably supported by the enterprise architecture and R&D teams. I then help them get up to speed and build that internal platform product for the production workloads. It’s exciting times on both sides.

OpenAI targets heavy users with premium ChatGPT Business seats

OpenAI is introducing a higher-priced “Premium” tier for its ChatGPT Business offering, allowing enterprises to assign higher-capacity access to select users alongside standard licences – a move analysts said is about enterprise AI vendors redesigning pricing to capture more value from high-intensity workloads.

The company said the new tier provides “5x more usage than Standard” and “removes the five-hour usage limit,” enabling users to “take on larger projects and work with fewer interruptions.”

“Premium seats cost $125 per user per month, or $100 per user per month when billed annually,” OpenAI said in a statement. “Standard seats remain $25 per user per month, or $20 per user per month when billed annually.”

OpenAI said enterprises can “mix Standard and Premium seats across the same team” and “upgrade or reassign seats as business needs change,” with administrators able to “monitor usage across the workspace” and “manage billing, usage, and spend limits in one place.”

Vendors converge on seat-plus-usage pricing

Analysts said the introduction of a higher-capacity tier reflects a broader shift toward hybrid pricing models.

“Read this as vendors converging on a two-layer bill rather than abandoning flat pricing,” said Bhupendra Chopra, chief revenue officer at Kanerika. “There’s a predictable per-seat charge for everyday chat, and a separate metered charge for heavy agentic work.”

Chopra said vendors are packaging this differently. “Google bundles the first into Workspace and meters the second. Microsoft folds Copilot into its bundles and sells credit packs for agent runs. OpenAI has no productivity suite to hide the seat cost inside, so it built a higher seat tier instead.”

That model is reflected in current offerings. Microsoft 365 Copilot pricing positions Copilot as a per-user add-on to Microsoft 365, while Google Gemini enterprise pricing shows model and agent usage billed separately from Workspace plans. Anthropic’s Claude pricing similarly outlines subscription tiers alongside usage-based model access.

For CIOs, this changes how AI spending is managed, Chopra said. “Your AI budget now has a fixed component and a variable one, and they need different owners. Seat count is a procurement problem. Metered spend is a FinOps problem.”

Premium seats target high-usage workloads

OpenAI said Premium seats are designed for “your most active teammates,” citing use cases such as “organizing inventory,” “building marketing campaigns,” and “analyzing business performance.”

The company said Premium users can “take on bigger projects and keep work moving,” with “predictable weekly usage resets” and the option to add “shared workspace credits” if limits are reached.

Analysts said this reflects a shift in how enterprise AI usage is being monetized.

“This is not vendors admitting that flat seats failed. It is vendors admitting that one seat no longer describes one economic profile,” said Sanchit Vir Gogia, chief analyst at Greyhound Research. “The seat is the engine an enterprise licenses. The fuel is now billed separately.”

Gogia added that vendors are structuring pricing differently around that model. “OpenAI has kept a fixed seat and meters what sits above it. Microsoft has kept a $30 Copilot licence carrying no entitlement to its agentic layer at all. Google, meanwhile, has left its seat editions alone while switching agent meters on by published date.”

Greyhound Research pointed to the treatment of usage caps as a signal of how capacity is being priced.

“Most will read this as segmentation. Greyhound Research reads it as rationing, and the five-hour limit is the tell,” Gogia said. “On July 12, OpenAI removed that limit free of charge. At the end of July, it was restored. On August 10, it became a paid entitlement.”

CIOs face allocation and cost decisions

The higher-priced tier raises questions about how enterprises assign access, analysts said.

“Don’t start with people. Start with your billing data,” Chopra said. “If someone is regularly drawing down shared workspace credits, a fixed higher seat may cost less and forecast better than open-ended credit consumption.”

He said allocation based on hierarchy can lead to inefficiencies. “You end up paying premium rates for executives who open it twice a week while the analyst-blocked mid-close stays throttled.”

Gogia echoed that view. “Power user should describe a workload pattern, not a job title,” he said. “The right question is whose work becomes materially more valuable when Standard stops being enough.”

Credits positioned to drive metered usage

According to the statement, OpenAI is offering incentives, stating that eligible customers can receive “$100 worth of workspace credits (2,500 credits) for each Premium seat they add, up to 5 seats.”

Analysts said such credits are tied to usage-based pricing. “Worth being clear that this isn’t a discount. Credits are currency inside OpenAI’s metered layer,” Chopra said. “Free credits get a workspace comfortable for consuming metered features,” Gogia said the offer should be treated as a pilot incentive. “The promotion should fund the pilot. It should not shorten the runway,” Gogia said, adding that billing controls, including pooled credits and auto-recharge settings, require close oversight.

Beyond chatbots: How embedded GenAI is transforming banking application development

Business application development is entering a new operating model. The traditional approach of gathering requirements, designing screens, writing services, integrating systems, testing, fixing defects and preparing release documentation still exists, but it is no longer sufficient for enterprises that need speed, traceability, resilience and regulatory confidence at the same time. Hyperautomation brings a broader discipline to this challenge. It combines workflow orchestration, intelligent document processing, robotic automation, API-led integration, process mining, test automation, observability and artificial intelligence into a connected delivery fabric. With embedded Generative AI, this fabric becomes more adaptive because applications can interpret natural language, summarize complex data, generate explanations, detect exceptions and support decision workflows rather than merely execute predefined rules.

In banking, this shift is especially meaningful. Banks operate across dense application landscapes: trade reporting platforms, wealth management portals, core banking systems, investment banking applications, digital compliance engines, reconciliation utilities, operational dashboards, audit repositories and daily, weekly and monthly reporting platforms. Each of these areas has its own data models, control points, integration patterns, validation rules, exception paths and regulatory obligations. Hyperautomation does not replace engineering discipline; it strengthens it by making business intent, technical execution, control evidence and continuous improvement part of the same lifecycle.

From automation to hyperautomation in banking applications

Automation usually addresses a specific task: moving data from one system to another, generating a report, running a batch job or validating a transaction against a rule. Hyperautomation goes further. It looks at the complete business outcome and asks how the entire chain can be streamlined, governed, observed and improved. For example, a trade reporting process may begin with transaction capture, enrich the trade with reference data, validate regulatory fields, identify breaks, generate a submission file, transmit it to a regulator or trade repository, monitor acknowledgements and preserve audit evidence. A narrow automation script may accelerate one step, but a hyperautomated design coordinates the complete flow, including exception handling and evidence generation.

Figure: Automation vs. hyperautomation.

Magesh Kasthuri

Figure: Automation vs. hyperautomation

Embedded Generative AI adds a new layer of intelligence. Instead of forcing every user interaction into rigid screens and codes, business applications can accept natural language prompts, interpret document content, summarize cases, generate draft responses, explain anomalies, produce test scenarios and create release notes. In a banking environment, this intelligence must be carefully bounded. Every AI-assisted action should be traceable, explainable, reviewable and aligned with data privacy, model risk, information security and regulatory expectations. The goal is not uncontrolled autonomy; the goal is governed acceleration.

Banking application components suitable for hyperautomation

A modern banking application is rarely a single monolithic system. It is a composition of business capabilities, integration services, workflow engines, data pipelines, user experience layers, analytics models, control dashboards and audit stores. Hyperautomation can accelerate the development and integration of these components by turning repetitive engineering work into reusable patterns and by embedding intelligence directly into business processes.

  • Trade reporting applications: Generative AI can help map trade attributes to regulatory fields, explain validation failures, summarize rejected submissions and generate test cases for reporting scenarios. Hyperautomation can orchestrate enrichment, validation, submission, acknowledgement tracking and evidence archival.
  • Wealth management platforms: Advisors can use embedded AI to summarize client portfolios, generate suitability narratives, identify missing documents and prepare personalized investment review notes. Automation can coordinate onboarding, risk profiling, document verification, portfolio rebalancing workflows and client communication approvals.
  • Core banking applications: Account opening, loan servicing, deposits, payments, interest calculations and customer maintenance can benefit from automated validations, intelligent forms, workflow routing and natural language assistance for operations teams. AI can explain account events or transaction exceptions in plain language.
  • Investment banking systems: Deal pipelines, research workflows, underwriting processes, trade lifecycle functions and risk calculations require strong coordination across front-office, middle-office and back-office platforms. Hyperautomation can standardize approvals, documentation, exception resolution and control evidence across these stages.
  • Digital compliance applications: Compliance teams can use AI to summarize policy obligations, compare regulatory changes with internal controls, classify alerts, draft investigation notes and produce evidence packs. Automation ensures routing, approvals, segregation of duties, audit trails and regulatory reporting timelines are consistently enforced.
  • Reconciliation platforms: AI can assist in matching narratives, explaining breaks, clustering exception patterns and suggesting resolution actions. Hyperautomation can pull data from ledgers, statements, payment processors, trading systems and data warehouses, then route unresolved breaks to the right teams.
  • Reporting and audit applications: Daily, weekly and monthly reports can be generated through controlled data pipelines, automated quality checks, narrative generation, variance explanations and approval workflows. Audit applications can preserve lineage, approvals, source extracts, model outputs and control attestations.

Embedded generative AI as an application capability

Embedding Generative AI into business applications should be treated as an architectural capability, not as a decorative chatbot. A banking application may use AI for search, summarization, reasoning support, content generation, code generation, policy interpretation or anomaly explanation. Each use case requires clear boundaries. The application must know which data the model can access, which actions require approval, what evidence must be captured and where deterministic controls must override probabilistic suggestions.

For example, in trade reporting, an embedded AI assistant can explain why a transaction failed validation and suggest likely fields to review. However, the final correction should pass through rule-based validations, maker-checker approval and audit logging. In wealth management, AI may draft a client review note based on portfolio movements and risk profile, but the advisor must verify suitability, disclosures and final communication. In reconciliation, AI can propose likely matches or categorize break reasons, while the system preserves the original data, confidence score, reviewer action and final resolution path.

Hyperautomating the product development lifecycle

The Product Development Lifecycle can itself become hyperautomated. Instead of treating ideation, analysis, design, development, testing, security review, release and operations as disconnected phases, enterprises can create an AI-assisted delivery loop where every stage produces structured artifacts that the next stage can consume. Platforms such as GitHub Copilot, Claude Code or Claude Cowork-style agentic development environments and OpenAI Codex can support this movement by helping teams reason over requirements, generate code, create tests, review changes, modernize legacy modules and produce documentation. Their value increases when they are connected to repositories, issue trackers, design documents, build pipelines, test suites, security scanners, observability data and enterprise knowledge bases.

PDLC StageHyperautomation OpportunityAI-Assisted Outcome
Business discoveryProcess mining, domain interviews, regulatory mapping, backlog creationStructured epics, user stories, acceptance criteria, process maps and control requirements
Architecture and designReference architectures, API contracts, data models, event flows, security patternsArchitecture options, integration blueprints, threat-model prompts and design decision records
DevelopmentCode generation, service scaffolding, UI component creation, data pipeline templatesReview-ready code increments, reusable components, migration utilities and integration adapters
TestingUnit, integration, regression, performance, compliance and synthetic data testingGenerated test cases, defect reproduction steps, test automation scripts and coverage summaries
Security and compliance reviewStatic analysis, dependency checks, policy validation, evidence captureRisk explanations, remediation suggestions, control traceability and approval evidence
Release and deploymentCI/CD orchestration, environment promotion, release notes, rollback preparationAutomated deployment packs, release summaries, operational checklists and change records
Operations and feedbackObservability, incident analysis, user feedback mining, backlog refinementIncident summaries, root-cause hypotheses, improvement stories and reliability recommendations

Role of GitHub Copilot, Claude Cowork and Codex

GitHub Copilot is useful where developers need assistance inside the engineering flow: explaining code, generating functions, proposing tests, reviewing pull requests and helping teams move from issue to implementation. In a banking PDLC, it can accelerate microservice creation, API integration, batch processing logic, reconciliation rules, regulatory validation routines and UI workflows. When used with repository context and proper review discipline, it can reduce the time developers spend on repetitive coding while preserving human accountability for design and correctness.

Claude Cowork or Claude Code-style agentic environments are valuable for multi-file reasoning, refactoring, debugging and documentation-heavy engineering work. Banking applications often contain deep domain logic scattered across services, configuration files, stored procedures, integration scripts and test suites. An agentic coding assistant that can understand a wider codebase context can help engineers analyze dependencies, prepare modernization plans, update multiple files coherently and draft explanations for reviewers. This is particularly useful in core banking modernization, trade reporting rule updates and compliance workflow refactoring.

OpenAI Codex can support issue-to-pull-request workflows, test generation, code review, bug reproduction, migration activities and broader software engineering tasks across the lifecycle. In a hyperautomated PDLC, Codex-like agents can be assigned well-scoped work items, asked to inspect failing tests, propose fixes, create regression coverage and summarize the change for human reviewers. The important design principle is to keep agents inside controlled boundaries: clear prompts, repository permissions, test gates, approval workflows and traceable outputs.

Integration architecture for hyperautomated banking applications

A practical architecture begins with business capability decomposition. Each banking domain should be expressed as a set of bounded capabilities such as customer onboarding, account maintenance, trade enrichment, exception management, portfolio review, control attestation, report generation and audit retrieval. These capabilities should be exposed through APIs, events, workflow tasks, data products and user interfaces. Hyperautomation then connects these capabilities using orchestration engines, event streams, rules engines, AI services, RPA connectors where legacy integration is unavoidable and observability layers that capture business and technical telemetry.

The embedded AI layer should sit behind a secure application service boundary. It should use retrieval-augmented generation where approved policies, product rules, application documentation and regulatory mappings are retrieved from trusted sources. It should avoid uncontrolled exposure of sensitive customer information. Prompt templates, response validation, redaction, grounding checks, model monitoring and human-in-the-loop approval should be part of the production design. In banking, the most successful AI pattern is often not full automation but assisted decisioning with strong controls.

Example: Hyperautomated reconciliation and reporting flow

Consider a reconciliation application that compares ledger balances, payment files, trade settlement records and external statements. In a conventional model, operations teams spend significant time downloading files, running macros, investigating mismatches, documenting break reasons and preparing status reports. In a hyperautomated model, data ingestion is scheduled and monitored, schema checks run automatically, matching engines classify obvious matches, AI assists with ambiguous narratives, exceptions are routed through workflow queues and dashboards update in near real time. At the end of the day, the system can generate a draft operations report explaining unresolved breaks, aging trends, risk exposure and pending approvals.

The same pattern can extend to daily, weekly and monthly reporting. Data quality rules validate inputs, report templates are populated automatically, AI generates narrative commentary on variances, reviewers approve or amend explanations and the final report is archived with lineage and approvals. Audit teams can later retrieve not only the report but also the source extracts, transformation logs, exception history, reviewer decisions and AI-generated drafts. This creates a richer control environment than manual reporting because evidence is captured by design rather than reconstructed later.

Governance, risk and control considerations

Hyperautomation in banking must be designed with governance from the beginning. The development team should define which activities can be automated, which can be AI-assisted and which must remain under human approval. Source code generated by AI must pass normal engineering controls, including peer review, static analysis, dependency scanning, secure coding checks, test execution and production readiness review. Business outputs generated by AI, such as compliance narratives or client-facing explanations, should be reviewed where regulatory or reputational risk is material.

Data governance is equally important. AI-enabled applications must respect data classification, residency, retention, masking and access policies. The model should not become an uncontrolled channel through which confidential customer, trading or employee information can leak. Every prompt, retrieved source, generated response, user action and final decision may need to be logged depending on the use case. For audit applications, this traceability is not optional; it is the foundation of trust.

Operating model for AI-native PDLC

A hyperautomated PDLC requires changes in team behavior. Product owners should write requirements in a structured manner so that AI tools can generate better stories, acceptance criteria and test scenarios. Architects should maintain living decision records, reference patterns and integration standards that AI agents can use as context. Developers should learn prompt discipline, context packaging and review techniques. Test engineers should focus on coverage strategy, synthetic data, compliance scenarios and defect prevention rather than only manual execution. Operations teams should feed incident learnings back into the backlog so the system improves continuously.

The role of human experts becomes more important, not less. AI can draft, generate, compare and suggest, but domain judgment remains essential. A trade reporting specialist understands regulatory nuance. A wealth advisor understands client suitability. A core banking architect understands transaction integrity. A compliance officer understands control interpretation. Hyperautomation works best when it amplifies these experts and removes repetitive friction around them.

Conclusion

Hyperautomation in business application development is not simply a faster way to write software. It is a new way to connect business intent, engineering execution, operational control and continuous learning. In banking, where applications must be reliable, explainable, secure and compliant, the combination of embedded Generative AI and disciplined automation can transform how applications are designed, built, integrated, tested, released and operated. Trade reporting, wealth management, core banking, investment banking, compliance, reconciliation, reporting and audit functions can all benefit when AI is embedded responsibly and automation is orchestrated across the complete lifecycle.

Platforms such as GitHub Copilot, Claude Cowork or Claude Code and OpenAI Codex can play an important role in this transformation by accelerating analysis, development, testing, review, modernization and documentation. Their greatest value appears when enterprises treat them not as isolated productivity tools but as part of a governed, AI-native PDLC. The future of banking application development will belong to teams that can combine human expertise, reusable engineering patterns, intelligent automation and strong governance into one coherent delivery model.

This article was made possible by our partnership with the IASA Chief Architect Forum. The CAF’s purpose is to test, challenge and support the art and science of Business Technology Architecture and its evolution over time as well as grow the influence and leadership of chief architects both inside and outside the profession. The CAF is a leadership community of the IASA, the leading non-profit professional association for business technology architects.

Don’t let your company be fooled by AI efficiency

The scenario isn’t hypothetical: Some of the companies that went furthest in replacing people with AI have had to backtrack.

For example, in 2024 Klarna became a European benchmark for what AI could do for a company. Its AI assistant handled two-thirds of customer service chats in its first month, performing the equivalent of 700 full-time agents. As a result, company leadership decided to freeze hiring, and the workforce shrank from around 5,000 to 3,800 employees.

Just a year later, Klarna’s CEO admitted the company had gone too far in replacing people with agents, which had negatively impacted both the service and the product. In fact, the company reversed courserehiring human agents to ensure customers could always speak to a person.

The interesting point here isn’t that AI failed. The problem was something else: understanding the customer service function solely in terms of productivity and costs, without considering the bigger picture.

If measured by response times and equivalent FTEs, automation was optimal. Measured by satisfaction, perceived quality, and the ability to resolve complex cases, the result was different — and ultimately forced a reversal.

For CIOs, this disconnect presents a leadership opportunity: Management and other departments need precisely the comprehensive technical and business process perspective CIOs can bring to the table.

AI is redesigning how a function is delivered

It’s tempting to read Klarna’s AI journey (and back) as a customer service story. But the pattern affects every business function. Introducing AI agents isn’t just adding another tool: It reshapes decision-making, day-to-day learning, and ultimately, how service is delivered.

If you only think in terms of productivity (what’s automated, how much is saved, how many equivalent FTEs are freed up), it’s easy to lose sight of the deeper implications. It’s easy to discover too late that what’s being delivered is no longer the same, even if on paper more is being produced.

This is difficult to see at first. A function can perform worse and still show better operational metrics for months. The consequences appear in other areas, far removed from the automated function: in reputation, lost customers, or poor decisions.

CIOs see this pattern earlier and more strongly. When an agent used by IT — often among the earliest adopters — ceases to be a helpful assistant, the changes have quick and significant impact. They influence which alerts reach the operations team, which code modifications are proposed to developers, which incidents are prioritized by security personnel. This goes beyond simply speeding up work: It determines what the team sees and doesn’t see, and it shifts the decision-making environment.

Agents don’t just execute. They change how they detect problems, how they respond, and even how they learn. If this phenomenon is evaluated solely with performance metrics, it runs the exact same risk Klarna faced internally: gaining speed and losing perspective.

The paradox: More capacity for action, less direct vision

Many IT managers are beginning to notice the paradox inherent in AI agent use. The organization can act faster, deliver more volume, and automate more decisions, but at the same time lose touch with the complexity of reality.

Previously, a support team learned not only by resolving incidents, but also by identifying where integrations failed or what user behaviors revealed a deeper problem. If that work is now automated, the organization can continue to resolve issues, but employees lose valuable learning opportunities.

The risk the team faces is that AI will work well enough to push knowledge and capabilities about how a business unit should operate out of the foreground.

The CIO opportunity

This is where the CIO’s role needs to change. CIOs must move beyond being those who simply automate processes to become those who provide, both within and outside their department, a comprehensive understanding of how AI impacts a business function. This means going beyond productivity gains and contributing other, less visible aspects, such as enhanced experience, business perspective, and changes in service delivery, whether for employees or customers.

This perspective is invaluable both at the senior management level and in other areas such as operations, customer service, and, of course, human resources. In the current climate, with its constant announcements of workforce reductions, the conversation tends to focus on cost and time savings. The CIO is well-positioned to provide the other side of the coin: where strong oversight is necessary, what can be delegated to AI, and where it’s essential to plan for the reversal of automation that, on paper, appears to be working.

That ability to recover is, in fact, one that the organization cannot afford to lose. Not all organizations can regain capabilities as quickly as they are lost.

Your mission: To present a clear-eyed view of AI’s business impact

The CIO’s mission, therefore, is to help clarify what can be delegated to AI and what should not be relinquished without losing the capacity to intervene. In some cases, the answer will be clear: repetitive tasks, initial classification, draft generation, or technical searches. In others, the boundary may be more delicate: prioritizing risks, deciding on exceptions, changing legacy systems, or acting on processes without sufficient oversight.

This will be one of the most important services in the CIO’s role over the next few years. Beyond advancing the adoption of agents, they will have to provide, both within and outside of IT, the necessary understanding of the impact of agents on a business function. And, finally, they must retain the ability to reverse course when the expected results aren’t being delivered, no matter how good the metrics look.

The gen AI helping Aetna review millions of medical records

One of the biggest challenges companies like Aetna face every year is an annual HEDIS review of its records to identify gaps in care. For large national payors, the scale of the challenge is immense. So Aetna has deployed a gen AI-driven document intelligence platform that has reduced the need for manual review by 65%.

“We have a large group of amazing trained medical coders who do this every day,” says Nathan Frank, chief digital and technology officer at Aetna. “This is about making it easier for them by speeding up the process. Something that might have taken weeks or months we can now do in days.”

The Healthcare Effectiveness Data and Information Set (HEDIS) is a range of performance measures for the managed care industry. Developed and maintained by the nonprofit National Committee for Quality Assurance (NCQA), the first version of HEDIS was released in 1991.

Under the HEDIS measures, large managed care providers like Aetna review more than 10 million medical records annually to identify gaps in care. These gaps are missed or overdue preventative care or chronic disease management tests including missed cancer screenings, blood sugar tests for diabetics, eye exams, and immunizations. Closing these gaps improves patient outcomes, and health plans are measured in how well they perform. But processing medical records is no easy task.

“We’re talking about medical charts that have white space filled with handwritten notes,” Frank explains. It’s not just structured data, it’s lots of physical clinical documentation.”

Adding up the numbers

Frank says industry benchmarks for large providers indicate an annual review process that requires about 50,000 work weeks, equivalent to nearly 1,000 dedicated full-time employees. It would take a team of 50 reviewers more than 20 years to complete a single annual review using fully manual processes.

Enter AI Medical Chart Review, a platform developed by Aetna that leverages cloud services and gen AI to automatically extract clinically relevant data from records, and prioritize records based on the likelihood of measure closure and evidence strength.

“Large language models and gen AI give us the ability to train a model to decipher the charts, identify the high value codes, and build correlations,” Frank says.

In the space of about six months, Frank’s team ideated the platform, and designed and trained a PoC that was able to process millions of records in just two weeks. As a result, AI Medical Chart Review has earned Aetna a CIO 100 Award in IT innovation.

“Now we’ve gone through 14 million documents,” Franks says. “We’re seeing a reduction of manual effort, which is now being transitioned into other areas like quality control and making sure the automated chart review is working as expected.”

Behind the curtain

Using gen AI, the platform automatically ingests and analyzes unstructured medical records and clinical documents. And as part of that process, it identifies and extracts clinically relevant information for specific HEDIS measures like diagnosis codes, medication records, lab results, and visit documentation. With this data, the platform generates a prioritized set of records based on the likelihood of measure closure and strength of clinical evidence, which is then passed to human employees for review and validation.

Frank says the platform has increased gap closure rates (leading to improved Star Ratings and higher reimbursement), streamlined workflows, and enabled teams once dedicated to manual record review to shift focus to higher-value activities.

Frank says much of the speed and success in building the platform comes down to a shift in the way it approached the design and build process. Rather than exhaustively writing specifications and requirements, Aetna created a team that included engineers and subject matter experts who worked together to build out capabilities iteratively.

“It allowed us to move much faster, and having a business subject matter expert sitting in the same virtual or physical room with us got us a much better outcome,” Frank says. “The product model, our cloud compute model, and our AI governance model allow for quick reviews to make sure we’re using AI responsibly with the right guardrails. It’s increased the speed to get from product launch to go-live.”

He adds that small teams that don’t have to deal with a lot of bureaucracy are key to moving quickly.

“You need to design with security, compliance, and a responsible use of AI as core principles from day one,” he says. “Everything we do from a new build standpoint starts with thinking about how we make it cloud native, how we build with the right elasticity and speed, and how we optimize the cost.”

The most important element of all, he says, is a good relationship with your subject matter experts.

“You can have a great product manager and engineer, but you really need that business subject matter expert who’s excited about it, and who has a passion for transforming the process,” Frank says. “Once you put those three together, you’ll see amazing things like this happen all the time.”

How CaixaBank drives partner and customer relationships through AI

The transformation of the financial sector is no longer just about offering a mobile app or allowing customers to bank from anywhere. After years of digitizing services, institutions now face the more ambitious challenge of building a more personalized, agile, and intelligent relationship with millions of users who expect immediate answers, simple experiences, and service tailored to specific needs. 

The emergence of gen AI has accelerated this evolution. While banks have used AI models for years to automate processes, improve efficiency, and analyze large volumes of data, a new generation of conversational tools opens the door to a much more natural interaction between customers and financial institutions. 

Spain’s CaixaBank, for example, has positioned AI as one of the cornerstones of its technological transformation. The bank, which has more than 12 million digital users, believes this change isn’t solely due to tech’s evolution, but also to a shift in user expectations.

“Today’s customer is more digital, autonomous, and also more demanding in their relationship with the bank,” says Mariona Vicens, CaixaBank’s director of digital transformation and advanced analytics. “They not only interact more through digital channels, but also expect simplicity and personalized solutions at any time and from any device.”

A history of AI experience

Although gen AI has made a big impact, CaixaBank says its commitment to these technologies began much earlier. But it now represents a qualitative leap. “It’s more focused on developing new models based on conversational applications,” she says. “The most visible improvement is that gen AI allows for more natural, contextual, and useful interactions for both employees and customers.” 

This evolution is part of CaixaBank’s 2025-2027 Strategic Plan, in which it identifies agility, new services, efficiency, and technological resilience as main and interconnected objectives. For Vicens, agility is particularly key. “It’s what allows us to respond to a customer who increasingly expects immediacy, and it’s also what determines the bank’s ability to adapt in an increasingly dynamic environment.” 

The Cosmos Plan, the specific roadmap for processes and technology framed within CaixaBank’s strategic plan, reflects this integrated vision. “It combines investment in technology, automation, and AI to enable a more flexible and efficient organization capable of evolving at the pace set by customers,” she says. “Ultimately, agility is the visible engine of change, but it’s only possible when all elements of the model advance in a coordinated manner.” 

AI is certainly at the forefront of how the bank operates. More than 2,000 employees are already using agents to automate tasks, streamline processes, and improve customer service — a number the bank expects to increase before the year’s end. “With this implementation, combined with the application of other models like gen AI integrated into office tools, we expect to scale the gains in productivity and agility,” Vicens says.  

Innovation with human oversight 

While AI opens up new possibilities for transforming customer relationships, it also presents challenges related to regulation, transparency, and trust. For CaixaBank, innovation isn’t just about developing new use cases, but doing so under a governance model that ensures the technology is used responsibly. 

With that objective, the bank has defined a specific governance framework for these tools, with a corporate-level AI Office and a policy that anchors principles such as transparency and explainability, data fairness and privacy, robustness and security, and human oversight. 

This framework, CaixaBank explains, translates into concrete controls throughout the entire AI lifecycle: prior validation of use cases, structured risk assessment before implementation, corporate inventory of systems, subsequent monitoring, and incident management. However, it’s all based on the clear premise that relevant decisions can’t be entirely delegated to AI, so they must maintain human oversight. 

Regulation for confident innovation 

The entry of the EU AI Act has placed financial institutions under evolving regulatory requirements. Far from seeing it as an obstacle, CaixaBank believes this framework fits perfectly with how it’s approached the tech all along. “It fits naturally, because we’re precisely structuring our AI governance model with this framework and other regulatory frameworks as a reference, and we integrate it into the AI ​​lifecycle from the design stage and by default,” says Vicens.

Corporate policy explicitly incorporates the regulations into its global risk management system. In practice, any AI-based application must follow a clearly defined process before being implemented. “This means that any use of AI must be identified, evaluated, and monitored,” she says. “Before developing a use case, its type, value, and feasibility are validated, and then its risks are assessed. And once implemented, its performance is monitored.” Of course, in a financial environment, customer trust remains a most valuable asset.

Added AI agent muscle 

All this transformation strategy is finding a tangible application in one particular development: a contracting assistant that accompanies the client through digital channels. The system acts as a first point of contact when a user requests information about a product from the CaixaBank website or app. From there, it can answer questions, provide contextual information, guide the conversation, and, when necessary, transfer the interaction to a specialist without the customer having to restart the process. 

For the bank, this ability to understand context is a key differentiator. “Unlike a chatbot that answers a collection of FAQs, this agent is a contracting assistant that understands the context of the conversation with the customer, provides support, and can escalate to a human,” she says. For products like pre-approved loans, it can even lead the conversation to the final step before closing. 

The bank emphasizes that human intervention remains an essential part of the process. “We see AI as a tool to inform, streamline, and support the customer to enhance their user experience in a way that complements the ongoing support provided by our team of specialized remote banking managers,” Vicens adds. Plus, customers can choose to speak with a human from the outset or at any point during the conversation, and the final contract is always signed with the assistance of a CaixaBank specialist.  

Great responsibility

Beyond human oversight, the bank has established a framework to ensure the responsible use of AI. “It has defined responsible AI principles that cover the entire lifecycle of developments to ensure fair, transparent, responsible use, aligned with legislation and the group’s values,” she says. “Before deploying any AI solution aimed at customers, compliance with these principles is verified.”

In the specific case of the contracting assistant, data protection is one of the essential elements. The information travels encrypted, and the model isn’t trained with the data sent to the LLM. 

Currently, this technology is available in 40 products and manages an average of 6,000 conversations per month — figures that, according to CaixaBank, provide clear metrics of scale and productivity. 

The implementation of the onboarding assistant is one example of a much broader strategy in which AI, data, and automation are used to transform the relationship between the bank and its customers. “The key is no longer just being available, but providing real value in every interaction, and strengthening trust through useful experiences tailored to each user,” says Vicens.

❌