Visualização de leitura

What JPMorgan does differently with AI that any company can apply


In the summer of 2024, JPMorgan Chase deployed its internal AI platform LLM Suite, launching it very differently than most do: The company didn’t force anyone to use it.

When LLM Suite arrived at its first major division, asset and wealth management, employees were asked to think of it as a research analyst: someone to ask for data, a draft, or an idea. Leadership didn’t set usage objectives or provide a formal mandate.

Access was rolled out in phases and only to those who requested it, and the bank allowed the tool to circulate through word-of-mouth recommendations among colleagues. While half the industry rushed to count users and publish adoption rates, JPMorgan gave up on pursuing that number.

It became flooded with users. In eight months, 200,000 employees had signed up without a single order being issued, out of a workforce of over 300,000. In time, the bank established more than 450 use cases in production.

Two years after that summer launch, JPMorgan had everything to boast about. It had established itself as a global leader in the use of AI: It was the top bank on the Fortune AIQ 50 list, and the third company overall, ahead of all the tech giants except Alphabet.

It was then that the bank’s head of analytics, Derek Waldron, the person best positioned to sell the success, pointed out what still wasn’t working: There was a gap between what the technology was capable of doing and what the bank was actually capturing in its business results.

That gesture is what distinguishes JPMorgan. Although it has much to celebrate, it knows what it lacks, it says so publicly, and it keeps searching for it. Behind that statement lies a way of innovating and measuring that the bank has been developing for years.

Giving up the number everyone was chasing

The first thing JPMorgan did right was not to make adoption the goal. By not forcing anyone, it turned platform usage into a barometer. If a tool worked, it was filled without any campaign; if it didn’t, it was emptied, and that emptiness provided valuable information. If adoption had become a target to be pursued, the organization would have optimized the number instead of understanding what the number represents.

The bank itself acknowledges that if a tool is broadly used, it means it’s popular, but not necessarily effective. To determine its effectiveness, something more was needed. The answer came from two decisions that only work together: linking each project to a business outcome, and creating the metrics to demonstrate that outcome.

First, to find initiatives that could have a real impact, instead of creating an agenda from the top down, the bank surveyed its business units, asking where there was a problem to solve. Within a few weeks, an internal portal gathered, according to the bank’s figures, nearly a thousand ideas. Of these, only a few hundred moved forward and reached production. An organization doesn’t open a funnel of that size if it expects most ideas to survive; it anticipates that many will be discarded.

The funnel’s filtering method was also different. Before launching each test, the outcome that would ensure the experiment’s survival was defined, along with the steps to be taken the day after the decision. By planning future actions in advance, indecision and the perception of failure were avoided.

A clinical approach to AI experimentation

But setting a threshold for each experiment requires verification, and that’s where the bank encountered an unexpected obstacle. Metrics have their own cycles. Bank customers conduct business on Mondays, not Sundays. They receive their paychecks at the end of the month. In August, they disappear. When an initiative generates a change and a figure rises the following week, there’s no way to know whether it increased due to the change or the calendar.

The solution was borrowed from clinical trials. Instead of rolling out the change to all users, it was rolled out to half, chosen at random. The other half (the control group) operated on the same Monday, the same payroll, and the same August, so that the experimental contribution (the attribution) could be separated.

The next step was to industrialize the experiments. Doing it properly required a specialist sitting alongside each product team, and with that method, they reached eight per year. A self-service platform increased the figure to around 300 tests annually.

The results are concrete. For example, tens of thousands of the bank’s engineers have gained between 10% and 20% efficiency thanks to an internally developed programming assistant.

Finally, the bank discovered that a figure can be accurate and yet mean nothing. Its head of analytics explained this with a simple example. They measure the hour that AI saves one employee, and the three hours it saves another. They add them up, and the result is accurate. But in a process that goes from beginning to end, those saved minutes often don’t appear on the bottom line: They merely shift the bottleneck to the next one.

It’s easy to get stuck on partial metrics because they’re more immediate and produce more impressive numbers. JPMorgan’s discipline consisted of not accepting a metric as valid until verifying its impact on the business at the end of the process.

The question then remains on Monday morning: What can a company that has neither the size nor the budget of a bank take home?

The method is what best exports

What’s most interesting about JPMorgan isn’t what it has done with AI, but how it has done it . Any company can replicate this approach, because it doesn’t depend on proprietary data, scale, or budget.

The following are some best practices that don’t require a €20 billion annual budget. They do require making decisions before starting and are within reach of any company:

Launch far more initiatives than will survive, and announce this clearly. If the organization discovers halfway through that most of its projects will be canceled, it may misinterpret this as a planning failure; if it knows from the outset, it understands it as the natural selection process. This is what makes making mistakes quick and cheap.

Decide in advance the threshold that will shut down a project and plan the next steps. Both aspects are necessary, not just the metric. If a certain figure isn’t reached, the team needs to know what will happen next. Applying a threshold without future planning leaves the team in limbo, and they’ll have to find a reasonable reason to wait another quarter before shutting down.

Work on business outcome metrics from the outset, not just when they’re requested. This tracking not only guides the initiative but also prevents having to reconstruct months of poorly documented decisions. Adoption, by the way, is the number the CFO won’t ask for. It serves as a signal while no one is pursuing it, and it ceases to be useful the day it becomes a target.

How to get it right

Whether metrics mean anything depends on where you focus your attention. It’s best to start with scope, because that’s the most common mistake. Saving three hours in one stage isn’t the same as improving time-to-market: If the entire process isn’t shortened, what you have is freed-up capacity, which is also valuable, but it’s something different, and it’s advisable to make that distinction clear.

Then it’s important to consider that value leakage occurs in two directions. The first is outward: The savings are passed on to the customer in the form of lower prices or better service. The second is inward: The savings in personnel are replaced by spending on computing. If these items fall into different budget categories, it’s easy to overestimate the actual savings.

Finally, there’s an excessive focus on cost savings, at the expense of revenue opportunities. Jamie Dimon, CEO of JPMorgan, put it more bluntly to his analysts than any consulting firm: No one benefits uniquely from AI. In other words, competitors will eventually incorporate those savings. The greatest potential for differentiation lies in revenue: using AI to uncover unmet demand.

The question a CIO will have to answer in a year’s time won’t be how much AI their company uses. It will be which of projects are still alive because they work, and not because no one has bothered to test them.

How IT can scale self-service without losing control

Every IT leader knows the pattern. One team builds a report in a spreadsheet. Another spins up a workflow with slightly different logic to answer the same question. A dashboard is shared across three departments, and within a week, nobody can say for certain where the underlying numbers came from. 

Instead of freeing up capacity, self-service has quietly become another form of manual work: chasing down mystery logic, reconciling duplicated effort, and answering questions nobody wants to own. 

Self-service was never the risk 

It is tempting to read that scenario as an argument for tighter control — fewer people building, more requests routed through a central team, more approvals before anything ships. That reaction is understandable, but it solves the wrong problem. 

Self-service fails when there are no shared rules for access, quality, documentation, and ownership. Without those guardrails, speed doesn’t produce faster decisions, just more confusion distributed across spreadsheets and more shared drives. 

The real tension is that most organizations have been offered only two options. Either lock everything down, or let everyone build whatever they want and hope it holds together. Neither one scales. 

IT as the paved road, not the checkpoint 

Centralizing data was never the hard part. The real challenge is the last mile: turning that data into decisions and actions the business can actually trust. Closing that gap does not mean IT owns every rule, calculation, and exception that determines how work gets done. 

It means IT builds the paved road — trusted access, approved workflows, reusable templates, and visibility into what is being built — while the people closest to the work own and adapt the business logic that runs through it. 

That division of labor changes what “governance” means in practice. Instead of a gate every request has to pass through one at a time, governance becomes the infrastructure that keeps logic visible, understandable, repeatable, and auditable by design. When the fastest way to answer a question is also the most trusted way, analysts do not need to be talked into compliance, and IT does not need to inspect every workflow to know it will hold up. It is simply how the work gets done. 

Freedom and guardrails, together 

Governed self-service isn’t about choosing between speed and control, it’s about giving each side of the equation what it actually needs to trust the other. 

Governed self-service gives analysts: 

  • Access to trusted data 
  • Reusable templates and workflow patterns 
  • Clear rules for sharing and automation 
  • A way to document logic 
  • Support when a workflow needs to scale 

And it gives IT: 

  • Visibility into who is building what 
  • Better governance over access and data use 
  • Fewer one-off requests 
  • Less mystery logic floating around the business 
  • A cleaner path from individual workflow to team-wide process 

What this looks like in practice 

Papa Johns’ finance team offers a useful example of governed self-service in action. The team handles risk-sensitive, high-volume work — franchise billing, royalty calculations, aggregator commissions, and SOX-compliant period close — across a global, multi-currency franchise business. 

Historically, much of that logic lived in spreadsheets and disconnected tools, separate from the systems of record and hard to audit when workflows changed. 

Using Alteryx, Papa Johns rebuilt franchise billing and reconciliation as a governed workflow that runs directly against its Google BigQuery environment, so calculations execute where the data already lives rather than being copied out to another location. 

With Alteryx, complex calculations are visible, repeatable, and auditable. Finance users can ask natural language questions, such as comparing month-over-month figures, and receive immediate answers while also seeing how logic is applied. IT can support governance without becoming a bottleneck. 

The partnership between the business and IT was key to scaling success. Michael Wyant, VP of Enterprise Data and Corporate Solutions, and his team are responsible for governance and data pipelines. The finance team owns the business logic and can adapt it as requirements change. 

Each side owns the part of the problem it understands best. 

The result is a workflow that finance trusts, that IT can stand behind, and that scales as a template for other high-stakes processes across the business. That’s the kind of outcome that governed self-service is meant to produce. 

Fewer surprises, more trust 

None of this requires IT to slow analysts down or analysts to work around IT. When self-service is built on shared standards, analysts stop waiting on tickets, IT stops chasing down mystery logic, and the business gets answers that hold up the moment someone asks, “Where did this number come from?” 

Alteryx supports that model by giving business teams a governed way to build and adapt workflows themselves, while giving IT the visibility, controls, and security required to support it all at enterprise scale. 

The goal was never more control for control’s sake. It is fewer surprises, less rework, and more answers the business can actually trust. 

Ready to see what governed self-service could look like for your team? Explore the AI-Ready Starter Kits to get started. 

 To learn more, visit us here

The logic layer: the missing piece in modern AI tech stacks

There’s a scenario that plays out every day across the enterprise. A salesperson is about to close a major deal. They want to know what their commission will be. They type the question into ChatGPT or their favorite AI assistant. What comes back is a thoughtful, well-written explanation of how software companies typically structure sales compensation. 

The one thing it won’t tell them is what their commission will actually be if they close this specific deal. That gap between what AI can reason over and what it knows about your business is the defining challenge of enterprise AI adoption right now. 

I call it the logic layer. And without it, AI gives you impressive sounding outputs that are often disconnected from how your business runs. 

Why business logic lives with the analyst 

One of the more persistent myths in AI is that analysts are on the verge of becoming unnecessary. 

The reality is the opposite, and the logic layer is exactly why. 

In an AI-enabled enterprise, analysts become more essential because they are closest to the logic and context that governs the business. They know which definition of pipeline matters and which edge cases matter in audit, merchandising, finance, or marketing. 

I believe enterprises that succeed in the AI era will not be defined by how much AI they deploy but whether the people who understand the business own and control the intelligence that runs it. 

If that ownership defaults entirely to IT or to a vendor’s black box, companies risk scaling systems they cannot fully adapt or audit. Giving business teams the tools and mandate to own their logic is what makes the AI system trustworthy and responsive to how your business runs. 

That is why I see analysts as the architects of this next phase. 

What the logic layer looks like in practice 

Let me return to the commissions example, because it illustrates the concept precisely. Right now, when a salesperson needs to know their commission on a deal, they send a message to the commissions analyst. That analyst has their own spreadsheet — because comp plans change every quarter, with spiffs and special programs layered on top. They run the math manually and send back an answer. 

What if that same analyst built a simple, well-defined calculator that encoded their commission logic — the actual rules for your company, your plans, your programs — and connected it to the AI systems your salespeople are already using? Now when a rep asks what their commission will be on a specific deal, they get the right answer. Not a generic explanation of how commissions work. 

And here’s the compounding value: that same logic can then be used by the annual planning agent to model the operational cost implications of different comp plans. It can feed the scenario planning model that runs hundreds of simulations for financial planning. The analyst who built it enables an entire network of AI systems to act on accurate, business-specific logic. 

That’s the logic layer in practice: curated, purpose-built data assets and calculators that encapsulate how your business works, maintained by the people who understand it, deployable to every AI system that needs it. 

What the logic layer requires 

This is where I think most companies are still stuck. They’ve made the infrastructure investments. They have cloud data platforms and approved LLMs. But they’re asking those systems to do things they were never designed to do on their own. 

The logic layer requires three things: 

  • Purpose-built data assets. A narrow, clean, well-defined data set that reflects how you actually measure a specific business process. 
  • Encoded business logic. This is the part that lives in people’s heads right now — the policies, the edge cases, the context that makes data mean something. 
  • The ability to update it. Nobody runs a business to keep it the same. The logic layer has to be something that domain experts can update when the business changes. 

A pragmatic path forward 

The good news is that you don’t have to wait for a perfect architecture before you start building a logic layer. 

Start with your highest-value, most-repeated business processes — the ones where an analyst is currently fielding the same questions week after week. These are the processes where encoding logic into a curated, AI-ready data asset delivers immediate, measurable value. 

Then, empower your analysts to own that encoding — not IT. Give them low-code tools to do the work, and the mandate to treat that encoded logic as a strategic asset they own and evolve as the business changes. 

This is also where leadership posture matters. 

I have said for a while that this should not be framed as a choice between business and IT. It is both. IT should set standards, manage infrastructure, establish security boundaries, and make approved AI capabilities available across the organization. But IT should not become the bottleneck for every piece of business logic the company needs to operationalize. 

If this feels familiar, it should. We have seen this pattern before in enterprise technology. Infrastructure and platforms matter. But the last mile, the part that turns capability into business value, always depends on the people closest to the work. 

AI is no different. 

The companies that get the most from AI will be the ones that treat it like an operating model. They will automate core workflows, curate the right data, and empower analysts and domain experts to define the logic that makes AI useful and generate answers the business can use. 

I recently had a chance to go deeper on these ideas on the Talking AI podcast. If you want to hear more of my thinking on the analyst’s evolving role, how the logic layer connects to agentic workflows, and why I think the next 18 months will be pivotal for getting this right, it’s worth a listen. 

 To learn more, visit us here

Where enterprise intelligence really comes from

Every new frontier model release seems to spur a fresh round of doomsday articles. Just Google “the end of white-collar jobs,” and you’ll be bombarded with discourse on the end of modern work, the unraveling of the social contract between employees and organizations. 

What I don’t see anyone talking about, however, and what I believe is a far more productive conversation, is the opportunity for knowledge workers. 

Nobody understands critical business processes better than your line-of-business (LOB) employees. Not executives. Not IT. Not even the most advanced LLMs. These are your business analysts and RevOps professionals, your supply chain managers and finance leaders, and the employees whose expertise has been forged over decades. 

For an enterprise to become truly intelligent, these workers must be involved in how AI workflows are built and deployed. Their guiding hand is the only way AI can learn and truly understand your business. 

But what does this transition look like, and how can organizations start operationalizing AI in a meaningful way alongside knowledge workers? Let’s take a look. 

What enterprise intelligence requires 

Imagine walking your board through a set of financials and recommending specific actions. Then, in your next meeting, you walk everything back because your AI layer got the numbers wrong. 

There is no faster way to kill an AI initiative than by delivering wrong outputs. Without trust, the whole system falls apart. 

In our recent survey of 1,400 business and IT leaders, we found that while over 90% of organizations are using AI, only 28% trust it to support decision-making. As for how many organizations scaled their AI pilots into production, the number was just under 25%, suggesting a very strong correlation between trust and operationalization. 

An intelligent enterprise, then, is an organization that has trustworthy AI embedded across the business. 

At Alteryx, we say the results of any AI system must follow our VURA framework: an AI system and its outputs must be visible, understandable, repeatable, and auditable. In other words, two people need to be able to go to AI with a question and arrive at the same answer; anyone who uses AI in their workflows must be able to explain how their AI system arrived at that answer. 

Who’s responsible for operationalizing AI? 

Enterprise intelligence is about trustworthy AI deployed throughout key business processes, but who’s ultimately responsible for these AI systems and processes: IT teams or knowledge workers? 

Let’s say you want to use AI in your Sarbanes-Oxley process, e.g., your journal entries, revenue recognition, access controls, etc. Before IT can help you build a new AI workflow, IT must first understand your Sarbanes-Oxley process in great detail. Then, they have to code a tool your finance team can trust. 

It’s possible, sure. But creating this solution would take an inordinate amount of time. Then, when a new regulation comes along or you have an acquisition, the whole thing falls apart. You have to get back in line with IT to retune everything. 

Moreover, if your books don’t balance out or if you fall out of compliance, IT does not want to have that responsibility fall on them. You can see why ownership of AI systems and workflows must sit with LOB workers. They are the only ones with the expertise to ensure the veracity of AI’s outputs. They are the only ones who can successfully shape and define its logic and oversee its ongoing execution. 

Data is the fuel. Business logic is what keeps AI on course. 

Finally, there’s the question of data. We’ve all heard “bad inputs, bad outputs.” Seeing as I’m the CEO of a data analytics company, you might expect me to say that reliable data is the end-all, be-all when it comes to trustworthy AI outputs. 

And while it’s absolutely essential, it’s only the first step. 

Aggregating your enterprise data into a cloud data platform is immensely useful. All of that data becomes readily accessible. You gain a single source of truth across teams and workflows. But you can’t point your LLM at a cloud data platform and ask it to make sense of your data for a complex business process. 

Again, you need the people who understand these critical processes to guide your LLMs to interpret the right data in the right way. This is what will make your AI systems visible, understandable, repeatable, and auditable. Yes, you need clean, reliable data. But more than that, you need business logic around that data, and that can only come from your knowledge workers. 

The five pillars of enterprise intelligence 

At the highest level, enterprise intelligence rests on five core pillars: 

  1. Trustworthy, transparent data 
  1. Empowered business analysts 
  1. Shared responsibility across the C-suite 
  1. Cross-functional collaboration 
  1. Leadership that evolves alongside AI 
     

Each pillar reinforces the same core idea: AI only becomes valuable when it’s grounded in reliable data, shaped by real business expertise, supported by executive ownership, and scaled across teams that can put it to work to improve their daily processes. 

Tap into the intelligence all around you 

As a business leader looking to build an intelligent enterprise, the most important questions you can ask are the ones around operationalizing AI in key business processes. What would it take for you to trust AI’s outputs? What would make AI-powered processes superior to your current ones? 

Once you have those answers, engage your LOB workers immediately. Give them ownership and autonomy. Rather than asking AI to replace them, lean into their intelligence. Let your knowledge workers use their expertise to amplify, shape, and govern AI. Their business mastery is what makes enterprise intelligence possible. 

 To learn more, visit us here

VURA: A framework for trustworthy AI at scale

“You’re right,” the LLM says. “I was mistaken.” 

Have you ever read these words during an AI workflow? Nothing kills trust faster than incorrect outputs. It’s no wonder, then, that only a quarter of businesses today fully trust AI to support decision-making and forecasting. 

And yet, we know AI is business critical. Nine out of 10 businesses are using it; 64% say it’s powering innovation. 

So, how do you bridge the gap from experimentation to trustworthy deployment? How do you get verifiable, reproducible results from AI at scale? In this article, I’ll show you the framework that’s powering AI success for leading organizations. 

Why organizations still don’t trust AI 

We asked 1,400 IT and business leaders what their biggest barriers to success with AI workflows were. One in two (49%) said inaccurate or biased outputs; 38% said it was a reluctance to allow AI to make decisions without human oversight. 

Then, there was the data issue. Data readiness is an integral part of successful AI workflows. However, half of all organizations said they still faced poor quality or fragmented data. While you don’t need perfect data to start using LLMs, you absolutely need trustworthy data. 

VURA: The framework for trustworthy AI 

Closing this trust gap requires two things. First, organizations need a logic layer that connects AI systems to the people who understand the data and business best. Line-of-business teams and analysts cannot sit on the sidelines. They need to help build and validate AI workflows so the logic behind AI’s outputs reflects how the business actually operates. 

Second, AI workflows and processes should be visible, understandable, repeatable, and auditable. Together, these principles form VURA, a framework we developed to help organizations build and scale trustworthy AI systems. These guidelines will help build trust in your data and your AI’s outputs. You’ll need both if you want your business to build enterprise intelligence. What follows are the four pillars of VURA.  

  • Visible 

Visibility is transparency. Your AI workflows shouldn’t be a black box regarding the data used and the logic applied. Every employee using AI tools should be able to answer two questions: “Where did this answer come from?” and “How did we draw that conclusion?” Otherwise, employees may be working from incorrect information. They could give your customers faulty intel or make important decisions with serious downstream effects. 

If those answers are still unclear, you may need to tighten your governance or reconsider whether your current AI and data solutions are working. Visibility becomes especially important when AI is used across teams. 

  • Understandable 

It can almost feel like science fiction when tools like ChatGPT or Gemini take the most complicated or vague of prompts, parse through them, and give you an intelligent, thoughtful answer. 

However, this low threshold for asking and answering virtually any question in natural language isn’t an excuse for glossing over business fundamentals. Your AI systems must be able to explain the logic behind their outputs to even non-technical business users, and your business experts must be able to validate those outputs. 

  • Repeatable 

Repeatable means that with the same AI tools, data, prompts, and business logic, AI will give you the same answer every time. Two people should be able to go to AI with the same question and arrive at the same answer. If an AI system or workflow gives you an excellent answer followed by one that’s clearly wrong, it’s not ready for operationalization. You can’t trust it. 

Repeatability also requires documentation. When teams identify prompts or processes that help produce reliable outcomes, those should be recorded and shared. 

  • Auditable 

An auditable AI process means you can see what happened. There’s a trail. If there’s an answer or report that seems off, you should be able to identify who owns the workflow, what data and prompts were used, what logic the system followed, and where human judgment and oversight were involved. Auditability is a check and balance for both your AI systems and the human engineers working behind the scenes. 

Start building trustworthy AI systems today 

AI can only deliver scalable business value when it’s grounded in trustworthy data and business logic. To operationalize these systems, you’ll have to ensure your AI workflows are visible, understandable, repeatable, and auditable. 

Alteryx is the transformation and business logic layer that helps you move AI from experimental pilots to trustworthy production. It connects to data wherever it lives, helps business users apply their expertise to AI-powered workflows, and instills the guardrails needed for both your data and your AI systems. 

With Alteryx, the people closest to the business can shape how data is prepared and applied, while IT gains the governance and auditability required for enterprise use. That’s how AI outcomes become trustworthy. That’s how enterprise intelligence is built. 

 To learn more, visit us here

Scaling beyond spreadsheets: platforms built for large-scale data analysis

If you’re a senior analyst, you’ve probably faced a dataset that used to open in seconds but now takes minutes. Or maybe you’re up against a formula that worked fine last quarter, but now shows an error because someone renamed a tab in a file three layers upstream. Ever spent an afternoon figuring out whose numbers are right when colleagues send back a few different “final” versions of the same report? None of that is a personal failure, but it is a sign that the volume and complexity of your work has outgrown what a spreadsheet was built to handle. 

The cost is more than just your time and inconvenience. When reporting slows down, multiple versions of a number circulate before anyone catches it, or one person’s spreadsheet logic is the only process your team has for something that matters, that’s a risk to the business. 

Plenty of solid analysis still belongs in a spreadsheet, but as your data and stakeholders grow, the balance between preparing data and analyzing it changes. If prep now takes more of your week than analysis does, the tool has become the bottleneck — not you. 

Watch for concrete signs you’ve hit that ceiling, what a platform “built for scale” needs to do differently, and how to decide whether it’s time to move. 

When spreadsheets stop being enough for your data analysis 

Every spreadsheet has a hard ceiling, and it’s lower than people might expect. Microsoft’s own published specifications cap every worksheet at 1,048,576 rows by 16,384 columns, regardless of your computer’s memory or Excel version. Once a dataset crosses that line, rows don’t get flagged — they simply don’t load, and it’s easy to miss. 

The bigger risk is accuracy. A 2024 literature review published in Frontiers of Computer Science, covering more than 30 years of spreadsheet research, found that 94% of spreadsheets used in business decision-making contain errors that create real risk of financial losses and operational mistakes. Most analysts already understand that the more a spreadsheet grows past its original design, the harder it gets to trust every formula in it. 

Alteryx’s own research backs this up from the analyst’s side of the desk. The 2025 State of Data Analysts in the Age of AI report, a global survey of 1,400 data analysts, found that 76% still rely on spreadsheets for data preparation, even as AI tools reshape the rest of their workflow. Manual prep work isn’t a habit analysts choose, but it’s still the default because nothing else is in place yet. 

Spreadsheets are ultimately designed for individual calculation, not for shared, repeatable, large-scale analysis. Asking them to do that job is where the cracks start to form, and when you need to start thinking of an alternative. 

What “built for scale” means 

“Scale” gets used loosely in analytics marketing, so it’s worth being specific about what a platform needs to do differently than a spreadsheet. It comes down to four tasks: 

  • Connect to data where it already lives 

A spreadsheet only knows what you paste into it, which means every report starts with an export, a download, or a copy-paste job that’s already slightly out of date by the time it’s finished. Platforms that support scale should be designed to connect, transform, and prepare AI-ready data by connecting natively to a broad range of enterprise applications, databases, and cloud platforms. This way data can be pulled in and refreshed rather than manually re-exported every reporting cycle. For an analyst, that means less time reconciling which export is current and more time on the analysis itself. 

  • Prepare and blend without rebuilding from scratch 

Scalable platforms should also give analysts a drag-and-drop canvas for cleansing, blending, and reshaping data from multiple sources, with code-friendly options like Python and SQL available for analysts who want them. It should allow for logic to be built once and held to a standard we call VURA: visible, understandable, repeatable, and auditable. Analysts shouldn’t have to deal with a chain of formulas that only one person fully understands, and that visibility matters as much as the automation. When you build a workflow as a series of documented steps, colleagues can review, troubleshoot, or take over in a way a dense formula chain rarely allows. 

  • Automate the workflow, not just the calculation 

The real definition of scale for an analyst is a process that runs without being rebuilt by hand. An essential component of that is workflow automation and orchestration so analysts can schedule and reuse any workflow they build. 

  • Report without abandoning familiar formats 

Leaving spreadsheets behind for analysis doesn’t mean stakeholders lose the outputs they’re used to. Reporting tools can generate tables, charts, and formatted outputs in PDF, HTML, or Excel, closing the loop between analysis and the people who need to read the result. 

Spreadsheets vs. a scale-ready analytics platform 

The differences between spreadsheets and analytics platforms are less about features and more about the needs that develop as your data and your team grow. 

Consideration Spreadsheet Scale-ready analytics platform 
Data connections Manual export/import from each source; data goes stale as soon as it’s pasted in Native connections to databases, cloud platforms, and enterprise apps that can refresh on demand 
Repeat work Rebuilt or copied by hand each cycle Built once as a workflow, then reused and scheduled 
Row and file limits Fixed worksheet ceiling regardless of vendor Designed to process large volumes without a hard row cap in the tool itself 
Auditability Hard-to-trace formulas and edits Workflow steps that are visible, understandable, repeatable, and auditable (VURA) 
Collaboration Version conflicts, emailed copies, confusion over which file is current Shared workspace with a single source of truth for a given workflow 

Matching the platform to where your team is right now 

“Scale” doesn’t mean the same thing for a 5-person team tracking budgets as it does for an enterprise running hundreds of scheduled workflows. It’s important to consider your team’s specific size and overall org structure before you shortlist any options. 

If your team is still primarily working out of Excel or CSV files and wants to reduce manual, repetitive spreadsheet work, there are several platforms built just for these types of applications. Alteryx One Starter Edition, as an example, is built specifically for that transition — code-free data prep accessible from a browser, aimed at teams getting started rather than running complex automation. 

You don’t need to know your exact tier or platform before you start a conversation with your team, but the conversation can go faster when you can describe your situation in terms of how many people touch the data, how often it needs to run, and who needs to see the output (rather than starting from a feature list). 

Governance doesn’t disappear with spreadsheets 

It’s tempting to think that moving off spreadsheets automatically solves governance, but that just changes what governance looks like. TechTarget’s coverage of data and analytics governance requirements notes that organizations should look for scalable, modular platforms that can adapt as needs change, rather than assuming governance is solved by the platform switch alone. 

It’s important to investigate whether or not a platform supports this with governance and administration capabilities such as role-based access controls, audit logs, and version history. They’re built to give IT the oversight it needs while giving analysts the flexibility to build and run their own workflows. 

A quick readiness check 

Before you bring a platform comparison to your team or your leadership, it helps to be specific about what’s driving the need. What follows are a few questions worth answering honestly: 

  • Are you regularly working with datasets that approach or exceed Excel’s row limit, or that make Excel noticeably slow to open and calculate? Slow-loading files and truncated imports are usually the first visible sign, not the first real cause. 
  • How much of your week goes to gathering and cleaning data by hand, instead of analyzing it? If prep consistently outweighs analysis, that ratio is the problem — not a single unwieldy file. 
  • If you left tomorrow, could someone else pick up your spreadsheet-based process without you walking them through it? A process that exists only in one person’s head is a significant business continuity risk. 
  • Do stakeholders currently receive conflicting versions of the same report because multiple people are editing copies independently? That’s a flag that the workflow needs a single source of truth. 
  • Does your organization need an audit trail for how a number was calculated, not just what the number is? Regulated or audited environments tend to outgrow spreadsheet-based tracking quickly. 

If you answered yes to two or more of these, this is a reasonable signal that it’s worth having a conversation about moving beyond spreadsheets. 

For a deeper dive on how to build the internal case, check out this guide from Alteryx on evaluating workflow automation tools for analytics teams and another breakdown of evaluating business intelligence tools that scale without increasing complexity

See it on your own data 

The fastest way to know whether a platform fits your workflow is to run your own data through it rather than a demo dataset. You can start a free trial of Alteryx One to test connectivity, data prep, and workflow automation against the kind of analysis you do every week. 

[CTA]  

 To learn more, visit us here

AI built the report, but can your business trust it?

The old way of creating reports is almost cliché, but only because it remains so pervasive. 

It’s a familiar scene: A teammate pings you at 4:57 pm asking for a last-minute report. The data and business logic you need live across 10 spreadsheets, in five Microsoft Teams threads, and in an email from two months ago that you can’t seem to find. 

But that was the old way. What happens when you use AI for the same situation? 

Let’s find out. 

What working with AI often looks like 

Your stakeholder pings you, asking for a report based on a massive tax reconciliation spreadsheet. 

This spreadsheet is a beast, chock-full of tabs, formulas, and data that’s been copied and pasted from several enterprise data sources. 

“Sorry for the last-minute ask,” they say, “but can you just throw AI at this?” 

You go to your LLM prompt library, select a robust prompt, and input it into Claude, along with the spreadsheet. 

Four seconds later, you get over 1,700 lines of Python code. Somewhere inside, there appear to be all the data transformations, calculations, and visualizations you need to build your report. 

But there’s a hiccup. 

Your stakeholder remembers that your tax jurisdictions change four times a year and wants to ensure that you can make any necessary changes. 

Sure, you think, that shouldn’t be a problem. I can probably find that line of code somewhere … 

Also, there are three subsidiaries. Someone else handles those taxes, so you’ll need to filter those out. 

Finally, your stakeholder remembers that your CFO will want to sign off on this and that your auditor is coming tomorrow. They’ll both want to see the logic behind your report. 

Suddenly, parsing through and validating hundreds of lines of AI-generated code seems far more difficult and time-consuming than you’d hoped. 

VURA: The missing piece 

While AI can bring incredible levels of automation and speed, those are only force multipliers when directed strategically. 

“I can get an infinite number of PowerPoints out of the AI systems if I want that,” Ethan Mollick recently told me during our executive exchange. “It may even be good content, but if it doesn’t serve the purpose you need it to, the productivity gains become a trap.” 

Ethan’s point is that more isn’t always better; bringing four hundred PowerPoints to a sales call won’t help you close a deal. Likewise, instantly generating hundreds of lines of Python is unlikely to help your CFO feel confident in your AI’s vibe-coded report. 

For an AI workflow to be trusted, it has to be Visible, Understandable, Repeatable, and Auditable, or VURA. You need to know what’s happening at every step of the process: where the inputs came from, how business logic was applied, and whether the outputs were correct. 

So, how can you accomplish this? 

The transformation and business logic layer 

Let’s try a different AI-powered workflow. Same situation and model. Only this time, we’re going to add a visual transformation and business logic layer. 

First, we go into Claude and type up a prompt, but instead of Python, we ask for an Alteryx workflow. 

We open our workflow in Alteryx, and instead of hundreds of lines of AI-generated code, we see a visual canvas showing the entire tax reconciliation process. 

It’s still an AI-generated workflow, but now, anyone in the organization can inspect it. They can see what data was used. Your analysts and domain experts can validate the logic. And you can add governance and repeat the process. 

Suddenly, AI-generated workflows become far more trustworthy and scalable, giving you a foundation for enterprise intelligence. 

The future of enterprise AI workflows 

AI tools that can’t adapt when the business changes have short shelf lives, and rebuilding from scratch constantly drains tokens, time, and energy. Endless iterations create endless chances for inconsistencies and errors. 

With a visual business logic layer, the people who know your business best — your business analysts, sales professionals, finance team, and more — can apply their expertise to your AI workflows and validate its outputs. They can see what’s happening at every step of your AI workflows so that every process is Visible, Understandable, Repeatable, and Auditable. 

Speed and reliability are no longer mutually exclusive. Now, you can bring AI’s power and your business experts together to create something fast and reliable, the intelligent solution you need to create scalable business value. 

Learn more: See how Alteryx One can help you build AI workflows your business can trust. Or, watch a live workflow demo to see Alteryx in Action. 

 To learn more, visit us here

Why finance teams need to modernize the logic behind spreadsheets

There’s a version of this story you’ve probably lived. The close is approaching, someone pulls a number from a file that hasn’t been updated, and an hour later you’re untangling a discrepancy that shouldn’t exist. The fix takes 20 minutes. Finding the source took two days. 

This is the part where most articles would tell you to ‘ditch the spreadsheet.’ But that’s not the real problem, and honestly, it’s a little insulting to the work you’ve actually done. 

Your spreadsheet isn’t the issue. The process built around it is. 

The logic is real. The medium is the limitation. 

Think about what lives in the workbooks your team maintains. How revenue maps to each entity. What counts as a valid reconciling item. The variance threshold that triggers a review. The intercompany elimination logic that took a year to get right. None of that is just data — it’s institutional knowledge. It’s business logic, and it belongs to finance. 

The problem is that spreadsheets were never designed to share that logic, version it, or let anything else use it reliably. When a process lives in a file on someone’s desktop, it’s invisible to every system downstream. You can’t hand it off cleanly. You can’t audit it without opening every tab. And when the person who built it leaves, a piece of your operations leaves with them. 

Why this matters more now than it did two years ago 

A lot of finance teams are under pressure to adopt AI — for close acceleration, anomaly detection, forecast assistance, narrative reporting. The pitch is compelling. The results, so far, have been uneven. 

Here’s why, and this part is specific to finance: AI can process data at scale and surface patterns quickly, but it cannot enforce your cost allocation methodology, validate your intercompany eliminations, or know what your organization has decided counts as an exception. 

For a tax team, that means it can’t apply your jurisdiction mappings reliably. For an audit team, it can’t reproduce your evidence logic. For FP&A, it can’t honor the constraint assumptions built into your planning model. That requires logic that’s documented, governed, and repeatable — and if that logic is locked in spreadsheets AI can’t see, AI can’t apply it. So it guesses. In finance, a confident guess on a tax provision or a consolidation rule isn’t a minor error. It’s a liability. 

The teams getting real value from AI are the ones who built the foundation first and then let AI work on top of it. 

The question most teams haven’t answered yet 

The shift that helps isn’t about which tool you use but where your process logic lives and who can access it. When your reconciliation rules, transformation logic, and validation criteria exist in governed workflows rather than locked files, the close gets more consistent, errors surface earlier, and handoffs get simpler. 

But getting from here to there raises a real question most teams are still working through: what does that transition look like for a tax team, an audit function, or an FP&A group that has years of logic built up in Excel? What moves first, what stays, and what does a week of progress realistically look like? 

That’s where the specifics matter — and that’s what we’ll get into next. 

 To learn more, visit us here

How finance leaders can close the AI trust gap

Most finance leaders at large organizations have made the right investments. A modern ERP, cloud data platforms, planning tools, and more. And now, increasingly, AI — for forecasting support, anomaly detection, close acceleration, and reporting at scale. 

The technology stack looks right. But when the board starts asking about results, the returns are harder to point to than the investments were. 

What your ERP was built to do — and what it wasn’t 

Your ERP is excellent at what it was designed for: capturing transactions, enforcing accounting standards, managing the chart of accounts. It is the system of record, and it performs that job well. 

But it doesn’t encode how your organization has decided to handle intercompany eliminations across a complex entity structure. It doesn’t carry your FP&A team’s cost allocation methodology, refined over three budget cycles. It doesn’t know what variance threshold triggers a controller review versus a VP escalation, or how your tax team has mapped jurisdictions for Pillar Two. That logic — specific, documented, organization-defined — isn’t in your ERP. It’s not in your data warehouse either. 

For most finance organizations, it lives in spreadsheets. Sometimes in the heads of the people who built them. 

Where AI runs into trouble in finance 

There’s a finding that gets cited a lot in finance AI conversations: research from MIT found that 95% of organizations are seeing no measurable return on their gen AI investments. Bain & Company looked at the same picture and reached a different conclusion for finance specifically. The fastest payback from AI in finance comes from embedding it in workflows — not from running pilots. The distinction matters because it explains why so many finance AI efforts stall after the proof of concept. 

AI can process data at speed and surface patterns across large datasets. What it cannot do is infer your business logic from raw inputs. Without that context, AI outputs in finance look confident but aren’t defensible — and in a function where auditability is a baseline requirement, that gap is not a minor limitation. It validates that trustworthy AI is critical for scaling workflows and AI pilots. 

Our own survey of 1,400 IT and business leaders asked what their biggest barriers to success with AI workflows were. One in two (49%) said inaccurate or biased outputs. Further, 38% said it was a reluctance to allow AI to make decisions without human oversight. While you don’t need perfect data to start using LLMs, you absolutely need trustworthy data. 

The layer that’s actually missing 

The gap between your ERP and your AI ambitions isn’t a data gap. It’s a business logic gap — the layer where your organization’s specific rules, methodologies, and decision criteria live, and where AI needs to operate to produce outputs you can stand behind. 

When that layer is built correctly — logic documented, workflows repeatable, outputs traceable — AI has validated, structured inputs rather than raw data it has to interpret. Outputs can be explained to auditors and to the board. And the sequencing question resolves itself: getting the process right is how you adopt AI. 

What it takes to build that layer 

Closing the gap takes more than a mandate to “use AI responsibly.” It takes three specific things, built and owned inside finance rather than handed off to IT. 

  • A purpose-built data asset for each process. Not another warehouse but a narrow, well-defined data set scoped to one process that reflects how your team measures it, not just what your ERP happens to store. 
  • Encoded logic, not tribal knowledge. The allocation methodology or the variance threshold that triggers escalation — built into a repeatable workflow instead of a senior analyst’s spreadsheet. The shift is building it once; in a form AI can use. 
  • A way to update it when the business changes. Comp plans get revised, tax jurisdictions shift, and the chart of accounts gets restructured after an acquisition. Logic that can only be changed by submitting a ticket to IT will be stale before it’s deployed — the people who own the process need to be the ones who can adjust the rule. 

None of this requires waiting for a perfect architecture. The highest-value starting point is whatever process has your analysts fielding the same question, the same way, every single cycle. Encode that one workflow first, connect it to the AI tools your team is already using, and the logic compounds from there: the same governed calculation that answers one controller’s question can feed the scenario model that runs your next planning cycle. 

 To learn more, visit us here

The CFO’s playbook for building AI-ready finance data  

Every CFO I talk to right now is under some version of the same pressure: the board wants AI, the business wants faster answers, and the finance team is often still reconciling spreadsheets. The promise of AI in finance is real. But so is the gap between that promise and what most organizations are able to deliver. 

I believe finance leaders need to be asking not simply, “How do we use AI?” but “What would make our data trustworthy enough for AI?” 

That distinction matters. AI-ready finance data is intentionally shaped for a specific business outcome, so we can trust what AI produces from it. In finance terms, it’s the difference between having transactions and being able to defend the numbers. 

Finance data is uniquely messy, and important 

Finance data is messy for rational reasons. We pull from multiple systems — ERP, CRM, payroll, procurement, planning tools, banks, data warehouses, and yes, still spreadsheets. 

We live through reorgs, acquisitions, new products, and chart of accounts changes. And when the business cannot wait, we create manual workarounds to keep moving. 

That complexity is the context in which we’re now being asked to use AI. It’s no wonder that so many initiatives stall. 

The non-negotiables of AI-ready finance data 

When Alteryx talks about AI-ready data, I translate it into a few non-negotiables. For finance leaders, this is where the concept becomes practical. 

  • Purpose-built, not “all the data” – AI-ready data should be scoped to the decision or workflow at hand. If I am building a cash forecast, I do not need every field from every ledger table. 
  • Clean and standardized – AI does not politely ignore bad inputs; it often amplifies them. That means your data needs to be deduplicated, standardized across dates, currencies, and units, and mapped to consistent hierarchies. 
  • Combined across sources, with business context – Finance work is inherently cross-source. AI-ready data is joined and enriched so the dataset reflects business reality, not just system silos. 
  • Traceable and transparent – This is where finance leaders should push harder than anyone else. AI-ready data has lineage. It is auditable and explainable, not just at the output layer, but in the data shaping behind it. 
  • Governed and controlled – AI readiness is about data risk management as much as data quality. AI-ready data should live inside a governed process, not a series of hero spreadsheets and copy-paste steps. 
  • Maintainable as the business changes – This is one of the hidden killers of AI initiatives. A one-time cleaned dataset is not AI-ready if it breaks the minute a new subsidiary is added, a cost center structure changes, or a revenue stream appears. AI-ready data has to be built through workflows that can be updated and re-run reliably, not through one-off cleanups. 

Where AI-ready data creates value in finance 

This is where the concept becomes real. AI-ready data is the difference between value and noise in some of finance’s most important workflows, including: 

  • Close acceleration: When trial balance data, mappings, intercompany logic, and exception rules are standardized, finance can generate more dependable variance flags and automate more of the financial close and reconciliation process. 
  • Cash forecasting: Better-connected bank data, AR/AP, billing schedules, and seasonality drivers make forecasts less likely to be derailed by missing or misclassified transactions. 
  • Anomaly and fraud detection: Clean, aligned vendor master data, payment runs, approval chains, and PO matching help teams reduce false positives and investigate issues faster. 
  • Revenue quality and leakage: When contracts, invoices, usage, CRM data, and credit logic are brought together in a way that reflects the actual economics of the business, AI can help surface patterns that matter. 
  • Narrative reporting: Grounding LLMs in curated, reconciled variance drivers and approved definitions allows teams to draft commentary responsibly within clear guardrails. 

Filling the AI data readiness gap 

I’ve found that in most organizations, there’s a constant friction point between data engineering and finance. Engineering understands the architecture, pipelines, and platforms. Finance understands the business context and logic — how revenue is recognized, how allocations work, where the exceptions hide. 

The handoff between those groups is often slow and messy. Analysts build fragile workarounds. Engineering teams inherit backlogs of finance requests that are actually business critical. 

What resonates with me about Alteryx is that it sits in that gap. It enables finance and business analysts to build repeatable data workflows for extracting, cleaning, joining, enriching, and shaping data for specific finance use cases. 

It emphasizes transparency and traceability, and it supports a model where IT can govern, and finance can execute. Just as importantly, it helps organizations turn their existing ERP, warehouse, and cloud investments into outputs that are actually usable for analytics, automation, and AI. 

How to get started 

If you want to make progress without boiling the ocean, my practical advice is simple: start small and start right. 

  • Pick one workflow that is high pain and highly repeatable (recs, allocations, forecasting inputs, reporting packs). 
  • Define what “trusted” means: the reconciliation rules, thresholds, approvals, and audit trail you need. 
  • Build the AI-ready dataset first cleaned, joined, governed, and repeatable. 
  • Then add AI where it makes sense (classification, summarization, exception explanation) inside the workflow, not as a free-floating tool. 

My bottom line is this: AI-ready data is an operating standard. It is how we scale AI without scaling risk. And for CFOs, that should be the real objective, not chasing the latest tool, but building the trusted data foundation that makes smarter automation, better decisions, and more resilient finance performance possible. 

To learn more, visit us here.  

The test every AI explanation in finance has to pass

Say your reconciliation tool flags a break between two ledgers, and now there’s a number that needs an explanation. The AI-generated summary says the mismatch is a timing difference, transaction posted late on one side. Reasonable. You move on. 

Then your controller asks which transaction, on which date, and why it posted late instead of on time. And now you’re not looking at an explanation anymore. You’re looking at a sentence that sounded like one. 

The four part test behind every AI answer 

That gap is the same thing the last piece here named: can you explain where the answer came from, and would the explanation survive someone pulling on it? Most practitioners have been running that check for years, on spreadsheets, on junior staff’s work, on their own numbers before a review meeting. AI just hands you answers that sound complete far more often now, and faster than the checking can keep pace with. 

The test itself breaks into a few plain questions, and it’s worth naming them because most people run all four without thinking about them separately: 

  • Visible: Can you see where the number came from? 
  • Understandable: Do you actually understand the logic that produced it, or just the sentence describing it? 
  • Repeatable: Would the same input produce the same answer next time, or is this a one-off? 
  • Auditable: Could someone other than you retrace it if they had to? 

Four different failure modes, and an AI-generated explanation can fail any one of them while still reading like a good answer. 

Why the gap is widening faster than the checking 

The reconciliation example holds up because it’s ordinary. Nobody’s arguing AI shouldn’t touch reconciliation work. Matching balances, drafting a first-pass explanation for a variance, flagging what needs a human look — that’s real time back. The problem isn’t the AI doing that work. It’s that the logic behind “this is a timing difference” has to already be defined somewhere the AI can point to. If it isn’t, the model is pattern-matching its way to something plausible, and plausible is not the same as traceable. 

Deloitte’s Finance Trends 2026 survey of over 1,300 finance leaders found 63% have fully deployed AI in their departments, with only 21% reporting clear, measurable ROI. That’s a broader adoption figure than an explanation-quality study, but the gap it points to lines up with the reconciliation example: plenty of AI running, not much of it yet standing up to scrutiny. 

Where the logic has to live 

Closing that gap starts with what the AI is drawing from in the first place, before it ever produces an answer. Every explanation an AI generates borrows its logic from somewhere: a threshold for what counts as material, a rule for what makes something a timing difference, an assumption about which system wins when two ledgers disagree. When that logic lives only as a pattern the model has inferred from past examples, the explanation is a guess dressed in confident language. When it’s defined, owned, and applied the same way every time, the AI has something real to summarize. 

Finance has kept this kind of logic for as long as the job has existed, often in a spreadsheet somebody built years ago that everybody trusts without fully remembering why it works. That logic hasn’t changed. Who can now touch it, and how fast, has, and that means the definitions underneath it need to hold up to more traffic than they ever have before. 

Get that part right, and the reconciliation example flips. The AI’s explanation becomes a summary of logic that was already defined, applied consistently, and traceable back to where it came from — the version that survives the follow-up question. 

See trusted AI workflows in action 

If you want to see what that looks like in a live workflow rather than in the abstract, Alteryx’s AI-Ready Starter Kits are pre-built Alteryx workflows and synthetic datasets designed to demonstrate how Alteryx can be applied to specific business use cases. They prepare and structure data to produce analysis-ready outputs, which can be extended using external AI tools. 

The Reconciliation Exception Resolution AI-Ready Starter Kit shows the pattern from this piece in practice: exceptions routed to an owner, prioritized by materiality, and documented consistently enough that the resolution holds up when someone asks how you got there. 

To learn more, visit us here.

Why finance leaders don’t fully trust AI and what they’re really checking for

There’s a specific moment every finance leader knows. A number is about to leave the building — headed for the board deck, the earnings call, or the audit committee — and right before it goes, you pause. You want to know where it came from, and you want to know it will still make sense if someone asks how you got it. That pause happens no matter what produced the number. 

That instinct shows up in the data too. In a Gartner survey of more than 200 CFOs, confidence across finance leaders’ top 2026 priorities averaged around 63%, while confidence in driving enterprise AI impact came in at just 36%. 

Leaders aren’t lacking confidence broadly. They’re confident about cost discipline and growth investment. The drop is specific to AI. It’s a broader measure than any single number leaving the building, but it points in the same direction: AI is the one place finance leaders can’t yet count on the confidence that usually comes easily. 

The four things every number has to pass 

That pause is a fast version of a test. Before you’d trust a number, you check four things: 

  • Where it came from 
  • Whether you could explain it simply 
  • Whether it would come out the same way twice 
  • Whether you could trace it back through the data if someone asked 

Most finance leaders have never written that test down. They’ve never had much reason to, because until now, the systems producing their numbers usually held up well enough that the check rarely turned into a real problem. 

AI doesn’t automatically pass that test. It can produce a plausible answer to almost anything, including things it has no real basis for knowing, and the answer looks the same whether the logic underneath is solid or made up. That’s the real source of the confidence gap. 

Leaders don’t doubt that AI can help. They doubt whether they could explain the answer if someone pushed back on it. The four things finance leaders already check for come down to four words: visible, understandable, repeatable, and auditable, or VURA. Those words succinctly describe what leaders were already checking for instinctually. 

Who owns the logic underneath 

Naming the test doesn’t resolve where it gets applied, though. That takes a harder answer about where the logic itself lives. Deterministic logic is defined by finance, not inferred by AI. A model can draft a variance commentary, summarize a forecast, or flag an anomaly worth a second look. 

It should never be the one deciding what counts as an exception, how revenue gets recognized, or which threshold triggers an escalation. Those are calls finance makes, and AI’s job is to work within them, explain them, and apply them consistently, not to invent them when it doesn’t have enough to go on. 

That distinction is where most AI disappointment in finance actually starts. The model usually isn’t failing at what it’s good at. The failure happens earlier: nobody defined the logic it needed, so it guessed, and it delivered that guess with exactly the same confidence it would use for a right answer. Looking at the output alone, you can’t tell the difference. 

That’s exactly what the four-question test catches. Ask where the number came from, whether you can explain it, whether it repeats, and whether you can trace it back to the data, and you’ll find out fast whether the AI applied logic finance defined or made something up that looks close enough. 

Building the standard into the workflow 

This is an architecture decision as much as a governance one. The four questions get easy answers when there’s a layer between raw enterprise data and the AI consuming it, one that prepares the data, holds the logic finance owns, and keeps every output traceable back to both. 

That’s the role Alteryx plays. It doesn’t compete with the model doing the reasoning, and it doesn’t replace the ERP or EPM system the data lives in. It’s the business logic layer that makes sure what reaches the model is something finance already stands behind, so the model’s output can be too. 

Build that in, and the pause before the number goes out changes what it’s doing. Instead of hoping the number will hold up, you can check that it does, every time, because the answers to those four questions are already built into how the workflow works, not something you have to reconstruct from memory. 

If you’re looking for a concrete way to see what that looks like on a real workflow, take a look at our AI-Ready Starter Kits: pre-built Alteryx workflows and synthetic datasets designed to demonstrate how Alteryx can be applied to specific business use cases. They prepare and structure data to produce analysis-ready outputs, which you can then extend using external AI tools such as large language models. 

Understanding why finance leaders hesitate to trust AI is only the first step. The next is building the governed foundation that gives AI reliable business logic to work from. 

Learn more in Building Finance AI You Can Trust, where you’ll explore the principles and practical steps behind AI-ready finance workflows. 

 To learn more, visit us here

The real reason AI isn’t paying off in finance

If you work in finance, you’ve probably been handed an AI tool in the last year or so. Maybe a copilot in your spreadsheet, maybe something bolted onto the close, maybe a chatbot that promised to answer any question about the numbers. And maybe, if you’re being honest, it hasn’t changed your Tuesday very much. 

You’re not doing it wrong. The tool isn’t broken. What’s missing is the part nobody put on the slide: AI is only as good as the work it’s standing on, and most of the time, the work underneath it is a mess. 

Confident AI answers you can’t trust 

Here’s a familiar scene. Someone asks the AI assistant a reasonable question — “why did margin move in the East region last month?” — and it produces an answer that sounds great. Confident. Well-organized. Possibly even formatted with little bullet points. The only problem is that you have no idea whether it’s right, because you don’t know which data it pulled, whether it used the current cost allocation method, or whether it quietly grabbed last fiscal year’s calendar. 

So you do what any sensible finance person does. You check it by hand. Which means the AI didn’t save you the work. It added a step. 

This is the quiet truth about why so much finance AI stalls. It’s not that the models can’t reason. It’s that they’re reasoning over data that was never cleaned, rules that were never written down, and logic that lives in one analyst’s head and three tabs of a workbook nobody else can open. AI didn’t create that gap. It just made it impossible to ignore, because now something is making decisions on top of it. 

McKinsey looked at how finance teams are actually using gen AI and found a useful counterexample. Across the handful of finance functions where they saw AI adopted in earnest, professionals were spending 20 to 30% less time crunching data — and putting that time back into the analysis their job is supposed to be about. In one case, a global consumer goods company pointed a gen AI assistant at budget-variance work and saw roughly 30% of that manual effort disappear. That’s a real result. But notice what made it real: it was pointed at a specific, repeatable task, working from data the team had already organized around a shared definition of what “variance” even means. The AI didn’t figure that out on its own. The team handed it a problem that was ready to be automated. 

What separates the workflows that pay off 

The finance work where AI delivers tends to share a few traits. It’s bounded — a clear start and end, not “answer anything about the business.” It’s repeatable, the same shape every month. And it’s tied to something that matters: cash, margin, risk, a number someone downstream is going to act on. 

That’s the easy part to say. The harder part is what has to be true underneath. For AI to work on one of those tasks, the data feeding it has to be prepared and validated before the model ever sees it. The rules — what counts, what gets excluded, how things roll up — have to be defined by your team and applied consistently, not guessed at by a model that’s never read your policy manual. And when the output lands, you have to be able to trace it back: which numbers, which logic, who signed off. In finance, that traceability isn’t a nice-to-have. It’s the difference between an answer you can put in front of an auditor and one you can only put in front of people who won’t ask hard questions. 

Think about the difference between two versions of the same workflow. In one, the AI reaches into raw data, applies whatever it infers the rules to be, and gives you a number. In the other, the data gets cleaned and structured first, your team’s actual business logic gets applied to it, and only then does AI work on top of a foundation it can stand on. The first one feels faster right up until something’s wrong and you can’t tell why. The second one is the one you can defend in a meeting. 

That’s really the test worth applying to any AI effort on your desk: can you explain where the answer came from, and would the explanation survive someone pulling on it? If yes, you’ve got something worth scaling. If no, more AI won’t fix it — it’ll just produce wrong answers more quickly. 

Where this leaves you on Monday 

None of this means starting over. The business logic your team has built — the spreadsheets, the rules, the institutional memory of how things actually work here — is the valuable part. The goal isn’t to throw it out for an AI that doesn’t know any of it. It’s to get that logic into a form that’s governed and repeatable, so AI can finally do something useful with it. 

The most practical move is also the least dramatic. Pick one workflow. Not the whole close, not “AI across finance.” One bounded, repeatable, annoying task you’d happily never do by hand again — invoice matching, a recurring variance pull, a report you rebuild every month. Get the data right for that one thing, write the rules down, and put AI to work on top of it. When it works, you’ll have something real: a workflow you can trust, and a clear sense of what the second one should be. 

Two traps worth naming, because they’re the ones McKinsey watched teams fall into. One is waiting for perfect data before you do anything — you’ll be waiting forever, and the team next door will have shipped three workflows by the time your data is pristine. The other is the opposite mistake: automating a process that’s still a tangle of exceptions and one-offs. Drop AI on top of a fragmented workflow and it doesn’t simplify it, it just adds a confident-sounding layer to the mess. The move is in between: standardize the one thing first, then automate it. 

If you want a low-stakes way to see what that looks like before you commit, our AI-Ready Starter Kits are built for exactly this. AI-Ready Starter Kits are pre-built Alteryx workflows and synthetic datasets designed to demonstrate how Alteryx can be applied to specific business use cases. They prepare and structure data to produce analysis-ready outputs, which can be extended using external AI tools such as large language models (LLMs). They won’t run your finance function — that’s not what they’re for. But they make the shape of a workflow that actually pays off tangible enough to copy. 

The AI on your desk isn’t the problem. The work underneath it is. Fix that for one thing, and you’ll stop wondering why AI hasn’t paid off — because it finally will. 

Search our full AI-Ready Starter Kit library to find finance use cases fit for you and your team. 

To learn more, visit us here.

Human-in-the-loop AI is becoming the default, not the exception

Over the past few years, much of the conversation has focused on autonomous AI and how quickly organizations can remove humans from decision-making. In financial services, we’re seeing the opposite trend. The organizations making the most sustainable progress aren’t eliminating human oversight—they’re redesigning it.

The model taking hold within the banking industry isn’t AI that operates independently and makes decisions; it’s AI that operates with intent and oversight. Human-in-the-loop is quickly becoming the standard, combining the speed and scale of machine-driven insight with the accountability, judgment and control that organizations can’t afford to lose. The shift is increasingly aligned with how regulators and industry frameworks are shaping responsible AI adoption, from the NIST AI Risk Management Framework to the revised U.S. banking agencies’ model risk management guidance, both of which reinforce governance, monitoring and accountability over blind automation.

From my perspective, this is not innovation slowing down; it’s AI adoption growing up. The first wave of enthusiasm focused heavily on what could be automated, but now the more important question is where can AI create meaningful value while keeping the right human judgment, oversight and accountability in place? In financial services, that distinction matters. In an industry built on trust, those capabilities are not optional— they are foundational to how we serve customers, manage risk and earn confidence every day.

Banking offers one of the clearest examples of why human-in-the-loop AI is becoming the default operating model for enterprise AI more broadly. Some of the most valuable AI use cases sit in environments where mistakes carry real consequences, customer impacts are significant and explainability is essential. In those moments, human oversight is what allows institutions to scale AI responsibly.

In banking, AI usually doesn’t operate in a vacuum. Whether it supports customer service, fraud detection, compliance, underwriting or internal productivity, it is touching workflows that affect customers, colleagues, regulators and the reputation of the institution. That is why responsible scale matters. Global bodies including the Financial Stability Board and the Bank for International Settlements have recognized the efficiency and analytical benefits AI can bring, while also warning that it can amplify model, cyber, concentration and governance risks if controls do not keep pace. For financial institutions, the mandate is clear – move with ambition, but scale with discipline.

Where human oversight matters most

The next phase of enterprise AI adoption will be defined by how well institutions understand where AI can move work faster, and where human judgment still needs to lead. For financial institutions, that starts with materiality. The greater the potential impact on customers, regulatory obligations or financial resilience, the stronger the case for meaningful human oversight.

Customer service is a good example. AI can help teams summarize inquiries, recommend next-best actions and reduce manual handling time. But when the issue involves a disputed transaction, a vulnerable customer, a complaint or product suitability, human judgment must remain central. AI can make service faster. It can make it more consistent. But it cannot replace empathy, context or accountability.

Fraud and financial crime are areas where AI can create real value, but human oversight remains essential. AI can detect patterns, anomalies and suspicious behavior across large data sets at a speed and scale people cannot match, but fraud is dynamic. Typologies evolve, bad actors adapt quickly and authorities have warned that AI can also increase the sophistication of scams, fraud and disinformation. In that environment, analysts and investigators play a critical role — validating signals, reducing false positives, escalating the right cases and applying judgment as the threat landscape changes.

Risk, compliance and credit are similar. AI can help synthesize internal data, identify control gaps and strengthen monitoring. But when outcomes affect lending decisions, regulatory obligations, capital or liquidity, institutions need governance that preserves challenge, review and accountability. The EU AI Act’s human oversight requirements for high-risk systems point to a broader direction of travel — the more consequential the use case, the more important it is that people can understand the system’s limitations, override outputs and intervene when needed. For U.S. institutions, the specific rule may differ, but the principle is already part of how banking operates. High-impact decisions require accountable oversight.

At the same time, human-in-the-loop cannot mean putting a manual checkpoint in front of every AI-assisted task, which would slow adoption and reduce the value AI can create. The goal is risk-based oversight. Lower-risk use cases may be managed through periodic review, testing and monitoring, while higher-risk applications may require real-time review before action is taken. What matters is that institutions define those thresholds clearly, rather than assuming one oversight model fits every use case.

Why collaborative AI is winning in financial services

The financial institutions that are embracing human-in-the-loop AI do so because they understand both the opportunity and the stakes. AI can process transactions, summarize complex information and identify patterns at a scale humans cannot match. At the same time, consumer expectations make clear that scale alone is not enough.  TD Bank’s 2026 AI Insights Report found that 78% of Americans now use AI-powered tools in their daily lives, yet only 18% are comfortable allowing AI to make important financial decisions independently. That gap says a lot about where the market is heading. Consumers are not rejecting AI, but they are drawing a clear line around accountability.

That is why speed cannot be the only measure of success. When customer outcomes, regulatory obligations or enterprise risk are involved, people still need to challenge the output, apply context and remain accountable for the decision.

That oversight matters because AI does not always fail in obvious or familiar ways. Generative AI can produce confident but inaccurate answers. Machine learning models can drift as data changes. Even highly accurate systems can deliver biased or poorly reasoned outputs when the data, assumptions or prompts behind them are flawed. NIST’s Generative AI Profile highlights risks including confabulation, privacy concerns, misalignment and automation bias.

For financial institutions, the lesson is that responsible AI requires people who understand how to use the technology, and also when to question it.

Building the organization for responsible AI at scale

In addition to being a technology challenge, responsible AI is also an operating model challenge. The institutions that scale AI well tend to do three things with discipline: establish clear governance, redesign workflows around the technology and build the skills employees need to use AI responsibly.

Governance starts with ownership, but it cannot sit with one executive or one team alone. It requires coordination across business lines, risk, compliance, legal, technology and model risk functions. That cross-functional model is becoming more common as organizations move beyond experimentation.  McKinsey’s State of AI report found that AI governance is often jointly owned and that CEO involvement in governance is correlated with stronger reported bottom-line impact, suggesting that firms derive more value when AI oversight is treated as an enterprise priority rather than a side initiative.

Workflow redesign is just as important because the real value of AI comes from reimagining processes end-to-end. That means identifying where AI can handle summarization, pattern recognition or drafting and where people should focus on exception handling, complex decisions and relationship-driven work. Human-in-the-loop is not about preserving the old operating model; it’s about building a better one.

That also requires new capabilities across the workforce — employees need to know how to use AI tools effectively and how to challenge them. They need to understand prompt quality, output limitations, data handling expectations and the warning signs that a system may be producing unreliable results. The  World Economic Forum’s 2025 report on AI in financial services underscores that while adoption is accelerating, responsible scaling depends on workforce adaptation, governance maturity and a clear understanding of risks alongside value creation. Responsible adoption depends as much on human capability as it does on model performance.

Looking ahead, enterprise AI in financial services will become more embedded, more specialized and more agentic in targeted domains. But that does not mean the human role becomes less important; if anything, it becomes more important. As AI takes on more analytical and operational work, people will increasingly serve as orchestrators, reviewers and decision-makers at the points that matter most. They will set objectives, define controls, interpret edge cases and know whether the AI results can be relied upon.

The organizations that lead in AI will be the ones that make human judgment a deliberate part of the design—clear about where AI can accelerate work, where people must remain accountable and how both can operate together with discipline. For financial institutions, the call to action is to treat human-in-the-loop AI as the operating model that makes innovation more trusted, more durable and more worthy of the customers and communities it serves.

Beyond chatbots: How embedded GenAI is transforming banking application development

Business application development is entering a new operating model. The traditional approach of gathering requirements, designing screens, writing services, integrating systems, testing, fixing defects and preparing release documentation still exists, but it is no longer sufficient for enterprises that need speed, traceability, resilience and regulatory confidence at the same time. Hyperautomation brings a broader discipline to this challenge. It combines workflow orchestration, intelligent document processing, robotic automation, API-led integration, process mining, test automation, observability and artificial intelligence into a connected delivery fabric. With embedded Generative AI, this fabric becomes more adaptive because applications can interpret natural language, summarize complex data, generate explanations, detect exceptions and support decision workflows rather than merely execute predefined rules.

In banking, this shift is especially meaningful. Banks operate across dense application landscapes: trade reporting platforms, wealth management portals, core banking systems, investment banking applications, digital compliance engines, reconciliation utilities, operational dashboards, audit repositories and daily, weekly and monthly reporting platforms. Each of these areas has its own data models, control points, integration patterns, validation rules, exception paths and regulatory obligations. Hyperautomation does not replace engineering discipline; it strengthens it by making business intent, technical execution, control evidence and continuous improvement part of the same lifecycle.

From automation to hyperautomation in banking applications

Automation usually addresses a specific task: moving data from one system to another, generating a report, running a batch job or validating a transaction against a rule. Hyperautomation goes further. It looks at the complete business outcome and asks how the entire chain can be streamlined, governed, observed and improved. For example, a trade reporting process may begin with transaction capture, enrich the trade with reference data, validate regulatory fields, identify breaks, generate a submission file, transmit it to a regulator or trade repository, monitor acknowledgements and preserve audit evidence. A narrow automation script may accelerate one step, but a hyperautomated design coordinates the complete flow, including exception handling and evidence generation.

Figure: Automation vs. hyperautomation.

Magesh Kasthuri

Figure: Automation vs. hyperautomation

Embedded Generative AI adds a new layer of intelligence. Instead of forcing every user interaction into rigid screens and codes, business applications can accept natural language prompts, interpret document content, summarize cases, generate draft responses, explain anomalies, produce test scenarios and create release notes. In a banking environment, this intelligence must be carefully bounded. Every AI-assisted action should be traceable, explainable, reviewable and aligned with data privacy, model risk, information security and regulatory expectations. The goal is not uncontrolled autonomy; the goal is governed acceleration.

Banking application components suitable for hyperautomation

A modern banking application is rarely a single monolithic system. It is a composition of business capabilities, integration services, workflow engines, data pipelines, user experience layers, analytics models, control dashboards and audit stores. Hyperautomation can accelerate the development and integration of these components by turning repetitive engineering work into reusable patterns and by embedding intelligence directly into business processes.

  • Trade reporting applications: Generative AI can help map trade attributes to regulatory fields, explain validation failures, summarize rejected submissions and generate test cases for reporting scenarios. Hyperautomation can orchestrate enrichment, validation, submission, acknowledgement tracking and evidence archival.
  • Wealth management platforms: Advisors can use embedded AI to summarize client portfolios, generate suitability narratives, identify missing documents and prepare personalized investment review notes. Automation can coordinate onboarding, risk profiling, document verification, portfolio rebalancing workflows and client communication approvals.
  • Core banking applications: Account opening, loan servicing, deposits, payments, interest calculations and customer maintenance can benefit from automated validations, intelligent forms, workflow routing and natural language assistance for operations teams. AI can explain account events or transaction exceptions in plain language.
  • Investment banking systems: Deal pipelines, research workflows, underwriting processes, trade lifecycle functions and risk calculations require strong coordination across front-office, middle-office and back-office platforms. Hyperautomation can standardize approvals, documentation, exception resolution and control evidence across these stages.
  • Digital compliance applications: Compliance teams can use AI to summarize policy obligations, compare regulatory changes with internal controls, classify alerts, draft investigation notes and produce evidence packs. Automation ensures routing, approvals, segregation of duties, audit trails and regulatory reporting timelines are consistently enforced.
  • Reconciliation platforms: AI can assist in matching narratives, explaining breaks, clustering exception patterns and suggesting resolution actions. Hyperautomation can pull data from ledgers, statements, payment processors, trading systems and data warehouses, then route unresolved breaks to the right teams.
  • Reporting and audit applications: Daily, weekly and monthly reports can be generated through controlled data pipelines, automated quality checks, narrative generation, variance explanations and approval workflows. Audit applications can preserve lineage, approvals, source extracts, model outputs and control attestations.

Embedded generative AI as an application capability

Embedding Generative AI into business applications should be treated as an architectural capability, not as a decorative chatbot. A banking application may use AI for search, summarization, reasoning support, content generation, code generation, policy interpretation or anomaly explanation. Each use case requires clear boundaries. The application must know which data the model can access, which actions require approval, what evidence must be captured and where deterministic controls must override probabilistic suggestions.

For example, in trade reporting, an embedded AI assistant can explain why a transaction failed validation and suggest likely fields to review. However, the final correction should pass through rule-based validations, maker-checker approval and audit logging. In wealth management, AI may draft a client review note based on portfolio movements and risk profile, but the advisor must verify suitability, disclosures and final communication. In reconciliation, AI can propose likely matches or categorize break reasons, while the system preserves the original data, confidence score, reviewer action and final resolution path.

Hyperautomating the product development lifecycle

The Product Development Lifecycle can itself become hyperautomated. Instead of treating ideation, analysis, design, development, testing, security review, release and operations as disconnected phases, enterprises can create an AI-assisted delivery loop where every stage produces structured artifacts that the next stage can consume. Platforms such as GitHub Copilot, Claude Code or Claude Cowork-style agentic development environments and OpenAI Codex can support this movement by helping teams reason over requirements, generate code, create tests, review changes, modernize legacy modules and produce documentation. Their value increases when they are connected to repositories, issue trackers, design documents, build pipelines, test suites, security scanners, observability data and enterprise knowledge bases.

PDLC StageHyperautomation OpportunityAI-Assisted Outcome
Business discoveryProcess mining, domain interviews, regulatory mapping, backlog creationStructured epics, user stories, acceptance criteria, process maps and control requirements
Architecture and designReference architectures, API contracts, data models, event flows, security patternsArchitecture options, integration blueprints, threat-model prompts and design decision records
DevelopmentCode generation, service scaffolding, UI component creation, data pipeline templatesReview-ready code increments, reusable components, migration utilities and integration adapters
TestingUnit, integration, regression, performance, compliance and synthetic data testingGenerated test cases, defect reproduction steps, test automation scripts and coverage summaries
Security and compliance reviewStatic analysis, dependency checks, policy validation, evidence captureRisk explanations, remediation suggestions, control traceability and approval evidence
Release and deploymentCI/CD orchestration, environment promotion, release notes, rollback preparationAutomated deployment packs, release summaries, operational checklists and change records
Operations and feedbackObservability, incident analysis, user feedback mining, backlog refinementIncident summaries, root-cause hypotheses, improvement stories and reliability recommendations

Role of GitHub Copilot, Claude Cowork and Codex

GitHub Copilot is useful where developers need assistance inside the engineering flow: explaining code, generating functions, proposing tests, reviewing pull requests and helping teams move from issue to implementation. In a banking PDLC, it can accelerate microservice creation, API integration, batch processing logic, reconciliation rules, regulatory validation routines and UI workflows. When used with repository context and proper review discipline, it can reduce the time developers spend on repetitive coding while preserving human accountability for design and correctness.

Claude Cowork or Claude Code-style agentic environments are valuable for multi-file reasoning, refactoring, debugging and documentation-heavy engineering work. Banking applications often contain deep domain logic scattered across services, configuration files, stored procedures, integration scripts and test suites. An agentic coding assistant that can understand a wider codebase context can help engineers analyze dependencies, prepare modernization plans, update multiple files coherently and draft explanations for reviewers. This is particularly useful in core banking modernization, trade reporting rule updates and compliance workflow refactoring.

OpenAI Codex can support issue-to-pull-request workflows, test generation, code review, bug reproduction, migration activities and broader software engineering tasks across the lifecycle. In a hyperautomated PDLC, Codex-like agents can be assigned well-scoped work items, asked to inspect failing tests, propose fixes, create regression coverage and summarize the change for human reviewers. The important design principle is to keep agents inside controlled boundaries: clear prompts, repository permissions, test gates, approval workflows and traceable outputs.

Integration architecture for hyperautomated banking applications

A practical architecture begins with business capability decomposition. Each banking domain should be expressed as a set of bounded capabilities such as customer onboarding, account maintenance, trade enrichment, exception management, portfolio review, control attestation, report generation and audit retrieval. These capabilities should be exposed through APIs, events, workflow tasks, data products and user interfaces. Hyperautomation then connects these capabilities using orchestration engines, event streams, rules engines, AI services, RPA connectors where legacy integration is unavoidable and observability layers that capture business and technical telemetry.

The embedded AI layer should sit behind a secure application service boundary. It should use retrieval-augmented generation where approved policies, product rules, application documentation and regulatory mappings are retrieved from trusted sources. It should avoid uncontrolled exposure of sensitive customer information. Prompt templates, response validation, redaction, grounding checks, model monitoring and human-in-the-loop approval should be part of the production design. In banking, the most successful AI pattern is often not full automation but assisted decisioning with strong controls.

Example: Hyperautomated reconciliation and reporting flow

Consider a reconciliation application that compares ledger balances, payment files, trade settlement records and external statements. In a conventional model, operations teams spend significant time downloading files, running macros, investigating mismatches, documenting break reasons and preparing status reports. In a hyperautomated model, data ingestion is scheduled and monitored, schema checks run automatically, matching engines classify obvious matches, AI assists with ambiguous narratives, exceptions are routed through workflow queues and dashboards update in near real time. At the end of the day, the system can generate a draft operations report explaining unresolved breaks, aging trends, risk exposure and pending approvals.

The same pattern can extend to daily, weekly and monthly reporting. Data quality rules validate inputs, report templates are populated automatically, AI generates narrative commentary on variances, reviewers approve or amend explanations and the final report is archived with lineage and approvals. Audit teams can later retrieve not only the report but also the source extracts, transformation logs, exception history, reviewer decisions and AI-generated drafts. This creates a richer control environment than manual reporting because evidence is captured by design rather than reconstructed later.

Governance, risk and control considerations

Hyperautomation in banking must be designed with governance from the beginning. The development team should define which activities can be automated, which can be AI-assisted and which must remain under human approval. Source code generated by AI must pass normal engineering controls, including peer review, static analysis, dependency scanning, secure coding checks, test execution and production readiness review. Business outputs generated by AI, such as compliance narratives or client-facing explanations, should be reviewed where regulatory or reputational risk is material.

Data governance is equally important. AI-enabled applications must respect data classification, residency, retention, masking and access policies. The model should not become an uncontrolled channel through which confidential customer, trading or employee information can leak. Every prompt, retrieved source, generated response, user action and final decision may need to be logged depending on the use case. For audit applications, this traceability is not optional; it is the foundation of trust.

Operating model for AI-native PDLC

A hyperautomated PDLC requires changes in team behavior. Product owners should write requirements in a structured manner so that AI tools can generate better stories, acceptance criteria and test scenarios. Architects should maintain living decision records, reference patterns and integration standards that AI agents can use as context. Developers should learn prompt discipline, context packaging and review techniques. Test engineers should focus on coverage strategy, synthetic data, compliance scenarios and defect prevention rather than only manual execution. Operations teams should feed incident learnings back into the backlog so the system improves continuously.

The role of human experts becomes more important, not less. AI can draft, generate, compare and suggest, but domain judgment remains essential. A trade reporting specialist understands regulatory nuance. A wealth advisor understands client suitability. A core banking architect understands transaction integrity. A compliance officer understands control interpretation. Hyperautomation works best when it amplifies these experts and removes repetitive friction around them.

Conclusion

Hyperautomation in business application development is not simply a faster way to write software. It is a new way to connect business intent, engineering execution, operational control and continuous learning. In banking, where applications must be reliable, explainable, secure and compliant, the combination of embedded Generative AI and disciplined automation can transform how applications are designed, built, integrated, tested, released and operated. Trade reporting, wealth management, core banking, investment banking, compliance, reconciliation, reporting and audit functions can all benefit when AI is embedded responsibly and automation is orchestrated across the complete lifecycle.

Platforms such as GitHub Copilot, Claude Cowork or Claude Code and OpenAI Codex can play an important role in this transformation by accelerating analysis, development, testing, review, modernization and documentation. Their greatest value appears when enterprises treat them not as isolated productivity tools but as part of a governed, AI-native PDLC. The future of banking application development will belong to teams that can combine human expertise, reusable engineering patterns, intelligent automation and strong governance into one coherent delivery model.

This article was made possible by our partnership with the IASA Chief Architect Forum. The CAF’s purpose is to test, challenge and support the art and science of Business Technology Architecture and its evolution over time as well as grow the influence and leadership of chief architects both inside and outside the profession. The CAF is a leadership community of the IASA, the leading non-profit professional association for business technology architects.

How AI helps the US Senate Federal Credit Union better manage risk

The United States Senate Federal Credit Union (USSFCU) is a nonprofit financial cooperative that provides traditional retail banking services to entities within the US government, such as the Senate and the Supreme Court.At present, the credit union’s headcount stands at nearly 150 people, managing around $1.6 billion in assets.

A few years back, when it started to expand its use of technology, cybersecurity was a key focus area, but the financial institution faced two major challenges in boosting security as it scaled. The USSFCU was carrying significant technical debt, and there were holes in the organization’s defenses.

“We found gaps where we needed more systems, tools, and people, and then there were instances where we had technologies in place that weren’t being used effectively,” says Mark Fournier, CIO at the credit union. “We weren’t buying a bunch of shiny new things without thinking about it. We were actually quite prescriptive every year, performing a number of different exercises to identify our shortcomings and then finding the right solution to fill the gaps. But over time this adds up. It was clear we couldn’t keep hiring more people and bringing in new solutions.”

The USSFCU needed a more efficient way to bring everything together and make its cyber estate easier to manage. For Fournier and his team, vulnerability management was the hardest hill to climb since they have to deal with about 100 new possible breach points every day.

“When we looked at the problem more closely, the impact of these vulnerabilities was far greater than we realized,” he says. “Not only because of the volume but because of a lack of clear understanding around the potential impact of each one across the broader business.”

Improved risk management

The USSFCU didn’t lack security tools, however. In fact, it had plenty, from scanners and endpoint tools to asset records, tickets, and internal documentation. But each tool saw only a slice of the environment, so there was little to no context. This made it difficult for the security team to separate real business risk from noise.

So for each new vulnerability, the security team had to run a manual investigation, which could take days. And while doing this, they still had to triage the next wave of findings. The organization, therefore, needed a way to know what mattered, why it mattered, who owned it, and whether taking the time to make a fix actually reduced risk. The USSFCU also required a solution to be deployed entirely in-house, leveraging its internal inferences.

Working with Tonic Security, the organization deployed an exposure management solution that pulls together data from different tools and data sources to create a clear picture of business risk. “One of the key functions of the platform is the ability to ingest anything,” says Fournier. “Breaking down silos between disparate systems is essential to unlock valuable contextual information.”

For the USSFCU, transparency and explainability are critical, he adds. This tool uses an AI data fabric to extract context from structured and unstructured data. This context drives prioritization, ensuring the right owner gets the right evidence, not a vague ticket. And once the work is done, the solution checks whether the exposure was reduced.

Because the AI is grounded in the customer’s own environment, it isn’t just guessing from a generic risk model. It reasons over USSFCU’s assets, owners, services, tickets, controls, and business context. But it isn’t using this data to train external models.

Describing one particular incident, Fournier explains that shortly after the initial deployment, various stakeholders met to assess progress. “We thought we were smart because we found an error with the platform,” he says. “The solution had labelled an asset as internet exposed, which we knew was incorrect.” But after a review and lengthy discussion, they were proven wrong. “Almost immediately, the value of bringing this information together became apparent.”

A template for bigger things

Before this solution, a high-severity finding could send an analyst on a lengthy scavenger hunt because of data located in so many different places. They’d check the scanner, asset inventory, tickets, and maybe even ask around to find the owner. But now they can find the asset, the owner, the business relevance, the exposure path, and the recommended action in one place. The solution has reduced the time taken to resolve a vulnerability by 75%. And with a clearer idea of what is and isn’t important, and what adds practical value, the number of incidents someone needs to respond to has reduced from about 100 a month to just 10.

Sharing his lessons from the project, Fournier says one needs to keep an open mind because the problem you think you have is often very different from the one you actually have. “This project has also been an eye-opener around how people can collaborate and operate across different areas of the business,” he says. “When I talk to my peers, they regularly highlight the disconnect between different departments and business functions. But with a project like this, when you’re crossing traditional boundaries, you need to have open lines of communication to succeed.”

How CaixaBank drives partner and customer relationships through AI

The transformation of the financial sector is no longer just about offering a mobile app or allowing customers to bank from anywhere. After years of digitizing services, institutions now face the more ambitious challenge of building a more personalized, agile, and intelligent relationship with millions of users who expect immediate answers, simple experiences, and service tailored to specific needs. 

The emergence of gen AI has accelerated this evolution. While banks have used AI models for years to automate processes, improve efficiency, and analyze large volumes of data, a new generation of conversational tools opens the door to a much more natural interaction between customers and financial institutions. 

Spain’s CaixaBank, for example, has positioned AI as one of the cornerstones of its technological transformation. The bank, which has more than 12 million digital users, believes this change isn’t solely due to tech’s evolution, but also to a shift in user expectations.

“Today’s customer is more digital, autonomous, and also more demanding in their relationship with the bank,” says Mariona Vicens, CaixaBank’s director of digital transformation and advanced analytics. “They not only interact more through digital channels, but also expect simplicity and personalized solutions at any time and from any device.”

A history of AI experience

Although gen AI has made a big impact, CaixaBank says its commitment to these technologies began much earlier. But it now represents a qualitative leap. “It’s more focused on developing new models based on conversational applications,” she says. “The most visible improvement is that gen AI allows for more natural, contextual, and useful interactions for both employees and customers.” 

This evolution is part of CaixaBank’s 2025-2027 Strategic Plan, in which it identifies agility, new services, efficiency, and technological resilience as main and interconnected objectives. For Vicens, agility is particularly key. “It’s what allows us to respond to a customer who increasingly expects immediacy, and it’s also what determines the bank’s ability to adapt in an increasingly dynamic environment.” 

The Cosmos Plan, the specific roadmap for processes and technology framed within CaixaBank’s strategic plan, reflects this integrated vision. “It combines investment in technology, automation, and AI to enable a more flexible and efficient organization capable of evolving at the pace set by customers,” she says. “Ultimately, agility is the visible engine of change, but it’s only possible when all elements of the model advance in a coordinated manner.” 

AI is certainly at the forefront of how the bank operates. More than 2,000 employees are already using agents to automate tasks, streamline processes, and improve customer service — a number the bank expects to increase before the year’s end. “With this implementation, combined with the application of other models like gen AI integrated into office tools, we expect to scale the gains in productivity and agility,” Vicens says.  

Innovation with human oversight 

While AI opens up new possibilities for transforming customer relationships, it also presents challenges related to regulation, transparency, and trust. For CaixaBank, innovation isn’t just about developing new use cases, but doing so under a governance model that ensures the technology is used responsibly. 

With that objective, the bank has defined a specific governance framework for these tools, with a corporate-level AI Office and a policy that anchors principles such as transparency and explainability, data fairness and privacy, robustness and security, and human oversight. 

This framework, CaixaBank explains, translates into concrete controls throughout the entire AI lifecycle: prior validation of use cases, structured risk assessment before implementation, corporate inventory of systems, subsequent monitoring, and incident management. However, it’s all based on the clear premise that relevant decisions can’t be entirely delegated to AI, so they must maintain human oversight. 

Regulation for confident innovation 

The entry of the EU AI Act has placed financial institutions under evolving regulatory requirements. Far from seeing it as an obstacle, CaixaBank believes this framework fits perfectly with how it’s approached the tech all along. “It fits naturally, because we’re precisely structuring our AI governance model with this framework and other regulatory frameworks as a reference, and we integrate it into the AI ​​lifecycle from the design stage and by default,” says Vicens.

Corporate policy explicitly incorporates the regulations into its global risk management system. In practice, any AI-based application must follow a clearly defined process before being implemented. “This means that any use of AI must be identified, evaluated, and monitored,” she says. “Before developing a use case, its type, value, and feasibility are validated, and then its risks are assessed. And once implemented, its performance is monitored.” Of course, in a financial environment, customer trust remains a most valuable asset.

Added AI agent muscle 

All this transformation strategy is finding a tangible application in one particular development: a contracting assistant that accompanies the client through digital channels. The system acts as a first point of contact when a user requests information about a product from the CaixaBank website or app. From there, it can answer questions, provide contextual information, guide the conversation, and, when necessary, transfer the interaction to a specialist without the customer having to restart the process. 

For the bank, this ability to understand context is a key differentiator. “Unlike a chatbot that answers a collection of FAQs, this agent is a contracting assistant that understands the context of the conversation with the customer, provides support, and can escalate to a human,” she says. For products like pre-approved loans, it can even lead the conversation to the final step before closing. 

The bank emphasizes that human intervention remains an essential part of the process. “We see AI as a tool to inform, streamline, and support the customer to enhance their user experience in a way that complements the ongoing support provided by our team of specialized remote banking managers,” Vicens adds. Plus, customers can choose to speak with a human from the outset or at any point during the conversation, and the final contract is always signed with the assistance of a CaixaBank specialist.  

Great responsibility

Beyond human oversight, the bank has established a framework to ensure the responsible use of AI. “It has defined responsible AI principles that cover the entire lifecycle of developments to ensure fair, transparent, responsible use, aligned with legislation and the group’s values,” she says. “Before deploying any AI solution aimed at customers, compliance with these principles is verified.”

In the specific case of the contracting assistant, data protection is one of the essential elements. The information travels encrypted, and the model isn’t trained with the data sent to the LLM. 

Currently, this technology is available in 40 products and manages an average of 6,000 conversations per month — figures that, according to CaixaBank, provide clear metrics of scale and productivity. 

The implementation of the onboarding assistant is one example of a much broader strategy in which AI, data, and automation are used to transform the relationship between the bank and its customers. “The key is no longer just being available, but providing real value in every interaction, and strengthening trust through useful experiences tailored to each user,” says Vicens.

❌