Meta won’t judge employees by how much they use AI when it comes to annual performance reviews, despite early efforts to drive AI adoption focusing on so-called token maxing.
The company has told employees that it “will not use AI adoption dashboards or token counts to evaluate impact,” according to a report by The Information.
The announcement came in an internal memo from executives Maher Saba and Santosh Janardhan, which said that instead of measuring AI usage, “
Meta won’t judge employees by how much they use AI when it comes to annual performance reviews, despite early efforts to drive AI adoption focusing on so-called token maxing.
The company has told employees that it “will not use AI adoption dashboards or token counts to evaluate impact,” according to a report by The Information.
The announcement came in an internal memo from executives Maher Saba and Santosh Janardhan, which said that instead of measuring AI usage, “managers should look at output quality, velocity, problem complexity and scope taken on.”
This marks a culture change for Meta, where engineers had previously competed to consume the most AI tokens, displaying their scores on a leaderboard. Meta then discovered that its employees were being diverted from regular work because they were using AI to carry out additional tasks to boost their scores.
Amazon had similar results when it implemented a leaderboard to track AI use; it also found that some employees were trying to game the system by using AI to complete unnecessary tasks, and it has now deleted it.
Meta had already started to look askance at the concept of using AI metrics as a tool to assess employees. Earlier this year, Chief Technology Officer Andrew Bosworth told employees in a memo that “nobody should be using AI tools just for the sake of using them,” adding that “token usage alone is not a measure of impact of any kind.”
As spending on cloud technologies grows, so does waste. The Flexera 2026 State of the Cloud Report found that 27% of organizations expect to spend more on cloud this year, with 17% already exceeding their budgets over the previous 12 months. The estimated share of wasted cloud spend has already crept up to 29%, undoing several years of progress, undoing several years of progress.
Companies are rapidly investing in cloud technology, but often understand less about how to
As spending on cloud technologies grows, so does waste. The Flexera 2026 State of the Cloud Report found that 27% of organizations expect to spend more on cloud this year, with 17% already exceeding their budgets over the previous 12 months. The estimated share of wasted cloud spend has already crept up to 29%, undoing several years of progress, undoing several years of progress.
Companies are rapidly investing in cloud technology, but often understand less about how to use it fully and efficiently. That is not a coincidence, and it is exactly the gap FinOps is meant to close. It is also why the practice is moving out of the finance department and into the strategy conversation.
What is FinOps?
FinOps is a blended operational framework that maximizes technology value by uniting engineering (DevOps), finance, and business teams. It involves close collaboration to break down silos between tech and finance, with shared ownership of cloud spend across engineering, finance, and business teams.
What distinguishes FinOps is real-time visibility into what is being spent and why. It also treats optimization as continuous work rather than an annual cleanup exercise.
FinOps is important because cloud spending isn’t like a typical budget line. It’s more variable and usage-based, so relying on an annual review doesn’t work. Engineers can quickly create infrastructure, scale it, and tear it back down in a day, making forecasting more challenging than in the past.
What FinOps does is change who sees what. Engineers have more insight into the actual cost of a build. Finance gets numbers it can trust. Business leaders can tie spending directly to business outcomes. It removes much of the guesswork and turns cost data into a shared language rather than a monthly surprise.
As Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise, puts it, “For a long time, FinOps meant only cutting the bill: find the unused stuff, resize a few instances, and report the savings. That still matters, but it’s not what separates companies today. The ones pulling ahead are using FinOps to make faster, smarter calls about where their tech spend actually pays off. That’s a different job, and it shows up directly in how fast a company can move.”
How FinOps spending has changed
Only a few years ago, FinOps was mostly focused on cloud infrastructure spending. Today, that focus increasingly includes AI-specific investment. The FinOps Foundation 2026 State of FinOps Report found that 98% of organizations now manage AI spend specifically. FinOps has also expanded well beyond cloud infrastructure. It’s more common now to see FinOps coverage extend to licensing (64%), private cloud (57%), and data centers (48%). Around 90% also manage SaaS spend or plan to do so within the next year.
It’s also worth noting that the same FinOps Foundation report found that 78% of teams report directly to the CTO or CIO rather than operating solely within finance departments. To us, that reporting line says a lot. It suggests that companies increasingly see technology spending as a strategic lever rather than simply a line item to reconcile.
How AI and cloud spending are moving in the same direction
The trend toward bringing AI and a broader range of technology spending into FinOps is backed up by a Gartner report, which estimates that global IT spending will hit $6.31 trillion by the end of 2026. That’s up 13.5% from the previous year. Data center systems spending is expected to grow 55.8%, with generative AI model spending more than doubling over the same timeframe. Gartner, in a separate forecast, expects public cloud services to grow by 21.3% in 2026, with the market reaching $1.48 trillion in value by the end of 2029.
We see these figures as two sides of the same shift. AI workloads are also usage-based, which makes them more unpredictable, partly because some teams haven’t had to consider unit economics before. A fine-tuning run or a forgotten inference endpoint can quickly become one of the biggest items on a cloud bill. Most teams don’t have the tagging, forecasting, or accountability needed to catch those costs before they get out of control.
“AI spend just behaves differently from a normal application workload. It spikes, it’s hard to pin on one team or feature, and you often don’t know the real cost per outcome until the invoice lands. Companies that already had solid FinOps habits before AI adoption took off are adjusting faster because visibility and ownership were already part of how they worked. Companies that treated FinOps as an annual cleanup are the ones getting caught out.”
Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise
Why visibility and shared spend ownership create an advantage
Flexera numbers on discount usage point to the same issue: fewer than 50% of organizations are using the most basic cost optimization tools. The adoption of tools like reserved instances or savings plans is slow, with only 48% of companies using Google Committed Use Discounts and 45% using AWS Reserved Instances. Too many others are leaving low-risk savings on the table.
In many cases, the real problem is a lack of ownership and visibility. If no team owns the cost of a workload, no one has enough reason or enough information to choose the right pricing model. That is where the competitive gap starts to open: some companies can explain and act on their spend quickly, while others cannot.
“A mistake we still see a lot is trying to optimize the bill instead of the system behind it. Deleting unused resources saves money once. Redesigning how workloads scale, how environments get spun up, and who’s on the hook for what keeps costs under control for good. That’s the difference that turns into a real competitive edge later.”
Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise
What visibility and shared ownership look like
FinOps operates well when at least three structures are in place.
Every workload or inference endpoint has a clear owner tied to its cost.
Cost and usage data are shared and available to everyone who needs them before they have to ask.
There is an ongoing review cadence designed around continuous optimization.
Teams that jump into dashboards before assigning ownership and establishing the data flow often end up with visibility but no accountability. Teams that start with ownership, even with basic tooling, get a different result. They tend to see savings stick instead of resetting every few months. That proper order is the single biggest predictor we have seen across cloud and AI cost engagements.
What business leaders need to know
There is a simple test a CEO or CFO should apply to FinOps. It’s whether the company can clearly state what a workload costs and whether it is worth that cost right now. Does it take finance three weeks to answer? Can the company provide an answer in real time? Are there live numbers and clear ownership behind every workload? That is what will enable leaders to make confident calls on where to invest and where to pull back.
FinOps is becoming a proxy for how well a company manages technology. With shared ownership comes high visibility. Add in continuous optimization and companies gain an advantage. These are not just cost-saving tactics. They represent operational discipline, separating those who can move quickly on AI from those who spend heavily only to find themselves still behind the rest of the pack.
Organizations often celebrate an AI launch at the moment the real work begins. The platform is available, the policy is published and employees have completed training. But none of those milestones tells a CIO whether work has improved, decisions are stronger or employees know when human judgment must override an AI recommendation.
This gap is visible in Kyndryl’s 2026 People Readiness Report. In a survey of 1,100 senior business and technology leaders across eight coun
Organizations often celebrate an AI launch at the moment the real work begins. The platform is available, the policy is published and employees have completed training. But none of those milestones tells a CIO whether work has improved, decisions are stronger or employees know when human judgment must override an AI recommendation.
This gap is visible in Kyndryl’s 2026 People Readiness Report. In a survey of 1,100 senior business and technology leaders across eight countries, 57% said AI was embedded in core processes or deployed broadly, while only 23% described their workforce as fully ready to use it successfully. Just 32% said their organizations had achieved at least one of their top two AI objectives. Technology deployment is advancing faster than the organizational capacity needed to turn it into value.
In transformation work, I have learned to be cautious when activity is presented as evidence of adoption. License activation, training attendance and prompt volume are easy to count. They do not show whether people can apply AI responsibly in a workflow or whether that workflow produces a better outcome.
Many CIOs now recognize that usage does not equal value. The next challenge is more difficult: creating an evidence chain that explains not only whether results changed, but why. That chain connects four layers – readiness, demonstrated capability, workflow behavior and business results.
Why deployment measures are insufficient
Many programs still treat workforce readiness as a downstream activity. Leaders select a platform, configure technical controls and announce availability. Training is then expected to solve every remaining problem: unclear use cases, employee anxiety, weak manager support, policy uncertainty and processes that were never redesigned.
When employees hesitate, leaders may interpret that hesitation as resistance. In my experience, it is often a rational response to ambiguity. People may not know which data they can use, whether an output must be verified, who remains accountable for a decision or how AI will affect the value of their role. A generic demonstration cannot answer questions that are specific to a job and workflow.
One practical readiness test I use is to ask people in different roles to describe the same AI-enabled workflow. Can they agree on its purpose, the information the system may use, the person who owns the outcome and the point at which a human must intervene? If not, the organization is not ready to scale. That disagreement is valuable evidence: It gives leaders a specific agenda for process design, communication, governance or learning.
Human involvement also should not be defined uniformly. A Stanford Digital Economy Lab study collected preferences from 1,500 domain workers and assessments from AI experts covering more than 844 tasks across 104 occupations. It found varied expectations for the level of human agency different tasks should retain. The practical implication is that leaders should not frame every use case as a choice between full automation and no automation. They should define the degree of human judgment each task requires.
Build an evidence chain for changed work
A useful AI adoption scorecard should answer four executive questions.
Readiness: Do people understand the purpose and boundaries? Readiness is more than awareness that a tool exists. Employees should be able to explain what the use case is intended to improve, which data is permitted, what outputs require validation, who owns the final decision and how to escalate a concern. Measure this with short scenario-based checks rather than confidence surveys alone. Present a realistic situation involving restricted data, an uncertain output or an exception to the normal process. Ask employees what they would do and why. A high self-reported comfort score is not a substitute for a correct decision.
Capability: Can people demonstrate the required judgment? Enterprise AI literacy provides a common foundation, but adoption requires role-based practice. A finance analyst, field supervisor and HR partner may share responsible-use principles, but they should not receive identical exercises or be assessed against identical criteria. Capability evidence should come from a demonstration in a realistic environment. Can the employee identify a plausible error, validate an important claim, document the basis for a decision and recognize when the case exceeds the system’s approved scope? This moves measurement from course completion to observable proficiency.
Behavior: Is the approved workflow being followed? Behavior measures whether the new practice has become part of the work. Platform analytics can contribute evidence, but they are not enough. CIOs also need to know whether people are completing required reviews, documenting decisions, escalating exceptions and avoiding unapproved workarounds. The target should not automatically be maximum usage. Some cases should remain human-only, and a high override or escalation rate may signal good judgment rather than poor adoption. Metrics must be interpreted in the context of the workflow and its risk.
Results: Did performance improve without unacceptable tradeoffs? Results should be defined before a pilot begins and compared with a credible pre-AI baseline or control group. Depending on the workflow, the relevant measures might include cycle time, first-pass quality, rework, error rates, cost, safety, risk events or stakeholder experience. Efficiency should always be paired with a quality or risk guardrail. Faster output is not progress if it creates more corrections, weakens decisions or transfers hidden work to another team.
In practice, consider an AI-assisted security-alert triage workflow. The desired outcome might be a reduction in the time required to classify high-priority alerts. The human accountability point is explicit: An analyst approves the severity classification and response action.
Readiness means analysts understand which information may enter the system and when escalation is mandatory. Capability means they can detect a plausible but incorrect severity recommendation. Behavior means eligible alerts move through the approved review path, with overrides and escalations recorded. Results mean triage time improves without increasing false negatives or delaying containment.
I recommend assigning an owner, evidence source, review cadence and decision threshold to each layer. The pilot should scale only when the desired behavior appears and the business outcome improves without breaching its quality, safety or risk guardrail. If usage rises but capability or results do not, the response should not automatically be more training. The use case, workflow, controls or management support may need to change.
This approach also makes cross-functional accountability clearer. IT enables the platform, data and controls. Business leaders define the work and desired result. Human resource and learning leaders build capability. Legal, compliance and security clarify boundaries. Managers reinforce behavior, while employees contribute the operating knowledge needed to make the workflow effective. The CIO’s orchestration role is to keep those contributions connected to the same outcome.
A 30-day test CIOs can start now
The World Economic Forum’s Future of Jobs Report 2025 found that 63% of surveyed employers viewed skills gaps as a leading barrier to business transformation. In response to expected AI disruption, 77% planned to reskill or upskill existing employees by 2030. More learning activity alone will not close the gap. Leaders must determine whether learning changes decisions, practices and results.
Over the next 30 days, ask each participating business unit to select one workflow and do six things:
Establish its current performance baseline.
Define one outcome AI is expected to improve.
Name the person accountable for the workflow result.
Identify one behavior that must change and one human decision that must remain.
Set a quality, safety or risk guardrail that cannot be traded for speed.
Review evidence from all four layers weekly and decide whether to scale, redesign or stop.
This creates a much stronger management conversation than reporting licenses, course completions or prompt counts. It shows where the evidence chain is breaking. A team may understand the rules but cannot challenge outputs. Employees may be capable but unable to use the approved tool within the actual process. The behavior may change while the business result remains flat. Each pattern calls for a different intervention.
Durable AI value will not come from the highest volume of activity. It will come from making expectations clear, giving employees realistic opportunities to practice, instrumenting how work changes and holding each use case to an explicit outcome and guardrail. A deployment turns the system on. Adoption changes how work is done. The evidence chain tells a CIO whether that change deserves to scale.
A single-page visual breakdown of the 108 cyber attacks recorded between August 1–15, 2026 — from the motivations behind them and the sectors hit hardest, to the tactics attackers used to get in.
A single-page visual breakdown of the 108 cyber attacks recorded between August 1–15, 2026 — from the motivations behind them and the sectors hit hardest, to the tactics attackers used to get in.
Cortney Pagel has learned to expect a particular question whenever she proposes changing how work gets done. Pagel, a senior business analyst and change manager at ENGIE Impact, often hears it from the digital side of the organization before she has finished explaining the change: How much money are we looking to save here? “I don’t always have an answer,” she says. “And so that is very frustrating for me.”
The frustration comes from a familiar structural problem. Most
Cortney Pagel has learned to expect a particular question whenever she proposes changing how work gets done. Pagel, a senior business analyst and change manager at ENGIE Impact, often hears it from the digital side of the organization before she has finished explaining the change: How much money are we looking to save here? “I don’t always have an answer,” she says. “And so that is very frustrating for me.”
The frustration comes from a familiar structural problem. Most financial systems connect money to departments, accounts, cost centers and, in more mature implementations, activities; however, they rarely connect economics to the anatomy of the work itself. Pagel describes processes that are still too manual to be tracked cleanly, while in other cases the financial detail exists but was never tied to the process model. For a period, she resorted to putting cost estimates in comment bubbles on process diagrams because there was no systematic place for them. “There’s been no great system or method to do it,” she says. “It’s definitely not just me.”
The question she keeps being asked therefore exposes a broader lacuna in management accounting. Finance can usually explain what a process consumes, and it can often estimate what a proposed change might save. Yet it is far less equipped to show which process components contribute value, which destroy it, which absorb risk and which create information or options whose economic effects may surface much later. Consequently, a transformation business case can become highly precise about one side of the equation while leaving the other largely narrative.
A ledger on the ledge of usefulness
The general ledger reports what a department or cost center consumes. Activity-based costing (ABC), where organizations have implemented it with sufficient discipline, pushes that resolution further by assigning costs to activities. Both approaches remain useful; nonetheless, their analytical center of gravity is consumption rather than contribution. They can tell management where resources were spent with increasing granularity, while offering much less visibility into what an individual activity economically produced.
Double-entry accounting, dating back 500 years to Venetian merchants, earns its reputation for symmetry, although the symmetry belongs primarily to bookkeeping. The two sides of an entry describe the same financial event, while revenue generally appears when a transaction is recognized rather than carrying a lineage back through the many process components that helped create it. A renewal, expansion, avoided loss, faster decision or improved customer relationship may depend on dozens of steps, yet the contribution of any one step rarely has an account to which it can be posted.
This creates an analytical asymmetry that can influence investment decisions more than finance leaders may realize. When a CFO or operating executive evaluates a proposed process change, the cost side often arrives quantified while the value side arrives as prose, judgment or a collection of indirect metrics. The quantified side therefore tends to carry disproportionate weight because it is already denominated in the unit in which the decision is made: money. Indeed, acknowledged uncertainty may be safer than one-sided precision, because the latter can carry the authority of a number while obscuring what the model omitted.
One process, many economic artifacts
Consider the process of customer onboarding. Operationally, it is a sequence of tasks needed to establish a customer, configure services, obtain approvals and move the relationship into a steady state. Economically, however, those same steps may establish relationship patterns that influence retention, create the account depth that enables a later cross-sell, generate behavioral and preference data whose usefulness compounds over the customer lifecycle, and reduce churn risk through investments made well before the customer has a reason to leave.
Embedded in that same process may be approval controls whose original compliance rationale has waned, manual handoffs between systems that were never integrated, duplicate checks and wait times that gradually erode the loyalty the process was intended to build. Some components may therefore create value; others may protect it and still others may quietly consume it. Yet a conventional cost model compresses this heterogeneous mesh (or mess!) of economic activity into a single process cost, which is useful but incomplete.
Improving or transforming the process requires a more discriminating account of what each component is doing economically: which steps build value, which erode it, which create unnecessary friction, which absorb risk, which generate ancillary benefits and which perform economic work that becomes visible only after the step is removed. Without that component-level view, an efficiency initiative can readily eliminate something valuable simply because its cost was easier to suss than its contribution.
Sample customer onboarding — economic process model.
Economic process modeling (EPM) supplies that missing layer. As I described in a recent column, Business Transformation Needs a True Economic Approach Rather Than Guesswork, EPM decomposes a process into its constituent components—the information flows, human decisions, system actions and organizational touchpoints that make up the actual work—and then attributes economic effects to each across five dimensions: revenue contribution, cost and friction, risk exposure, option value and information value.
The component level matters because the economically significant finding often sits buried within the process as a whole. Two steps that look roughly equivalent on a process diagram may carry very different economic profiles once attribution is applied. A seemingly minor validation step, for example, may generate information that reduces downstream risk, while a more conspicuous approval step may be largely vestigial. A cost-only review can easily misread the two because it sees effort more readily than consequence.
Pagel describes the capability she wants in similarly practical terms: the ability to see processes at an organizational level, understand what they cost in aggregate and then break those economics down step by step. That level of resolution helps because process transformation decisions are rarely made at the level of an abstract end-to-end flow; they are made by automating, eliminating, combining, outsourcing or redesigning individual components. Consequently, finance needs an economic view at the same level where the design decision is actually being made.
The oft-ignored value of information itself
Information value is where this analytical oversight may be most acute, particularly because most processes today generate data as a byproduct of execution. For example, a credit review produces repayment-behavior signals, a claims intake creates fraud indicators, and a procurement approval accumulates supplier-performance evidence. Those outputs may have future utility well beyond the transaction or process that generated them, even though conventional cost accounting typically has no place to represent them.
Most organizations, however, still treat much of this data primarily as documentation, exhaust or a compliance burden rather than as a potentially monetizable asset. A process redesign can therefore appear efficient while externalizing, degrading or destroying information whose economic contribution was absent from the business case. Infonomics, the discipline of treating information as an economic asset with attributable value, provides the grounding for this dimension of EPM and helps expose value that can otherwise disappear during an ostensibly sensible transformation.
From cost review to capital allocation
EPM extends cost accounting by adding an economic perspective the ledger wasn’t designed for. Sure, cost remains an indispensable computation. However, the business case becomes materially more complete when the components proposed for automation or elimination are also evaluated for revenue contribution, risk absorption, optionality, information yield and the friction they create or remove.
This can change the quality of the capital-allocation discussion. A step that costs $500,000 annually certainly may be a strong automation candidate, yet the savings figure is incomplete if the same step prevents $2 million in avoidable losses, preserves a customer relationship, generates valuable information or creates an option the business may need later. Conversely, a relatively inexpensive step can still be economically destructive if it adds delay, rework or customer attrition. The point is not to manufacture spurious precision around every benefit; rather, it is to make the relevant sources of value visible, estimable and subject to the same scrutiny as cost.
Which brings us back to Pagel and the question she hears whenever she proposes a change: “How much money are we looking to save here?” Savings are only one side of the economic case. The more consequential question may be what each affected component contributes today, what value may disappear if it is changed, and what new value the redesigned process could create.
Indeed, a transformation can look compelling when the savings are visible and the value at risk remains invisible. Economic process modeling gives finance a way to juxtapose both in the analysis, so that a proposed change can be judged not merely by what the organization expects to spend less, but by what the work itself is actually worth.
The advent of AI is precisely why organizations need technology economists, not just IT finance professionals.
IT finance is primarily concerned with budgeting, accounting, cost allocation, depreciation, chargebacks and financial reporting. These disciplines remain important, but they assume a relatively stable relationship between technology spending and business outcomes. AI breaks that assumption.
Technology economics asks a fundamentally different question: How d
The advent of AI is precisely why organizations need technology economists, not just IT finance professionals.
IT finance is primarily concerned with budgeting, accounting, cost allocation, depreciation, chargebacks and financial reporting. These disciplines remain important, but they assume a relatively stable relationship between technology spending and business outcomes. AI breaks that assumption.
Technology economics asks a fundamentally different question: How do technology investments create, destroy, shift or delay economic value?
AI introduces a set of economic dynamics that traditional IT finance was never designed to evaluate.
AI creates non-linear economics
In traditional IT, spending $10 million typically produced a somewhat predictable capacity increase or operational improvement.
With AI, a $10 million investment might generate $100 million in value. It might generate no value at all. It could increase costs while appearing successful. It could also create strategic advantages that do not show up in financial statements for years.
A technology economist studies the relationship between technology inputs, organizational capability, productivity outcomes and economic value creation.
IT finance largely records the spending.
AI changes the economics of labor
AI is not merely another technology platform. It acts as a form of digital labor.
Organizations now face questions such as:
Should work be done by humans, AI, automation or a combination?
What is the marginal cost of an AI-generated transaction versus a human-generated one?
How does AI affect productivity elasticity?
When does AI create labor substitution versus labor augmentation?
These are economic questions, not accounting questions.
AI simultaneously creates technology inflation and deflation
A fascinating paradox is emerging: AI can reduce costs in some areas while dramatically increasing costs elsewhere.
For example, fewer coding hours. More GPU costs. Lower service desk costs. Higher cybersecurity costs. Reduced consulting expenses. Increased data management expenses.
Technology economists study entire economic systems and value chains.
IT finance often sees only line items.
AI requires measuring economic outcomes, not technology outputs
Historically, organizations measured projects delivered, systems implemented, budgets achieved and uptime percentages.
The AI era requires measuring:
Revenue generated
Margin improvement
Risk reduction
Productivity gains
Decision quality improvement
Time-to-market acceleration
Innovation capacity
Technology economists focus on these outcome measures.
This is one reason why AI performance measurement frameworks, including AI-focused balanced scorecard approaches, are becoming increasingly important.
AI introduces massive opportunity costs
One of the largest AI risks is not technological failure.
It is investing in the wrong AI initiatives.
A bank might spend $50 million building an AI solution that saves $5 million annually while ignoring another opportunity that could have generated $500 million in new revenue.
Technology economics focuses on capital allocation efficiency, opportunity cost, marginal returns and portfolio optimization.
These concepts sit outside traditional IT finance.
AI makes technology a strategic production function
Historically, technology supported the business.
Increasingly, technology is the business.
In many industries, AI determines customer experience, operating efficiency, innovation speed and competitive advantage.
Technology is becoming a primary production factor alongside labor, capital and natural resources.
Organizations therefore need experts who understand the economics of technology as a production asset.
AI creates new forms of technical and economic debt
Many organizations are deploying AI rapidly without understanding:
Long-term infrastructure costs
Model maintenance costs
Data quality costs
Governance costs
Security costs
Regulatory costs
A technology economist examines the total lifecycle economics.
The cheapest AI solution today may become the most expensive solution over the next decade.
Why this matters
The central challenge of the AI era is no longer “Can we build it?”
The challenge is, “Should we build it, where should we deploy it, what value will it create, what risks will it introduce and what is the optimal economic allocation of technology capital?”
Those are technology economics questions.
IT finance professionals are essential for controlling and reporting technology spending.
Technology economists are essential for determining whether that spending creates sustainable economic value.
As AI becomes embedded into every business process, the organizations that outperform will not necessarily be those with the biggest AI budgets. They will be those that best understand the economics of technology itself — how AI, data, infrastructure, labor, risk and innovation combine to create measurable business value. That is the domain of technology economics.
Somebody in finance has already forwarded you the pricing comparison. A downloadable frontier-class model at a fraction of what you’re paying now, with the obvious question attached: why are we still on flagship rates? The honest answer comes in two halves. Open weights are the best thing to happen to enterprise AI buyers since the category existed, and switching to one still won’t lower your bill this year. Both of those are true, and the gap between them is where the use
Somebody in finance has already forwarded you the pricing comparison. A downloadable frontier-class model at a fraction of what you’re paying now, with the obvious question attached: why are we still on flagship rates? The honest answer comes in two halves. Open weights are the best thing to happen to enterprise AI buyers since the category existed, and switching to one still won’t lower your bill this year. Both of those are true, and the gap between them is where the useful work sits.
How we got here
In mid-July, a 2.8-trillion-parameter open-weight model shipped with performance close to the commercial frontier, and the full weights followed ten days later under a custom license. Markets moved before Washington did. Semiconductors were hit hardest that session and one widely held chip ETF finished the week almost 9% lower (CNBC). Washington began weighing restrictions on open-weight models soon after, and the industry answered inside a fortnight.
On July 24, twenty-five companies published a letter asking policymakers to leave downloadable model weights alone (Tom’s Hardware). Not a single founding signatory sold access to a closed-frontier model. Three major labs were absent at launch, two signed within 72 hours (TheNextWeb) and the roster passed 270 organizations inside ten days (Forbes). The sole holdout published its position days later, agreeing with much of the letter while disputing two safety claims.
What open weights hand you
Start with what genuinely changed, because it’s larger than the coverage suggests. A downloadable model at frontier-class capability puts a permanent public floor under what that capability can be sold for. But no supplier prices at whatever the market will bear once a comparable input is obtainable elsewhere, and that shift doesn’t reverse.
Stakeholders already know how to think about this. They just haven’t been filing AI under the right heading, which is concentration risk. A single provider holding a load-bearing production input, controlling both pricing and release schedule, would sit on the risk register in any other procurement category. The only reason AI was able to bypass this was that there was no alternative worth naming. Now there is one. The leverage shows up at renewal whether or not you ever deploy an open model, since the negotiating position changes the moment the alternative becomes credible.
It changes what you can responsibly commit to, as well. Until now, a multi-year AI investment has meant betting the program on one supplier’s pricing decisions and deprecation schedule, and that’s a hard paper to take into an investment committee. Commitments get easier when the input underneath them has a substitute. Workloads governed by data residency rules come back into scope too, and for some companies that means markets they’d written off.
Investors’ point of view is a little different in this scenario, and probably more accurate. Valuations built on sustained pricing power at the model layer assume something the capability data no longer supports. As models converge, the primary durable margin moves toward distribution, proprietary data and internal workflows that the customers cannot rip out. This happens to be the ground that the coalition’s founding signatories already hold.
None of that requires a single enterprise to switch models. So, the case against restrictions is a real one, whatever mix of principle and self-interest sits behind it. And note one of its own asks: public funding for shared evaluation frameworks, which the signatories evidently agree don’t exist yet.
The monopoly is breaking, just not on the scoreboard everyone watches
Stanford’s 2026 AI Index puts the leading closed model ahead of the leading open model by 3.3% as of March 2026, having been 0.5% ahead in August 2024 (Stanford HAI). The same chapter records six labs clustered inside 25 Arena Elo points at the top and reads that convergence as pushing competition toward cost and reliability. For most enterprise work, a 3.3% capability gap is not a reason to pay a multiple.
Market share tells a different story. Menlo Ventures, surveying 495 US enterprise AI decision-makers, puts three vendors at 88% of the enterprise LLM API market between them, on 40%, 27% and 21% (Menlo Ventures). The same research found enterprises tend to stay with whichever vendor they picked, upgrading within that provider even where switching costs are low.
Both are true and reconciling them is the point. Suppliers price differently when they know you can leave, and that holds whether or not you ever. The alternative never has to be used to change what you pay. So, the pricing monopoly is gone while market share sits exactly where it was. Pricing power was the monopoly that mattered to buyers, and open weights broke it.
The headline price is not the cost
This part is arithmetic. On published rates one recent open model looks roughly a third the price of a leading commercial system. Cost per completed task tells a different story, and the firm that measures it states the mechanism plainly: because cost tracks real token usage, models producing longer answers or more reasoning bill more per task even at identical per-token prices (Artificial Analysis). One current frontier model cost about twice its predecessor per task on that measure, driven entirely by token consumption and not by any price change.
It’s not possible to move a production workload to a budget-friendly mode without having a per-tasking definition of good enough, and literally no one has that handy. Ask what accuracy a workflow requires and you’ll hear crickets, or a number invented on the spot. Ask what it currently achieves and you’ll get the same silence. Same shrug, different meetings. Until both questions have answers, a pricing table is just somebody else’s workload dressed up as your business case.
I work on AI infrastructure, and the primary hurdle that I keep hitting isn’t a technical one. Writing down what good enough means requires somebody to put their name on a number they’ll be held to later. That’s an organizational decision, and not an engineering one, which is exactly why these documents don’t exist in most companies. Teams spend multiple quarters comparing models but barely spend a week agreeing what exactly they’re comparing them for. The related thing I’d say from that seat is that most groups believe they evaluated a model when what they did was try it. Someone ran twenty prompts, liked what came back, and the decision got made in the room. That’s a demo. Demos flatter every model about equally, which is why they can’t tell you whether the cheap one is costing you anything.
Self-hosting won’t rescue the math for most buyers either. A frontier-scale checkpoint runs well past a terabyte, so outside the regulated cases above, the win shows up as hosted providers competing for your workload.
What will actually move your bill
Two things will, and neither of them is a model release. First is the eval infrastructure, since it turns any price difference into a decision that you can defend. Scrape a few hundred real queries from the peak-load and freeze them as your golden test set. Get the people who own the business outcome to write down what a good response looks like, in specifics instead of adjectives. Test the existing model first and generate the scores. Most teams underestimate this step, and it’s the one that makes every future comparison possible.
The second is your own compliance position, which the August decision did nothing to simplify. Every proposal in circulation points at documentation and audit trails, and the holdout lab’s own position points the same direction from the opposite side of the debate. What reaches the buyer either way is a demand for evidence about what your systems can do and what they did. Most 2027 budgets don’t carry that line.
And on the stock market question
Expect volatility on release days and don’t mistake it for repricing. A capable open model lands and chip stocks sell off within hours, one such session costing a single chipmaker close to $600 billion (CNBC). They recover over the following weeks.
My read is that markets keep filing these as demand shocks when they are supply-side price events. Cheaper capability has driven adoption and compute consumption up together every time, which is the opposite of what a selloff assumes. So, keep the two conversations apart. A chip selloff tells you about supplier margins and nothing about your own AI spend, and boards that conflate them freeze budgets during a dip or wave them through during a rally. What would genuinely reprice this sector is a regulatory outcome raising the cost of shipping capability, or an adoption curve that flattens.
Every IT leader knows the pattern. One team builds a report in a spreadsheet. Another spins up a workflow with slightly different logic to answer the same question. A dashboard is shared across three departments, and within a week, nobody can say for certain where the underlying numbers came from.
Instead of freeing up capacity, self-service has quietly become another form of manual work: chasing down mystery logic, reconciling duplicated effort, and answering question
Every IT leader knows the pattern. One team builds a report in a spreadsheet. Another spins up a workflow with slightly different logic to answer the same question. A dashboard is shared across three departments, and within a week, nobody can say for certain where the underlying numbers came from.
Instead of freeing up capacity, self-service has quietly become another form of manual work: chasing down mystery logic, reconciling duplicated effort, and answering questions nobody wants to own.
Self-service was never the risk
It is tempting to read that scenario as an argument for tighter control — fewer people building, more requests routed through a central team, more approvals before anything ships. That reaction is understandable, but it solves the wrong problem.
Self-service fails when there are no shared rules for access, quality, documentation, and ownership. Without those guardrails, speed doesn’t produce faster decisions, just more confusion distributed across spreadsheets and more shared drives.
The real tension is that most organizations have been offered only two options. Either lock everything down, or let everyone build whatever they want and hope it holds together. Neither one scales.
IT as the paved road, not the checkpoint
Centralizing data was never the hard part. The real challenge is the last mile: turning that data into decisions and actions the business can actually trust. Closing that gap does not mean IT owns every rule, calculation, and exception that determines how work gets done.
It means IT builds the paved road — trusted access, approved workflows, reusable templates, and visibility into what is being built — while the people closest to the work own and adapt the business logic that runs through it.
That division of labor changes what “governance” means in practice. Instead of a gate every request has to pass through one at a time, governance becomes the infrastructure that keeps logic visible, understandable, repeatable, and auditable by design. When the fastest way to answer a question is also the most trusted way, analysts do not need to be talked into compliance, and IT does not need to inspect every workflow to know it will hold up. It is simply how the work gets done.
Freedom and guardrails, together
Governed self-service isn’t about choosing between speed and control, it’s about giving each side of the equation what it actually needs to trust the other.
Governed self-service gives analysts:
Access to trusted data
Reusable templates and workflow patterns
Clear rules for sharing and automation
A way to document logic
Support when a workflow needs to scale
And it gives IT:
Visibility into who is building what
Better governance over access and data use
Fewer one-off requests
Less mystery logic floating around the business
A cleaner path from individual workflow to team-wide process
What this looks like in practice
Papa Johns’ finance team offers a useful example of governed self-service in action. The team handles risk-sensitive, high-volume work — franchise billing, royalty calculations, aggregator commissions, and SOX-compliant period close — across a global, multi-currency franchise business.
Historically, much of that logic lived in spreadsheets and disconnected tools, separate from the systems of record and hard to audit when workflows changed.
Using Alteryx, Papa Johns rebuilt franchise billing and reconciliation as a governed workflow that runs directly against its Google BigQuery environment, so calculations execute where the data already lives rather than being copied out to another location.
With Alteryx, complex calculations are visible, repeatable, and auditable. Finance users can ask natural language questions, such as comparing month-over-month figures, and receive immediate answers while also seeing how logic is applied. IT can support governance without becoming a bottleneck.
The partnership between the business and IT was key to scaling success. Michael Wyant, VP of Enterprise Data and Corporate Solutions, and his team are responsible for governance and data pipelines. The finance team owns the business logic and can adapt it as requirements change.
Each side owns the part of the problem it understands best.
The result is a workflow that finance trusts, that IT can stand behind, and that scales as a template for other high-stakes processes across the business. That’s the kind of outcome that governed self-service is meant to produce.
Fewer surprises, more trust
None of this requires IT to slow analysts down or analysts to work around IT. When self-service is built on shared standards, analysts stop waiting on tickets, IT stops chasing down mystery logic, and the business gets answers that hold up the moment someone asks, “Where did this number come from?”
Alteryx supports that model by giving business teams a governed way to build and adapt workflows themselves, while giving IT the visibility, controls, and security required to support it all at enterprise scale.
The goal was never more control for control’s sake. It is fewer surprises, less rework, and more answers the business can actually trust.
Ready to see what governed self-service could look like for your team? Explore the AI-Ready Starter Kits to get started.
There’s a scenario that plays out every day across the enterprise. A salesperson is about to close a major deal. They want to know what their commission will be. They type the question into ChatGPT or their favorite AI assistant. What comes back is a thoughtful, well-written explanation of how software companies typically structure sales compensation.
The one thing it won’t tell them is what their commission will actually be if they close this specific deal. That gap b
There’s a scenario that plays out every day across the enterprise. A salesperson is about to close a major deal. They want to know what their commission will be. They type the question into ChatGPT or their favorite AI assistant. What comes back is a thoughtful, well-written explanation of how software companies typically structure sales compensation.
The one thing it won’t tell them is what their commission will actually be if they close this specific deal. That gap between what AI can reason over and what it knows about your business is the defining challenge of enterprise AI adoption right now.
I call it the logic layer. And without it, AI gives you impressive sounding outputs that are often disconnected from how your business runs.
Why business logic lives with the analyst
One of the more persistent myths in AI is that analysts are on the verge of becoming unnecessary.
The reality is the opposite, and the logic layer is exactly why.
In an AI-enabled enterprise, analysts become more essential because they are closest to the logic and context that governs the business. They know which definition of pipeline matters and which edge cases matter in audit, merchandising, finance, or marketing.
I believe enterprises that succeed in the AI era will not be defined by how much AI they deploy but whether the people who understand the business own and control the intelligence that runs it.
If that ownership defaults entirely to IT or to a vendor’s black box, companies risk scaling systems they cannot fully adapt or audit. Giving business teams the tools and mandate to own their logic is what makes the AI system trustworthy and responsive to how your business runs.
That is why I see analysts as the architects of this next phase.
What the logic layer looks like in practice
Let me return to the commissions example, because it illustrates the concept precisely. Right now, when a salesperson needs to know their commission on a deal, they send a message to the commissions analyst. That analyst has their own spreadsheet — because comp plans change every quarter, with spiffs and special programs layered on top. They run the math manually and send back an answer.
What if that same analyst built a simple, well-defined calculator that encoded their commission logic — the actual rules for your company, your plans, your programs — and connected it to the AI systems your salespeople are already using? Now when a rep asks what their commission will be on a specific deal, they get the right answer. Not a generic explanation of how commissions work.
And here’s the compounding value: that same logic can then be used by the annual planning agent to model the operational cost implications of different comp plans. It can feed the scenario planning model that runs hundreds of simulations for financial planning. The analyst who built it enables an entire network of AI systems to act on accurate, business-specific logic.
That’s the logic layer in practice: curated, purpose-built data assets and calculators that encapsulate how your business works, maintained by the people who understand it, deployable to every AI system that needs it.
What the logic layer requires
This is where I think most companies are still stuck. They’ve made the infrastructure investments. They have cloud data platforms and approved LLMs. But they’re asking those systems to do things they were never designed to do on their own.
The logic layer requires three things:
Purpose-built data assets. A narrow, clean, well-defined data set that reflects how you actually measure a specific business process.
Encoded business logic. This is the part that lives in people’s heads right now — the policies, the edge cases, the context that makes data mean something.
The ability to update it. Nobody runs a business to keep it the same. The logic layer has to be something that domain experts can update when the business changes.
A pragmatic path forward
The good news is that you don’t have to wait for a perfect architecture before you start building a logic layer.
Start with your highest-value, most-repeated business processes — the ones where an analyst is currently fielding the same questions week after week. These are the processes where encoding logic into a curated, AI-ready data asset delivers immediate, measurable value.
Then, empower your analysts to own that encoding — not IT. Give them low-code tools to do the work, and the mandate to treat that encoded logic as a strategic asset they own and evolve as the business changes.
This is also where leadership posture matters.
I have said for a while that this should not be framed as a choice between business and IT. It is both. IT should set standards, manage infrastructure, establish security boundaries, and make approved AI capabilities available across the organization. But IT should not become the bottleneck for every piece of business logic the company needs to operationalize.
If this feels familiar, it should. We have seen this pattern before in enterprise technology. Infrastructure and platforms matter. But the last mile, the part that turns capability into business value, always depends on the people closest to the work.
AI is no different.
The companies that get the most from AI will be the ones that treat it like an operating model. They will automate core workflows, curate the right data, and empower analysts and domain experts to define the logic that makes AI useful and generate answers the business can use.
I recently had a chance to go deeper on these ideas on the Talking AI podcast. If you want to hear more of my thinking on the analyst’s evolving role, how the logic layer connects to agentic workflows, and why I think the next 18 months will be pivotal for getting this right, it’s worth a listen.
Every new frontier model release seems to spur a fresh round of doomsday articles. Just Google “the end of white-collar jobs,” and you’ll be bombarded with discourse on the end of modern work, the unraveling of the social contract between employees and organizations.
What I don’t see anyone talking about, however, and what I believe is a far more productive conversation, is the opportunity for knowledge workers.
Nobody understands critical business processes better
Every new frontier model release seems to spur a fresh round of doomsday articles. Just Google “the end of white-collar jobs,” and you’ll be bombarded with discourse on the end of modern work, the unraveling of the social contract between employees and organizations.
What I don’t see anyone talking about, however, and what I believe is a far more productive conversation, is the opportunity for knowledge workers.
Nobody understands critical business processes better than your line-of-business (LOB) employees. Not executives. Not IT. Not even the most advanced LLMs. These are your business analysts and RevOps professionals, your supply chain managers and finance leaders, and the employees whose expertise has been forged over decades.
For an enterprise to become truly intelligent, these workers must be involved in how AI workflows are built and deployed. Their guiding hand is the only way AI can learn and truly understand your business.
But what does this transition look like, and how can organizations start operationalizing AI in a meaningful way alongside knowledge workers? Let’s take a look.
What enterprise intelligence requires
Imagine walking your board through a set of financials and recommending specific actions. Then, in your next meeting, you walk everything back because your AI layer got the numbers wrong.
There is no faster way to kill an AI initiative than by delivering wrong outputs. Without trust, the whole system falls apart.
In our recent survey of 1,400 business and IT leaders, we found that while over 90% of organizations are using AI, only 28% trust it to support decision-making. As for how many organizations scaled their AI pilots into production, the number was just under 25%, suggesting a very strong correlation between trust and operationalization.
An intelligent enterprise, then, is an organization that has trustworthy AI embedded across the business.
At Alteryx, we say the results of any AI system must follow our VURA framework: an AI system and its outputs must be visible, understandable, repeatable, and auditable. In other words, two people need to be able to go to AI with a question and arrive at the same answer; anyone who uses AI in their workflows must be able to explain how their AI system arrived at that answer.
Who’s responsible for operationalizing AI?
Enterprise intelligence is about trustworthy AI deployed throughout key business processes, but who’s ultimately responsible for these AI systems and processes: IT teams or knowledge workers?
Let’s say you want to use AI in your Sarbanes-Oxley process, e.g., your journal entries, revenue recognition, access controls, etc. Before IT can help you build a new AI workflow, IT must first understand your Sarbanes-Oxley process in great detail. Then, they have to code a tool your finance team can trust.
It’s possible, sure. But creating this solution would take an inordinate amount of time. Then, when a new regulation comes along or you have an acquisition, the whole thing falls apart. You have to get back in line with IT to retune everything.
Moreover, if your books don’t balance out or if you fall out of compliance, IT does not want to have that responsibility fall on them. You can see why ownership of AI systems and workflows must sit with LOB workers. They are the only ones with the expertise to ensure the veracity of AI’s outputs. They are the only ones who can successfully shape and define its logic and oversee its ongoing execution.
Data is the fuel. Business logic is what keeps AI on course.
Finally, there’s the question of data. We’ve all heard “bad inputs, bad outputs.” Seeing as I’m the CEO of a data analytics company, you might expect me to say that reliable data is the end-all, be-all when it comes to trustworthy AI outputs.
And while it’s absolutely essential, it’s only the first step.
Aggregating your enterprise data into a cloud data platform is immensely useful. All of that data becomes readily accessible. You gain a single source of truth across teams and workflows. But you can’t point your LLM at a cloud data platform and ask it to make sense of your data for a complex business process.
Again, you need the people who understand these critical processes to guide your LLMs to interpret the right data in the right way. This is what will make your AI systems visible, understandable, repeatable, and auditable. Yes, you need clean, reliable data. But more than that, you need business logic around that data, and that can only come from your knowledge workers.
The five pillars of enterprise intelligence
At the highest level, enterprise intelligence rests on five core pillars:
Trustworthy, transparent data
Empowered business analysts
Shared responsibility across the C-suite
Cross-functional collaboration
Leadership that evolves alongside AI
Each pillar reinforces the same core idea: AI only becomes valuable when it’s grounded in reliable data, shaped by real business expertise, supported by executive ownership, and scaled across teams that can put it to work to improve their daily processes.
Tap into the intelligence all around you
As a business leader looking to build an intelligent enterprise, the most important questions you can ask are the ones around operationalizing AI in key business processes. What would it take for you to trust AI’s outputs? What would make AI-powered processes superior to your current ones?
Once you have those answers, engage your LOB workers immediately. Give them ownership and autonomy. Rather than asking AI to replace them, lean into their intelligence. Let your knowledge workers use their expertise to amplify, shape, and govern AI. Their business mastery is what makes enterprise intelligence possible.
“You’re right,” the LLM says. “I was mistaken.”
Have you ever read these words during an AI workflow? Nothing kills trust faster than incorrect outputs. It’s no wonder, then, that only a quarter of businesses today fully trust AI to support decision-making and forecasting.
And yet, we know AI is business critical. Nine out of 10 businesses are using it; 64% say it’s powering innovation.
So, how do you bridge the gap from experimentation to trustworthy deploymen
Have you ever read these words during an AI workflow? Nothing kills trust faster than incorrect outputs. It’s no wonder, then, that only a quarter of businesses today fully trust AI to support decision-making and forecasting.
And yet, we know AI is business critical. Nine out of 10 businesses are using it; 64% say it’s powering innovation.
So, how do you bridge the gap from experimentation to trustworthy deployment? How do you get verifiable, reproducible results from AI at scale? In this article, I’ll show you the framework that’s powering AI success for leading organizations.
Why organizations still don’t trust AI
We asked 1,400 IT and business leaders what their biggest barriers to success with AI workflows were. One in two (49%) said inaccurate or biased outputs; 38% said it was a reluctance to allow AI to make decisions without human oversight.
Then, there was the data issue. Data readiness is an integral part of successful AI workflows. However, half of all organizations said they still faced poor quality or fragmented data. While you don’t need perfect data to start using LLMs, you absolutely need trustworthy data.
VURA: The framework for trustworthy AI
Closing this trust gap requires two things. First, organizations need a logic layer that connects AI systems to the people who understand the data and business best. Line-of-business teams and analysts cannot sit on the sidelines. They need to help build and validate AI workflows so the logic behind AI’s outputs reflects how the business actually operates.
Second, AI workflows and processes should be visible, understandable, repeatable, and auditable. Together, these principles form VURA, a framework we developed to help organizations build and scale trustworthy AI systems. These guidelines will help build trust in your data and your AI’s outputs. You’ll need both if you want your business to build enterprise intelligence. What follows are the four pillars of VURA.
Visible
Visibility is transparency. Your AI workflows shouldn’t be a black box regarding the data used and the logic applied. Every employee using AI tools should be able to answer two questions: “Where did this answer come from?” and “How did we draw that conclusion?” Otherwise, employees may be working from incorrect information. They could give your customers faulty intel or make important decisions with serious downstream effects.
If those answers are still unclear, you may need to tighten your governance or reconsider whether your current AI and data solutions are working. Visibility becomes especially important when AI is used across teams.
Understandable
It can almost feel like science fiction when tools like ChatGPT or Gemini take the most complicated or vague of prompts, parse through them, and give you an intelligent, thoughtful answer.
However, this low threshold for asking and answering virtually any question in natural language isn’t an excuse for glossing over business fundamentals. Your AI systems must be able to explain the logic behind their outputs to even non-technical business users, and your business experts must be able to validate those outputs.
Repeatable
Repeatable means that with the same AI tools, data, prompts, and business logic, AI will give you the same answer every time. Two people should be able to go to AI with the same question and arrive at the same answer. If an AI system or workflow gives you an excellent answer followed by one that’s clearly wrong, it’s not ready for operationalization. You can’t trust it.
Repeatability also requires documentation. When teams identify prompts or processes that help produce reliable outcomes, those should be recorded and shared.
Auditable
An auditable AI process means you can see what happened. There’s a trail. If there’s an answer or report that seems off, you should be able to identify who owns the workflow, what data and prompts were used, what logic the system followed, and where human judgment and oversight were involved. Auditability is a check and balance for both your AI systems and the human engineers working behind the scenes.
Start building trustworthy AI systems today
AI can only deliver scalable business value when it’s grounded in trustworthy data and business logic. To operationalize these systems, you’ll have to ensure your AI workflows are visible, understandable, repeatable, and auditable.
Alteryx is the transformation and business logic layer that helps you move AI from experimental pilots to trustworthy production. It connects to data wherever it lives, helps business users apply their expertise to AI-powered workflows, and instills the guardrails needed for both your data and your AI systems.
With Alteryx, the people closest to the business can shape how data is prepared and applied, while IT gains the governance and auditability required for enterprise use. That’s how AI outcomes become trustworthy. That’s how enterprise intelligence is built.
If you’re a senior analyst, you’ve probably faced a dataset that used to open in seconds but now takes minutes. Or maybe you’re up against a formula that worked fine last quarter, but now shows an error because someone renamed a tab in a file three layers upstream. Ever spent an afternoon figuring out whose numbers are right when colleagues send back a few different “final” versions of the same report? None of that is a personal failure, but it is a sign that the volume an
If you’re a senior analyst, you’ve probably faced a dataset that used to open in seconds but now takes minutes. Or maybe you’re up against a formula that worked fine last quarter, but now shows an error because someone renamed a tab in a file three layers upstream. Ever spent an afternoon figuring out whose numbers are right when colleagues send back a few different “final” versions of the same report? None of that is a personal failure, but it is a sign that the volume and complexity of your work has outgrown what a spreadsheet was built to handle.
The cost is more than just your time and inconvenience. When reporting slows down, multiple versions of a number circulate before anyone catches it, or one person’s spreadsheet logic is the only process your team has for something that matters, that’s a risk to the business.
Plenty of solid analysis still belongs in a spreadsheet, but as your data and stakeholders grow, the balance between preparing data and analyzing it changes. If prep now takes more of your week than analysis does, the tool has become the bottleneck — not you.
Watch for concrete signs you’ve hit that ceiling, what a platform “built for scale” needs to do differently, and how to decide whether it’s time to move.
When spreadsheets stop being enough for your data analysis
Every spreadsheet has a hard ceiling, and it’s lower than people might expect. Microsoft’s own published specifications cap every worksheet at 1,048,576 rows by 16,384 columns, regardless of your computer’s memory or Excel version. Once a dataset crosses that line, rows don’t get flagged — they simply don’t load, and it’s easy to miss.
The bigger risk is accuracy. A 2024 literature review published in Frontiers of Computer Science, covering more than 30 years of spreadsheet research, found that 94% of spreadsheets used in business decision-making contain errors that create real risk of financial losses and operational mistakes. Most analysts already understand that the more a spreadsheet grows past its original design, the harder it gets to trust every formula in it.
Alteryx’s own research backs this up from the analyst’s side of the desk. The 2025 State of Data Analysts in the Age of AI report, a global survey of 1,400 data analysts, found that 76% still rely on spreadsheets for data preparation, even as AI tools reshape the rest of their workflow. Manual prep work isn’t a habit analysts choose, but it’s still the default because nothing else is in place yet.
Spreadsheets are ultimately designed for individual calculation, not for shared, repeatable, large-scale analysis. Asking them to do that job is where the cracks start to form, and when you need to start thinking of an alternative.
What “built for scale” means
“Scale” gets used loosely in analytics marketing, so it’s worth being specific about what a platform needs to do differently than a spreadsheet. It comes down to four tasks:
Connect to data where it already lives
A spreadsheet only knows what you paste into it, which means every report starts with an export, a download, or a copy-paste job that’s already slightly out of date by the time it’s finished. Platforms that support scale should be designed to connect, transform, and prepare AI-ready data by connecting natively to a broad range of enterprise applications, databases, and cloud platforms. This way data can be pulled in and refreshed rather than manually re-exported every reporting cycle. For an analyst, that means less time reconciling which export is current and more time on the analysis itself.
Prepare and blend without rebuilding from scratch
Scalable platforms should also give analysts a drag-and-drop canvas for cleansing, blending, and reshaping data from multiple sources, with code-friendly options like Python and SQL available for analysts who want them. It should allow for logic to be built once and held to a standard we call VURA: visible, understandable, repeatable, and auditable. Analysts shouldn’t have to deal with a chain of formulas that only one person fully understands, and that visibility matters as much as the automation. When you build a workflow as a series of documented steps, colleagues can review, troubleshoot, or take over in a way a dense formula chain rarely allows.
Automate the workflow, not just the calculation
The real definition of scale for an analyst is a process that runs without being rebuilt by hand. An essential component of that is workflow automation and orchestration so analysts can schedule and reuse any workflow they build.
Report without abandoning familiar formats
Leaving spreadsheets behind for analysis doesn’t mean stakeholders lose the outputs they’re used to. Reporting tools can generate tables, charts, and formatted outputs in PDF, HTML, or Excel, closing the loop between analysis and the people who need to read the result.
Spreadsheets vs. a scale-ready analytics platform
The differences between spreadsheets and analytics platforms are less about features and more about the needs that develop as your data and your team grow.
Consideration
Spreadsheet
Scale-ready analytics platform
Data connections
Manual export/import from each source; data goes stale as soon as it’s pasted in
Native connections to databases, cloud platforms, and enterprise apps that can refresh on demand
Repeat work
Rebuilt or copied by hand each cycle
Built once as a workflow, then reused and scheduled
Row and file limits
Fixed worksheet ceiling regardless of vendor
Designed to process large volumes without a hard row cap in the tool itself
Auditability
Hard-to-trace formulas and edits
Workflow steps that are visible, understandable, repeatable, and auditable (VURA)
Collaboration
Version conflicts, emailed copies, confusion over which file is current
Shared workspace with a single source of truth for a given workflow
Matching the platform to where your team is right now
“Scale” doesn’t mean the same thing for a 5-person team tracking budgets as it does for an enterprise running hundreds of scheduled workflows. It’s important to consider your team’s specific size and overall org structure before you shortlist any options.
If your team is still primarily working out of Excel or CSV files and wants to reduce manual, repetitive spreadsheet work, there are several platforms built just for these types of applications. Alteryx One Starter Edition, as an example, is built specifically for that transition — code-free data prep accessible from a browser, aimed at teams getting started rather than running complex automation.
You don’t need to know your exact tier or platform before you start a conversation with your team, but the conversation can go faster when you can describe your situation in terms of how many people touch the data, how often it needs to run, and who needs to see the output (rather than starting from a feature list).
Governance doesn’t disappear with spreadsheets
It’s tempting to think that moving off spreadsheets automatically solves governance, but that just changes what governance looks like. TechTarget’s coverage of data and analytics governance requirements notes that organizations should look for scalable, modular platforms that can adapt as needs change, rather than assuming governance is solved by the platform switch alone.
It’s important to investigate whether or not a platform supports this with governance and administration capabilities such as role-based access controls, audit logs, and version history. They’re built to give IT the oversight it needs while giving analysts the flexibility to build and run their own workflows.
A quick readiness check
Before you bring a platform comparison to your team or your leadership, it helps to be specific about what’s driving the need. What follows are a few questions worth answering honestly:
Are you regularly working with datasets that approach or exceed Excel’s row limit, or that make Excel noticeably slow to open and calculate? Slow-loading files and truncated imports are usually the first visible sign, not the first real cause.
How much of your week goes to gathering and cleaning data by hand, instead of analyzing it? If prep consistently outweighs analysis, that ratio is the problem — not a single unwieldy file.
If you left tomorrow, could someone else pick up your spreadsheet-based process without you walking them through it? A process that exists only in one person’s head is a significant business continuity risk.
Do stakeholders currently receive conflicting versions of the same report because multiple people are editing copies independently? That’s a flag that the workflow needs a single source of truth.
Does your organization need an audit trail for how a number was calculated, not just what the number is? Regulated or audited environments tend to outgrow spreadsheet-based tracking quickly.
If you answered yes to two or more of these, this is a reasonable signal that it’s worth having a conversation about moving beyond spreadsheets.
The fastest way to know whether a platform fits your workflow is to run your own data through it rather than a demo dataset. You can start a free trial of Alteryx One to test connectivity, data prep, and workflow automation against the kind of analysis you do every week.
The old way of creating reports is almost cliché, but only because it remains so pervasive.
It’s a familiar scene: A teammate pings you at 4:57 pm asking for a last-minute report. The data and business logic you need live across 10 spreadsheets, in five Microsoft Teams threads, and in an email from two months ago that you can’t seem to find.
But that was the old way. What happens when you use AI for the same situation?
Let’s find out.
What working with AI ofte
The old way of creating reports is almost cliché, but only because it remains so pervasive.
It’s a familiar scene: A teammate pings you at 4:57 pm asking for a last-minute report. The data and business logic you need live across 10 spreadsheets, in five Microsoft Teams threads, and in an email from two months ago that you can’t seem to find.
But that was the old way. What happens when you use AI for the same situation?
Let’s find out.
What working with AI often looks like
Your stakeholder pings you, asking for a report based on a massive tax reconciliation spreadsheet.
This spreadsheet is a beast, chock-full of tabs, formulas, and data that’s been copied and pasted from several enterprise data sources.
“Sorry for the last-minute ask,” they say, “but can you just throw AI at this?”
You go to your LLM prompt library, select a robust prompt, and input it into Claude, along with the spreadsheet.
Four seconds later, you get over 1,700 lines of Python code. Somewhere inside, there appear to be all the data transformations, calculations, and visualizations you need to build your report.
But there’s a hiccup.
Your stakeholder remembers that your tax jurisdictions change four times a year and wants to ensure that you can make any necessary changes.
Sure, you think, that shouldn’t be a problem. I can probably find that line of code somewhere …
Also, there are three subsidiaries. Someone else handles those taxes, so you’ll need to filter those out.
Finally, your stakeholder remembers that your CFO will want to sign off on this and that your auditor is coming tomorrow. They’ll both want to see the logic behind your report.
Suddenly, parsing through and validating hundreds of lines of AI-generated code seems far more difficult and time-consuming than you’d hoped.
VURA: The missing piece
While AI can bring incredible levels of automation and speed, those are only force multipliers when directed strategically.
“I can get an infinite number of PowerPoints out of the AI systems if I want that,” Ethan Mollick recently told me during our executive exchange. “It may even be good content, but if it doesn’t serve the purpose you need it to, the productivity gains become a trap.”
Ethan’s point is that more isn’t always better; bringing four hundred PowerPoints to a sales call won’t help you close a deal. Likewise, instantly generating hundreds of lines of Python is unlikely to help your CFO feel confident in your AI’s vibe-coded report.
For an AI workflow to be trusted, it has to be Visible, Understandable, Repeatable, and Auditable, or VURA. You need to know what’s happening at every step of the process: where the inputs came from, how business logic was applied, and whether the outputs were correct.
So, how can you accomplish this?
The transformation and business logic layer
Let’s try a different AI-powered workflow. Same situation and model. Only this time, we’re going to add a visual transformation and business logic layer.
First, we go into Claude and type up a prompt, but instead of Python, we ask for an Alteryx workflow.
We open our workflow in Alteryx, and instead of hundreds of lines of AI-generated code, we see a visual canvas showing the entire tax reconciliation process.
It’s still an AI-generated workflow, but now, anyone in the organization can inspect it. They can see what data was used. Your analysts and domain experts can validate the logic. And you can add governance and repeat the process.
Suddenly, AI-generated workflows become far more trustworthy and scalable, giving you a foundation for enterprise intelligence.
The future of enterprise AI workflows
AI tools that can’t adapt when the business changes have short shelf lives, and rebuilding from scratch constantly drains tokens, time, and energy. Endless iterations create endless chances for inconsistencies and errors.
With a visual business logic layer, the people who know your business best — your business analysts, sales professionals, finance team, and more — can apply their expertise to your AI workflows and validate its outputs. They can see what’s happening at every step of your AI workflows so that every process is Visible, Understandable, Repeatable, and Auditable.
Speed and reliability are no longer mutually exclusive. Now, you can bring AI’s power and your business experts together to create something fast and reliable, the intelligent solution you need to create scalable business value.
There’s a version of this story you’ve probably lived. The close is approaching, someone pulls a number from a file that hasn’t been updated, and an hour later you’re untangling a discrepancy that shouldn’t exist. The fix takes 20 minutes. Finding the source took two days.
This is the part where most articles would tell you to ‘ditch the spreadsheet.’ But that’s not the real problem, and honestly, it’s a little insulting to the work you’ve actually done.
Your sprea
There’s a version of this story you’ve probably lived. The close is approaching, someone pulls a number from a file that hasn’t been updated, and an hour later you’re untangling a discrepancy that shouldn’t exist. The fix takes 20 minutes. Finding the source took two days.
This is the part where most articles would tell you to ‘ditch the spreadsheet.’ But that’s not the real problem, and honestly, it’s a little insulting to the work you’ve actually done.
Your spreadsheet isn’t the issue. The process built around it is.
The logic is real. The medium is the limitation.
Think about what lives in the workbooks your team maintains. How revenue maps to each entity. What counts as a valid reconciling item. The variance threshold that triggers a review. The intercompany elimination logic that took a year to get right. None of that is just data — it’s institutional knowledge. It’s business logic, and it belongs to finance.
The problem is that spreadsheets were never designed to share that logic, version it, or let anything else use it reliably. When a process lives in a file on someone’s desktop, it’s invisible to every system downstream. You can’t hand it off cleanly. You can’t audit it without opening every tab. And when the person who built it leaves, a piece of your operations leaves with them.
Why this matters more now than it did two years ago
A lot of finance teams are under pressure to adopt AI — for close acceleration, anomaly detection, forecast assistance, narrative reporting. The pitch is compelling. The results, so far, have been uneven.
Here’s why, and this part is specific to finance: AI can process data at scale and surface patterns quickly, but it cannot enforce your cost allocation methodology, validate your intercompany eliminations, or know what your organization has decided counts as an exception.
For a tax team, that means it can’t apply your jurisdiction mappings reliably. For an audit team, it can’t reproduce your evidence logic. For FP&A, it can’t honor the constraint assumptions built into your planning model. That requires logic that’s documented, governed, and repeatable — and if that logic is locked in spreadsheets AI can’t see, AI can’t apply it. So it guesses. In finance, a confident guess on a tax provision or a consolidation rule isn’t a minor error. It’s a liability.
The teams getting real value from AI are the ones who built the foundation first and then let AI work on top of it.
The question most teams haven’t answered yet
The shift that helps isn’t about which tool you use but where your process logic lives and who can access it. When your reconciliation rules, transformation logic, and validation criteria exist in governed workflows rather than locked files, the close gets more consistent, errors surface earlier, and handoffs get simpler.
But getting from here to there raises a real question most teams are still working through: what does that transition look like for a tax team, an audit function, or an FP&A group that has years of logic built up in Excel? What moves first, what stays, and what does a week of progress realistically look like?
That’s where the specifics matter — and that’s what we’ll get into next.
Most finance leaders at large organizations have made the right investments. A modern ERP, cloud data platforms, planning tools, and more. And now, increasingly, AI — for forecasting support, anomaly detection, close acceleration, and reporting at scale.
The technology stack looks right. But when the board starts asking about results, the returns are harder to point to than the investments were.
What your ERP was built to do — and what it wasn’t
Your ERP is exc
Most finance leaders at large organizations have made the right investments. A modern ERP, cloud data platforms, planning tools, and more. And now, increasingly, AI — for forecasting support, anomaly detection, close acceleration, and reporting at scale.
The technology stack looks right. But when the board starts asking about results, the returns are harder to point to than the investments were.
What your ERP was built to do — and what it wasn’t
Your ERP is excellent at what it was designed for: capturing transactions, enforcing accounting standards, managing the chart of accounts. It is the system of record, and it performs that job well.
But it doesn’t encode how your organization has decided to handle intercompany eliminations across a complex entity structure. It doesn’t carry your FP&A team’s cost allocation methodology, refined over three budget cycles. It doesn’t know what variance threshold triggers a controller review versus a VP escalation, or how your tax team has mapped jurisdictions for Pillar Two. That logic — specific, documented, organization-defined — isn’t in your ERP. It’s not in your data warehouse either.
For most finance organizations, it lives in spreadsheets. Sometimes in the heads of the people who built them.
Where AI runs into trouble in finance
There’s a finding that gets cited a lot in finance AI conversations: research from MIT found that 95% of organizations are seeing no measurable return on their gen AI investments. Bain & Company looked at the same picture and reached a different conclusion for finance specifically. The fastest payback from AI in finance comes from embedding it in workflows — not from running pilots. The distinction matters because it explains why so many finance AI efforts stall after the proof of concept.
AI can process data at speed and surface patterns across large datasets. What it cannot do is infer your business logic from raw inputs. Without that context, AI outputs in finance look confident but aren’t defensible — and in a function where auditability is a baseline requirement, that gap is not a minor limitation. It validates that trustworthy AI is critical for scaling workflows and AI pilots.
Our own survey of 1,400 IT and business leaders asked what their biggest barriers to success with AI workflows were. One in two (49%) said inaccurate or biased outputs. Further, 38% said it was a reluctance to allow AI to make decisions without human oversight. While you don’t need perfect data to start using LLMs, you absolutely need trustworthy data.
The layer that’s actually missing
The gap between your ERP and your AI ambitions isn’t a data gap. It’s a business logic gap — the layer where your organization’s specific rules, methodologies, and decision criteria live, and where AI needs to operate to produce outputs you can stand behind.
When that layer is built correctly — logic documented, workflows repeatable, outputs traceable — AI has validated, structured inputs rather than raw data it has to interpret. Outputs can be explained to auditors and to the board. And the sequencing question resolves itself: getting the process right is how you adopt AI.
What it takes to build that layer
Closing the gap takes more than a mandate to “use AI responsibly.” It takes three specific things, built and owned inside finance rather than handed off to IT.
A purpose-built data asset for each process. Not another warehouse but a narrow, well-defined data set scoped to one process that reflects how your team measures it, not just what your ERP happens to store.
Encoded logic, not tribal knowledge. The allocation methodology or the variance threshold that triggers escalation — built into a repeatable workflow instead of a senior analyst’s spreadsheet. The shift is building it once; in a form AI can use.
A way to update it when the business changes. Comp plans get revised, tax jurisdictions shift, and the chart of accounts gets restructured after an acquisition. Logic that can only be changed by submitting a ticket to IT will be stale before it’s deployed — the people who own the process need to be the ones who can adjust the rule.
None of this requires waiting for a perfect architecture. The highest-value starting point is whatever process has your analysts fielding the same question, the same way, every single cycle. Encode that one workflow first, connect it to the AI tools your team is already using, and the logic compounds from there: the same governed calculation that answers one controller’s question can feed the scenario model that runs your next planning cycle.
Every CFO I talk to right now is under some version of the same pressure: the board wants AI, the business wants faster answers, and the finance team is often still reconciling spreadsheets. The promise of AI in finance is real. But so is the gap between that promise and what most organizations are able to deliver.
I believe finance leaders need to be asking not simply, “How do we use AI?” but “What would make our data trustworthy enough for AI?”
That distinction m
Every CFO I talk to right now is under some version of the same pressure: the board wants AI, the business wants faster answers, and the finance team is often still reconciling spreadsheets. The promise of AI in finance is real. But so is the gap between that promise and what most organizations are able to deliver.
I believe finance leaders need to be asking not simply, “How do we use AI?” but “What would make our data trustworthy enough for AI?”
That distinction matters. AI-ready finance data is intentionally shaped for a specific business outcome, so we can trust what AI produces from it. In finance terms, it’s the difference between having transactions and being able to defend the numbers.
Finance data is uniquely messy, and important
Finance data is messy for rational reasons. We pull from multiple systems — ERP, CRM, payroll, procurement, planning tools, banks, data warehouses, and yes, still spreadsheets.
We live through reorgs, acquisitions, new products, and chart of accounts changes. And when the business cannot wait, we create manual workarounds to keep moving.
That complexity is the context in which we’re now being asked to use AI. It’s no wonder that so many initiatives stall.
The non-negotiables of AI-ready finance data
When Alteryx talks about AI-ready data, I translate it into a few non-negotiables. For finance leaders, this is where the concept becomes practical.
Purpose-built, not “all the data” – AI-ready data should be scoped to the decision or workflow at hand. If I am building a cash forecast, I do not need every field from every ledger table.
Clean and standardized – AI does not politely ignore bad inputs; it often amplifies them. That means your data needs to be deduplicated, standardized across dates, currencies, and units, and mapped to consistent hierarchies.
Combined across sources, with business context – Finance work is inherently cross-source. AI-ready data is joined and enriched so the dataset reflects business reality, not just system silos.
Traceable and transparent – This is where finance leaders should push harder than anyone else. AI-ready data has lineage. It is auditable and explainable, not just at the output layer, but in the data shaping behind it.
Governed and controlled – AI readiness is about data risk management as much as data quality. AI-ready data should live inside a governed process, not a series of hero spreadsheets and copy-paste steps.
Maintainable as the business changes – This is one of the hidden killers of AI initiatives. A one-time cleaned dataset is not AI-ready if it breaks the minute a new subsidiary is added, a cost center structure changes, or a revenue stream appears. AI-ready data has to be built through workflows that can be updated and re-run reliably, not through one-off cleanups.
Where AI-ready data creates value in finance
This is where the concept becomes real. AI-ready data is the difference between value and noise in some of finance’s most important workflows, including:
Close acceleration: When trial balance data, mappings, intercompany logic, and exception rules are standardized, finance can generate more dependable variance flags and automate more of the financial close and reconciliation process.
Cash forecasting: Better-connected bank data, AR/AP, billing schedules, and seasonality drivers make forecasts less likely to be derailed by missing or misclassified transactions.
Anomaly and fraud detection: Clean, aligned vendor master data, payment runs, approval chains, and PO matching help teams reduce false positives and investigate issues faster.
Revenue quality and leakage: When contracts, invoices, usage, CRM data, and credit logic are brought together in a way that reflects the actual economics of the business, AI can help surface patterns that matter.
Narrative reporting: Grounding LLMs in curated, reconciled variance drivers and approved definitions allows teams to draft commentary responsibly within clear guardrails.
Filling the AI data readiness gap
I’ve found that in most organizations, there’s a constant friction point between data engineering and finance. Engineering understands the architecture, pipelines, and platforms. Finance understands the business context and logic — how revenue is recognized, how allocations work, where the exceptions hide.
The handoff between those groups is often slow and messy. Analysts build fragile workarounds. Engineering teams inherit backlogs of finance requests that are actually business critical.
What resonates with me about Alteryx is that it sits in that gap. It enables finance and business analysts to build repeatable data workflows for extracting, cleaning, joining, enriching, and shaping data for specific finance use cases.
It emphasizes transparency and traceability, and it supports a model where IT can govern, and finance can execute. Just as importantly, it helps organizations turn their existing ERP, warehouse, and cloud investments into outputs that are actually usable for analytics, automation, and AI.
How to get started
If you want to make progress without boiling the ocean, my practical advice is simple: start small and start right.
Pick one workflow that is high pain and highly repeatable (recs, allocations, forecasting inputs, reporting packs).
Define what “trusted” means: the reconciliation rules, thresholds, approvals, and audit trail you need.
Build the AI-ready dataset first cleaned, joined, governed, and repeatable.
Then add AI where it makes sense (classification, summarization, exception explanation) inside the workflow, not as a free-floating tool.
My bottom line is this: AI-ready data is an operating standard. It is how we scale AI without scaling risk. And for CFOs, that should be the real objective, not chasing the latest tool, but building the trusted data foundation that makes smarter automation, better decisions, and more resilient finance performance possible.
Say your reconciliation tool flags a break between two ledgers, and now there’s a number that needs an explanation. The AI-generated summary says the mismatch is a timing difference, transaction posted late on one side. Reasonable. You move on.
Then your controller asks which transaction, on which date, and why it posted late instead of on time. And now you’re not looking at an explanation anymore. You’re looking at a sentence that sounded like one.
The four part t
Say your reconciliation tool flags a break between two ledgers, and now there’s a number that needs an explanation. The AI-generated summary says the mismatch is a timing difference, transaction posted late on one side. Reasonable. You move on.
Then your controller asks which transaction, on which date, and why it posted late instead of on time. And now you’re not looking at an explanation anymore. You’re looking at a sentence that sounded like one.
The four part test behind every AI answer
That gap is the same thing the last piece here named: can you explain where the answer came from, and would the explanation survive someone pulling on it? Most practitioners have been running that check for years, on spreadsheets, on junior staff’s work, on their own numbers before a review meeting. AI just hands you answers that sound complete far more often now, and faster than the checking can keep pace with.
The test itself breaks into a few plain questions, and it’s worth naming them because most people run all four without thinking about them separately:
Visible: Can you see where the number came from?
Understandable: Do you actually understand the logic that produced it, or just the sentence describing it?
Repeatable: Would the same input produce the same answer next time, or is this a one-off?
Auditable: Could someone other than you retrace it if they had to?
Four different failure modes, and an AI-generated explanation can fail any one of them while still reading like a good answer.
Why the gap is widening faster than the checking
The reconciliation example holds up because it’s ordinary. Nobody’s arguing AI shouldn’t touch reconciliation work. Matching balances, drafting a first-pass explanation for a variance, flagging what needs a human look — that’s real time back. The problem isn’t the AI doing that work. It’s that the logic behind “this is a timing difference” has to already be defined somewhere the AI can point to. If it isn’t, the model is pattern-matching its way to something plausible, and plausible is not the same as traceable.
Deloitte’s Finance Trends 2026 survey of over 1,300 finance leaders found 63% have fully deployed AI in their departments, with only 21% reporting clear, measurable ROI. That’s a broader adoption figure than an explanation-quality study, but the gap it points to lines up with the reconciliation example: plenty of AI running, not much of it yet standing up to scrutiny.
Where the logic has to live
Closing that gap starts with what the AI is drawing from in the first place, before it ever produces an answer. Every explanation an AI generates borrows its logic from somewhere: a threshold for what counts as material, a rule for what makes something a timing difference, an assumption about which system wins when two ledgers disagree. When that logic lives only as a pattern the model has inferred from past examples, the explanation is a guess dressed in confident language. When it’s defined, owned, and applied the same way every time, the AI has something real to summarize.
Finance has kept this kind of logic for as long as the job has existed, often in a spreadsheet somebody built years ago that everybody trusts without fully remembering why it works. That logic hasn’t changed. Who can now touch it, and how fast, has, and that means the definitions underneath it need to hold up to more traffic than they ever have before.
If you want to see what that looks like in a live workflow rather than in the abstract, Alteryx’s AI-Ready Starter Kits are pre-built Alteryx workflows and synthetic datasets designed to demonstrate how Alteryx can be applied to specific business use cases. They prepare and structure data to produce analysis-ready outputs, which can be extended using external AI tools.
The Reconciliation Exception Resolution AI-Ready Starter Kit shows the pattern from this piece in practice: exceptions routed to an owner, prioritized by materiality, and documented consistently enough that the resolution holds up when someone asks how you got there.
There’s a specific moment every finance leader knows. A number is about to leave the building — headed for the board deck, the earnings call, or the audit committee — and right before it goes, you pause. You want to know where it came from, and you want to know it will still make sense if someone asks how you got it. That pause happens no matter what produced the number.
That instinct shows up in the data too. In a Gartner survey of more than 200 CFOs, confidence acros
There’s a specific moment every finance leader knows. A number is about to leave the building — headed for the board deck, the earnings call, or the audit committee — and right before it goes, you pause. You want to know where it came from, and you want to know it will still make sense if someone asks how you got it. That pause happens no matter what produced the number.
That instinct shows up in the data too. In a Gartner survey of more than 200 CFOs, confidence across finance leaders’ top 2026 priorities averaged around 63%, while confidence in driving enterprise AI impact came in at just 36%.
Leaders aren’t lacking confidence broadly. They’re confident about cost discipline and growth investment. The drop is specific to AI. It’s a broader measure than any single number leaving the building, but it points in the same direction: AI is the one place finance leaders can’t yet count on the confidence that usually comes easily.
The four things every number has to pass
That pause is a fast version of a test. Before you’d trust a number, you check four things:
Where it came from
Whether you could explain it simply
Whether it would come out the same way twice
Whether you could trace it back through the data if someone asked
Most finance leaders have never written that test down. They’ve never had much reason to, because until now, the systems producing their numbers usually held up well enough that the check rarely turned into a real problem.
AI doesn’t automatically pass that test. It can produce a plausible answer to almost anything, including things it has no real basis for knowing, and the answer looks the same whether the logic underneath is solid or made up. That’s the real source of the confidence gap.
Leaders don’t doubt that AI can help. They doubt whether they could explain the answer if someone pushed back on it. The four things finance leaders already check for come down to four words: visible, understandable, repeatable, and auditable, or VURA. Those words succinctly describe what leaders were already checking for instinctually.
Who owns the logic underneath
Naming the test doesn’t resolve where it gets applied, though. That takes a harder answer about where the logic itself lives. Deterministic logic is defined by finance, not inferred by AI. A model can draft a variance commentary, summarize a forecast, or flag an anomaly worth a second look.
It should never be the one deciding what counts as an exception, how revenue gets recognized, or which threshold triggers an escalation. Those are calls finance makes, and AI’s job is to work within them, explain them, and apply them consistently, not to invent them when it doesn’t have enough to go on.
That distinction is where most AI disappointment in finance actually starts. The model usually isn’t failing at what it’s good at. The failure happens earlier: nobody defined the logic it needed, so it guessed, and it delivered that guess with exactly the same confidence it would use for a right answer. Looking at the output alone, you can’t tell the difference.
That’s exactly what the four-question test catches. Ask where the number came from, whether you can explain it, whether it repeats, and whether you can trace it back to the data, and you’ll find out fast whether the AI applied logic finance defined or made something up that looks close enough.
Building the standard into the workflow
This is an architecture decision as much as a governance one. The four questions get easy answers when there’s a layer between raw enterprise data and the AI consuming it, one that prepares the data, holds the logic finance owns, and keeps every output traceable back to both.
That’s the role Alteryx plays. It doesn’t compete with the model doing the reasoning, and it doesn’t replace the ERP or EPM system the data lives in. It’s the business logic layer that makes sure what reaches the model is something finance already stands behind, so the model’s output can be too.
Build that in, and the pause before the number goes out changes what it’s doing. Instead of hoping the number will hold up, you can check that it does, every time, because the answers to those four questions are already built into how the workflow works, not something you have to reconstruct from memory.
If you’re looking for a concrete way to see what that looks like on a real workflow, take a look at our AI-Ready Starter Kits: pre-built Alteryx workflows and synthetic datasets designed to demonstrate how Alteryx can be applied to specific business use cases. They prepare and structure data to produce analysis-ready outputs, which you can then extend using external AI tools such as large language models.
Understanding why finance leaders hesitate to trust AI is only the first step. The next is building the governed foundation that gives AI reliable business logic to work from.
Learn more in Building Finance AI You Can Trust, where you’ll explore the principles and practical steps behind AI-ready finance workflows.
If you work in finance, you’ve probably been handed an AI tool in the last year or so. Maybe a copilot in your spreadsheet, maybe something bolted onto the close, maybe a chatbot that promised to answer any question about the numbers. And maybe, if you’re being honest, it hasn’t changed your Tuesday very much.
You’re not doing it wrong. The tool isn’t broken. What’s missing is the part nobody put on the slide: AI is only as good as the work it’s standing on, and most o
If you work in finance, you’ve probably been handed an AI tool in the last year or so. Maybe a copilot in your spreadsheet, maybe something bolted onto the close, maybe a chatbot that promised to answer any question about the numbers. And maybe, if you’re being honest, it hasn’t changed your Tuesday very much.
You’re not doing it wrong. The tool isn’t broken. What’s missing is the part nobody put on the slide: AI is only as good as the work it’s standing on, and most of the time, the work underneath it is a mess.
Confident AI answers you can’t trust
Here’s a familiar scene. Someone asks the AI assistant a reasonable question — “why did margin move in the East region last month?” — and it produces an answer that sounds great. Confident. Well-organized. Possibly even formatted with little bullet points. The only problem is that you have no idea whether it’s right, because you don’t know which data it pulled, whether it used the current cost allocation method, or whether it quietly grabbed last fiscal year’s calendar.
So you do what any sensible finance person does. You check it by hand. Which means the AI didn’t save you the work. It added a step.
This is the quiet truth about why so much finance AI stalls. It’s not that the models can’t reason. It’s that they’re reasoning over data that was never cleaned, rules that were never written down, and logic that lives in one analyst’s head and three tabs of a workbook nobody else can open. AI didn’t create that gap. It just made it impossible to ignore, because now something is making decisions on top of it.
McKinsey looked at how finance teams are actually using gen AI and found a useful counterexample. Across the handful of finance functions where they saw AI adopted in earnest, professionals were spending 20 to 30% less time crunching data — and putting that time back into the analysis their job is supposed to be about. In one case, a global consumer goods company pointed a gen AI assistant at budget-variance work and saw roughly 30% of that manual effort disappear. That’s a real result. But notice what made it real: it was pointed at a specific, repeatable task, working from data the team had already organized around a shared definition of what “variance” even means. The AI didn’t figure that out on its own. The team handed it a problem that was ready to be automated.
What separates the workflows that pay off
The finance work where AI delivers tends to share a few traits. It’s bounded — a clear start and end, not “answer anything about the business.” It’s repeatable, the same shape every month. And it’s tied to something that matters: cash, margin, risk, a number someone downstream is going to act on.
That’s the easy part to say. The harder part is what has to be true underneath. For AI to work on one of those tasks, the data feeding it has to be prepared and validated before the model ever sees it. The rules — what counts, what gets excluded, how things roll up — have to be defined by your team and applied consistently, not guessed at by a model that’s never read your policy manual. And when the output lands, you have to be able to trace it back: which numbers, which logic, who signed off. In finance, that traceability isn’t a nice-to-have. It’s the difference between an answer you can put in front of an auditor and one you can only put in front of people who won’t ask hard questions.
Think about the difference between two versions of the same workflow. In one, the AI reaches into raw data, applies whatever it infers the rules to be, and gives you a number. In the other, the data gets cleaned and structured first, your team’s actual business logic gets applied to it, and only then does AI work on top of a foundation it can stand on. The first one feels faster right up until something’s wrong and you can’t tell why. The second one is the one you can defend in a meeting.
That’s really the test worth applying to any AI effort on your desk: can you explain where the answer came from, and would the explanation survive someone pulling on it? If yes, you’ve got something worth scaling. If no, more AI won’t fix it — it’ll just produce wrong answers more quickly.
Where this leaves you on Monday
None of this means starting over. The business logic your team has built — the spreadsheets, the rules, the institutional memory of how things actually work here — is the valuable part. The goal isn’t to throw it out for an AI that doesn’t know any of it. It’s to get that logic into a form that’s governed and repeatable, so AI can finally do something useful with it.
The most practical move is also the least dramatic. Pick one workflow. Not the whole close, not “AI across finance.” One bounded, repeatable, annoying task you’d happily never do by hand again — invoice matching, a recurring variance pull, a report you rebuild every month. Get the data right for that one thing, write the rules down, and put AI to work on top of it. When it works, you’ll have something real: a workflow you can trust, and a clear sense of what the second one should be.
Two traps worth naming, because they’re the ones McKinsey watched teams fall into. One is waiting for perfect data before you do anything — you’ll be waiting forever, and the team next door will have shipped three workflows by the time your data is pristine. The other is the opposite mistake: automating a process that’s still a tangle of exceptions and one-offs. Drop AI on top of a fragmented workflow and it doesn’t simplify it, it just adds a confident-sounding layer to the mess. The move is in between: standardize the one thing first, then automate it.
If you want a low-stakes way to see what that looks like before you commit, our AI-Ready Starter Kits are built for exactly this. AI-Ready Starter Kits are pre-built Alteryx workflows and synthetic datasets designed to demonstrate how Alteryx can be applied to specific business use cases. They prepare and structure data to produce analysis-ready outputs, which can be extended using external AI tools such as large language models (LLMs). They won’t run your finance function — that’s not what they’re for. But they make the shape of a workflow that actually pays off tangible enough to copy.
The AI on your desk isn’t the problem. The work underneath it is. Fix that for one thing, and you’ll stop wondering why AI hasn’t paid off — because it finally will.
This essay was written with Kasra Rafi, and originally appeared in The Guardian.
Earlier this month, about 40 top mathematicians gathered at OpenAI’s offices to discuss the future of their profession. The meeting was off-the-record, but if recent articles by mathematicians are any guide, it was mostly pretty glum. People fear for their jobs, their careers and the work they love.
We think the contrary view is more likely, at least in the short-term. AI models are nowhere near as capable as experi
This essay was written with Kasra Rafi, and originally appeared in The Guardian.
Earlier this month, about 40 top mathematicians gathered at OpenAI’s offices to discuss the future of their profession. The meeting was off-the-record, but if recentarticlesbymathematicians are any guide, it was mostly pretty glum. People fear for their jobs, their careers and the work they love.
We think the contraryview is more likely, at least in the short-term. AI models are nowhere near as capable as experienced academic mathematicians.
This isn’t to say that AIs aren’t producing stunning mathematical results at the level of PhD researchers. In mid-May, OpenAI announced that its frontier AI model disproved the unit distance conjecture, a famous 80-year-old problem in discrete geometry. In July, Anthropic’s published two AI-derived results in academic cryptanalysis. Earlier this month, OpenAI published 10 new mathematical results from its latest AI model. And Anthropic published Claude’s attempt to prove the century-and-a-half-old Riemann hypothesis.
These results are both a vivid demonstration of the amazing capabilities of frontier AI in 2026 and an illustration of their limitations. In general, these AI-powered advances in mathematics fall into one of two categories. Some are counterexamples to mathematical statements that people had been trying to prove. Others are novel applications of known techniques to existing problems that human experts either did not know or did not think of using.
The counterexample to the Jacobian conjecture is the most notable example of the first kind. Once it had been found, checking it was quick and straightforward. The difficult part was finding it among a large number of possibilities. The AI seems to have combined some sort of intuition acquired through machine learning with extensive computational search, in order to find the right example.
An example of the second kind is the unit-distance conjecture. It was motivated by an elegant construction, and most mathematicians expected it to be essentially optimal—so they generally tried to prove rather than disprove it. The counterexample brings in ideas from elsewhere in mathematics: algebraic number theory. If an expert with that background deliberately set out to find a counterexample, they would probably have succeeded. But there was no reason for someone with precisely that expertise to focus on this problem. Because of its scope, AIs don’t have those same limitations.
These results are relatively low-hanging fruit for AI; none of them required developing an extensive new theory. This does not make the discoveries trivial, or the AI’s achievements less impressive. Choosing the right direction, and recognizing an unexpected connection between subjects, are themselves forms of creativity. They are the same sorts of capabilities that led to AIs playing the game of Go at the grandmaster level, or doing Nobel-prize level chemistry in the area of protein folding.
What we have not yet seen is an AI developing a substantial new conceptual framework in order to solve a mathematical problem. Much of mathematics proceeds by identifying the objects that are truly central to a question and then developing a theory that helps us understand them. Current AIs are very strong at searching and recombining existing ideas, but they are weak at building any deep and sustained new theory.
This speaks to a more general limitation of current AI systems. They are creative in the sense that they can recombine existing ideas in novel ways. But they are not creative in others: they have not yet developed conceptually new theories or structures. And while they have larger working memories than humans do, know more about more different things than any particular human does, and can process information faster than humans, can, true novelty is still largely beyond their reach.
Of course, that distinction may not survive for very long. Predictions are notoriously hard, especially about the future of AI. None of these mathematical capabilities were explicitly designed for, or planned. They’re all emergent properties of increasingly capable AI models. We are both confident that someday we will see AI models that are capable of the type of creativity required to do novel mathematics. Will that be in a few months, a few years or a few decades? Of course we don’t know, but our guess is sooner rather than later.