Visualização de leitura

IT infrastructure shortages are real and lasting. Here’s how to cope

Lead times of nine to 12 or even 18 months. Costs rising by 35%, 45%, even 50% to 200%. More than halfway through 2026, the market for IT infrastructure that’s crucial for enterprise projects, including those involving artificial intelligence, is strapped.

Memory is at the root of the shortages. Memory prices “have risen by 50% to 200%, resulting in PC prices increasing by 35% to 45% and some server prices rising over 125%,” according to Jon Forest, VP analyst at Gartner. Network switches also need memory, albeit in lesser amounts than servers, so they are not immune, with prices and lead times likewise rising dramatically.

Industry experts agree that most of the issues stem from hyperscalers gobbling up memory capacity, which trickles down to servers, storage systems, and networking devices. But while the source of the problem may be new, supply chain disruptions are far from unprecedented.

As a result, industry insiders are not short on advice on how best to deal with the situation, with tips including making better use of what you have, considering options beyond your usual scope, and lots of planning with your vendors and internal finance teams.

State of the problem

Just how bad is the current supply chain problem? “It’s pretty bad,” says Matt Kimball, vice president and principal analyst with Moor Insights & Strategy. Companies accustomed to 30- to 45-day lead times for various infrastructure are now looking at 6, 12, or even 18 months.

“It’s real, and I’m hearing it from companies of all sizes, from the 1000-server to the 10,000-server shops,” Kimball says.

“Memory costs are expected to rise sharply well into 2027 and will reach up to 25% of network hardware expenses by the end of 2027,” according to an email Gartner’s Forest sent to Network World. The figure below shows the timeline Gartner expects for memory prices, and Forest notes that the same timing applies across networking, storage, and compute infrastructure. 

Gartner NAND DRAM stats

Gartner

“Enterprise network equipment pricing is projected to increase by over 20% in 2026. This upward trend is anticipated to continue with a further rise of 3% to 5% entering 2027, with no signs of price reduction until the end of 2027.”

But “reduction” will likely look more like “stabilization.”

“That’s something a lot of people don’t like to talk about. But let’s say prices went up 40%, they may come down five,” says Phillip Privett, senior vice president of vendor management with the global distributor and value-added reseller TD SYNNEX. “They’re not going to come down 40%.”

Perhaps worse, compared with past disruptions caused by issues such as fires in chip fabrication factories or the Covid pandemic, Kimball says this one is “durable” because its cause—the AI wave—is more long-lasting and just getting started.

“This AI inference wave we’re hitting is just beginning. It’s going to be longer and bigger than the training wave,” he says. “It’s impacting everything, from AI infrastructure to the traditional stuff that’s standing up your virtualization and cloud infrastructure.”

No vendors seem to be immune, not even the likes of Cisco, which makes its own Cisco Silicon One chips. Or, at least, it designs the chips; they’re actually manufactured by the Taiwan Semiconductor Manufacturing Company (TSMC), the same company that makes many of the other chips that are in such demand. And that’s only one component of many that comprise a switch.

On the other hand, the margins Cisco gets from enterprise sales are far greater than those from hyperscalers because Cisco sells mainly just hardware to hyperscalers, whereas enterprise sales generally include software and services as well. So, Cisco has incentive to keep enterprise customers happy and maintain the 66% margins it reported in Q3, its latest quarter.

Still, Cisco must deal with the same shortages as other vendors.

“I wouldn’t say any company is faring better than others,” says Neil Anderson, vice president and CTO for cloud, infrastructure, and AI solutions at World Wide Technology (WWT). “There may be nuances that some suppliers are employing to balance it to some extent, but I fail to recognize a supplier that’s not having almost the same issue.”

Cloud storage vendor Backblaze is one company that’s facing equipment cost and availability issues. “There are different types of shortages occurring in multiple places, all driven by unusual market demands, really by just a handful of very large buyers,” says James Rowell, senior vice president of operations with Backblaze.

Backblaze is constantly forecasting and monitoring demand triggers, Rowell says. That involves close alignment with the sales team to forecast client needs, as well as paying attention to historical trendlines to predict upcoming demand from new deals and growth with existing clients. But the company also looks for “unnatural market-related triggers” that would cause a spike in utilization.

With hyperscalers buying up vast amounts of capacity, “This is definitely an unnatural phase,” Rowell says. “For about for the last 12 months, I would say there’s been somewhere between a 15% and 30% uptick in costs,” especially in terms of servers and compute disks.

On the positive side, at least for Backblaze, the company is also seeing an uptick in business from an interesting source: AI companies. “We reported in the last earnings period a 70% increase in AI companies using our platform,” says Patrick Thomas, vice president of marketing at Backblaze. “That’s massive.”

On top of that, the company is seeing an uptick in deals from enterprises that can’t get the storage capacity they need or want on-prem. “There’s a general market nervousness where we’ve got potential deals coming our way because those organizations are concerned about being able to do it themselves,” Rowell says.

While some expect new chip fabrication plants currently under construction will ease memory supply constraints, Privett doesn’t buy it. “I don’t see it getting better anytime soon,” he says. “Building a new fab is a two-year process.”

Advice: Start with the basics

Enterprises, then, must play the cards they’re dealt. For Moore Insights’ Kimball, who did stints as an IT exec with the states of Florida and Oregon, that starts with making the most of what you have.

Such a strategy is “shockingly not implemented much” across the companies he sees. “A simple capacity planning exercise can free up a lot of resources.” That includes virtualized servers running at just 20% to 30% utilization as well as extending the life of existing servers. While 15 or 20 years ago it was common to refresh every four years or so, companies can often get six or seven years out of today’s servers.

While such strategies won’t solve your AI compute challenges, they can certainly help support your ongoing operations and free up budget for AI and other modernization projects, he says.

“Sweat your assets,” agrees Privett of TD SYNNEX. “Work them as much as you can, add only what you need, get extensions on your licensing, renewals on your services agreements and things like that. Just sweat it out a little longer.”

If you have budget to spend but can’t get the hardware you’re after, buy something else, says WWT’s Anderson. “Look at things that are not tied to those components, like software projects or SaaS licensing,” he says.

Get friendly with finance teams

Numerous experts recommend regular meetings with your CFO or finance teams to keep them apprised of what you’re up against so the company can plan accordingly.

Gartner’s Forest advises using rolling 12- to 24‑month forecasts and engaging early with suppliers to identify constrained components and SKUs. Committing to quarterly or monthly buys can help you avoid long-term agreements that extend past the rapid increases we’re seeing in 2026, he says.

Also engage with the financing arm of your equipment vendors, some of which are offering financing incentives, Privett says. Compute vendors in particular are offering subsidized financing, deferred payments, and low-cost financing for the first year or so. “Those are huge opportunities to take advantage of,” he says.

By engaging with finance teams, IT groups can conduct budget allocation exercises and try to come up with ways to make the financials work. The last thing you want to do is surprise them with additional budget requests out of the blue.

Kimball recalls his days with the state of Florida, when all budget requests were examined by a technical review working group—which was designed to be hostile.

“I can’t imagine going to them and saying, ‘Oh, did I say that was a million dollars? It’s actually $2 million. I need you to write me a bigger check,’” he says. “I would walk into one of the swamps in Tallahassee and get eaten by the alligators instead of doing that.”

Work with your vendors and VARs

As you put plans together, lean on your vendors for help, including channel partners such as value-added resellers (VAR) and national resellers. “Work with them to map things out and understand what your workloads will look like,” Kimball says.

That’s what Backblaze’s Rowell regularly does with his suppliers. He lays out his forecast for the year, with commitments on what Backblaze will definitely buy, as well as scenarios that account for rapid growth, say, 2x. “And they’ll come back with, ‘Well, okay, no problem,’ or maybe they say we need to put in an allocation right away, or we won’t be able to get what we may need,” he says.

Similarly, he sits down with his CFO regularly to map out predictive models that factor in inflation, price hikes, and the like. The idea is to plan out multiple scenarios, so you don’t get blindsided.

“If you don’t do that, you’ll get caught with your pants down, on the upside-down end of spectrum,” he said – meaning not having the capacity to take advantage of market opportunities.

Acquiring the capacity you need to meet project demand may also mean being flexible in terms of your equipment choices. If you’re a Dell shop but can’t get Dell servers, maybe you go with Lenovo, Kimball says.

“You’ve got to figure out how to use all this silicon and infrastructure in a heterogenous way to serve your needs,” he says. That’s especially true when it comes to AI infrastructure. “If you think you’re going to go with 100% Nvidia for everything from RAG [retrieval augmented generation] to inferencing at the edge, you’re kind of crazy, not because of cost but because of availability.”

Look at alternatives, including AMD and cloud solutions, while staying mindful of how it all plays together. You may not be able to get Nvidia GPUs, but AWS, Azure, and Oracle Cloud have them, Kimball notes.

Be strategic, perhaps by using cloud offerings to handle certain tuning or inference workloads, then bringing them back in-house when appropriate. “Have a better understanding of what absolutely has to be on prem and what can be in the cloud,” he says.

That’s good advice, says Backblaze’s Thomas. When it comes to AI, think about performance tiers and the range of use cases you have. They don’t all need top-tier performance.

“People get wrapped around axle of needing the top end. There’s a lot of flexibility in the edges, innovation in different hardware and software,” Thomas says.

Gartner likewise advises companies to increase configuration flexibility and expand sourcing paths. That may include buying from secondary markets and lease-return programs to preserve continuity with existing infrastructure until the shortages pass, Forest says.

Get started somewhere

Even if you can’t acquire or have to wait for the infrastructure you need, don’t let that keep you from getting started with AI or other modernization projects.

Options include public cloud and neocloud providers, Anderson says. WWT also provides capacity in its own lab so customers can get started with proof-of-concept projects. “Don’t just throw your hands up. We can help you find access to capacity,” Anderson says. “Production-scale AI may be delayed, but don’t let that derail your strategy.”

Colocation providers may likewise be an option, especially if enterprises are struggling to acquire high-end networking equipment. Networking is a key value proposition for colocation providers, in that they have built-in connections to various cloud providers and other ecosystem players.

Equinix, for example, has 280 data centers in 77 metropolitan areas, says Phil Read, senior director, colocation product management for the company. If you have the compute infrastructure, Equinix can help you with the high-end connectivity required both intra- data center and at edge facilities.

It also has partnerships with the likes of Cisco and Nvidia for “ready-to-go AI connectivity,” Read says. That means Equinix offers the right infrastructure to meet the requirements of high-end compute solutions in terms of power density and cooling. Such power densities are significant, requiring 120k VA per rack and up. “There’s plenty of talk about a megawatt rack,” he says.

Power is a significant issue in this entire discussion, Privett says. Older installed computing infrastructure likely consumes far more power than newer systems, which is an argument for upgrading as soon as possible.

“If you modernize today, you could substantially reduce the number of servers needed to support the same applications at a much lower power consumption rate,” Privett says. He advises sitting down with folks from the OT side of the house to make sure power is available for whatever you want to do. In many areas, power is at a premium.

If your plans include installing GPU environments in your own data center, WWT advises you not to delay. “We’re telling customers, you need to talk with us and get that designed, get that ordered, because it will take quite a bit of time until it actually ships and we’re able to install it,” Anderson says.

Moor Insights’ Kimball agrees. “You have to order these parts today if you want to see them hitting your dock, your warehouse, or your office 12 months from now.”

How cost visibility becomes a competitive advantage in FinOps in 2026

As spending on cloud technologies grows, so does waste. The Flexera 2026 State of the Cloud Report found that 27% of organizations expect to spend more on cloud this year, with 17% already exceeding their budgets over the previous 12 months. The estimated share of wasted cloud spend has already crept up to 29%, undoing several years of progress, undoing several years of progress.

Companies are rapidly investing in cloud technology, but often understand less about how to use it fully and efficiently. That is not a coincidence, and it is exactly the gap FinOps is meant to close. It is also why the practice is moving out of the finance department and into the strategy conversation.

What is FinOps?

FinOps is a blended operational framework that maximizes technology value by uniting engineering (DevOps), finance, and business teams. It involves close collaboration to break down silos between tech and finance, with shared ownership of cloud spend across engineering, finance, and business teams.

What distinguishes FinOps is real-time visibility into what is being spent and why. It also treats optimization as continuous work rather than an annual cleanup exercise.

FinOps is important because cloud spending isn’t like a typical budget line. It’s more variable and usage-based, so relying on an annual review doesn’t work. Engineers can quickly create infrastructure, scale it, and tear it back down in a day, making forecasting more challenging than in the past.

What FinOps does is change who sees what. Engineers have more insight into the actual cost of a build. Finance gets numbers it can trust. Business leaders can tie spending directly to business outcomes. It removes much of the guesswork and turns cost data into a shared language rather than a monthly surprise.

As Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise, puts it, “For a long time, FinOps meant only cutting the bill: find the unused stuff, resize a few instances, and report the savings. That still matters, but it’s not what separates companies today. The ones pulling ahead are using FinOps to make faster, smarter calls about where their tech spend actually pays off. That’s a different job, and it shows up directly in how fast a company can move.”

How FinOps spending has changed

Only a few years ago, FinOps was mostly focused on cloud infrastructure spending. Today, that focus increasingly includes AI-specific investment. The FinOps Foundation 2026 State of FinOps Report found that 98% of organizations now manage AI spend specifically. FinOps has also expanded well beyond cloud infrastructure. It’s more common now to see FinOps coverage extend to licensing (64%), private cloud (57%), and data centers (48%). Around 90% also manage SaaS spend or plan to do so within the next year.

It’s also worth noting that the same FinOps Foundation report found that 78% of teams report directly to the CTO or CIO rather than operating solely within finance departments. To us, that reporting line says a lot. It suggests that companies increasingly see technology spending as a strategic lever rather than simply a line item to reconcile.

How AI and cloud spending are moving in the same direction

The trend toward bringing AI and a broader range of technology spending into FinOps is backed up by a Gartner report, which estimates that global IT spending will hit $6.31 trillion by the end of 2026. That’s up 13.5% from the previous year. Data center systems spending is expected to grow 55.8%, with generative AI model spending more than doubling over the same timeframe. Gartner, in a separate forecast, expects public cloud services to grow by 21.3% in 2026, with the market reaching $1.48 trillion in value by the end of 2029.

We see these figures as two sides of the same shift. AI workloads are also usage-based, which makes them more unpredictable, partly because some teams haven’t had to consider unit economics before. A fine-tuning run or a forgotten inference endpoint can quickly become one of the biggest items on a cloud bill. Most teams don’t have the tagging, forecasting, or accountability needed to catch those costs before they get out of control.

“AI spend just behaves differently from a normal application workload. It spikes, it’s hard to pin on one team or feature, and you often don’t know the real cost per outcome until the invoice lands. Companies that already had solid FinOps habits before AI adoption took off are adjusting faster because visibility and ownership were already part of how they worked. Companies that treated FinOps as an annual cleanup are the ones getting caught out.”

Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise

Why visibility and shared spend ownership create an advantage

Flexera numbers on discount usage point to the same issue: fewer than 50% of organizations are using the most basic cost optimization tools. The adoption of tools like reserved instances or savings plans is slow, with only 48% of companies using Google Committed Use Discounts and 45% using AWS Reserved Instances. Too many others are leaving low-risk savings on the table.

In many cases, the real problem is a lack of ownership and visibility. If no team owns the cost of a workload, no one has enough reason or enough information to choose the right pricing model. That is where the competitive gap starts to open: some companies can explain and act on their spend quickly, while others cannot.

“A mistake we still see a lot is trying to optimize the bill instead of the system behind it. Deleting unused resources saves money once. Redesigning how workloads scale, how environments get spun up, and who’s on the hook for what keeps costs under control for good. That’s the difference that turns into a real competitive edge later.”

Siarhei Sukhadolski, Chief Delivery Officer & Head of Competence Center at Innowise

What visibility and shared ownership look like

FinOps operates well when at least three structures are in place.

  • Every workload or inference endpoint has a clear owner tied to its cost.
  • Cost and usage data are shared and available to everyone who needs them before they have to ask.
  • There is an ongoing review cadence designed around continuous optimization.

Teams that jump into dashboards before assigning ownership and establishing the data flow often end up with visibility but no accountability. Teams that start with ownership, even with basic tooling, get a different result. They tend to see savings stick instead of resetting every few months. That proper order is the single biggest predictor we have seen across cloud and AI cost engagements.

What business leaders need to know

There is a simple test a CEO or CFO should apply to FinOps. It’s whether the company can clearly state what a workload costs and whether it is worth that cost right now. Does it take finance three weeks to answer? Can the company provide an answer in real time? Are there live numbers and clear ownership behind every workload? That is what will enable leaders to make confident calls on where to invest and where to pull back.

FinOps is becoming a proxy for how well a company manages technology. With shared ownership comes high visibility. Add in continuous optimization and companies gain an advantage. These are not just cost-saving tactics. They represent operational discipline, separating those who can move quickly on AI from those who spend heavily only to find themselves still behind the rest of the pack.

Economic process modeling: Business cases beyond cost accounting

Cortney Pagel has learned to expect a particular question whenever she proposes changing how work gets done. Pagel, a senior business analyst and change manager at ENGIE Impact, often hears it from the digital side of the organization before she has finished explaining the change: How much money are we looking to save here? “I don’t always have an answer,” she says. “And so that is very frustrating for me.”

The frustration comes from a familiar structural problem. Most financial systems connect money to departments, accounts, cost centers and, in more mature implementations, activities; however, they rarely connect economics to the anatomy of the work itself. Pagel describes processes that are still too manual to be tracked cleanly, while in other cases the financial detail exists but was never tied to the process model. For a period, she resorted to putting cost estimates in comment bubbles on process diagrams because there was no systematic place for them. “There’s been no great system or method to do it,” she says. “It’s definitely not just me.”

The question she keeps being asked therefore exposes a broader lacuna in management accounting. Finance can usually explain what a process consumes, and it can often estimate what a proposed change might save. Yet it is far less equipped to show which process components contribute value, which destroy it, which absorb risk and which create information or options whose economic effects may surface much later. Consequently, a transformation business case can become highly precise about one side of the equation while leaving the other largely narrative.

A ledger on the ledge of usefulness

The general ledger reports what a department or cost center consumes. Activity-based costing (ABC), where organizations have implemented it with sufficient discipline, pushes that resolution further by assigning costs to activities. Both approaches remain useful; nonetheless, their analytical center of gravity is consumption rather than contribution. They can tell management where resources were spent with increasing granularity, while offering much less visibility into what an individual activity economically produced.

Double-entry accounting, dating back 500 years to Venetian merchants, earns its reputation for symmetry, although the symmetry belongs primarily to bookkeeping. The two sides of an entry describe the same financial event, while revenue generally appears when a transaction is recognized rather than carrying a lineage back through the many process components that helped create it. A renewal, expansion, avoided loss, faster decision or improved customer relationship may depend on dozens of steps, yet the contribution of any one step rarely has an account to which it can be posted.

This creates an analytical asymmetry that can influence investment decisions more than finance leaders may realize. When a CFO or operating executive evaluates a proposed process change, the cost side often arrives quantified while the value side arrives as prose, judgment or a collection of indirect metrics. The quantified side therefore tends to carry disproportionate weight because it is already denominated in the unit in which the decision is made: money. Indeed, acknowledged uncertainty may be safer than one-sided precision, because the latter can carry the authority of a number while obscuring what the model omitted.

One process, many economic artifacts

Consider the process of customer onboarding. Operationally, it is a sequence of tasks needed to establish a customer, configure services, obtain approvals and move the relationship into a steady state. Economically, however, those same steps may establish relationship patterns that influence retention, create the account depth that enables a later cross-sell, generate behavioral and preference data whose usefulness compounds over the customer lifecycle, and reduce churn risk through investments made well before the customer has a reason to leave.

Embedded in that same process may be approval controls whose original compliance rationale has waned, manual handoffs between systems that were never integrated, duplicate checks and wait times that gradually erode the loyalty the process was intended to build. Some components may therefore create value; others may protect it and still others may quietly consume it. Yet a conventional cost model compresses this heterogeneous mesh (or mess!) of economic activity into a single process cost, which is useful but incomplete.

Improving or transforming the process requires a more discriminating account of what each component is doing economically: which steps build value, which erode it, which create unnecessary friction, which absorb risk, which generate ancillary benefits and which perform economic work that becomes visible only after the step is removed. Without that component-level view, an efficiency initiative can readily eliminate something valuable simply because its cost was easier to suss than its contribution.

Economoic process model: sample customer onboarding.

Sample customer onboarding — economic process model.

LINQ.it

Putting economics on the process map

Economic process modeling (EPM) supplies that missing layer. As I described in a recent column, Business Transformation Needs a True Economic Approach Rather Than Guesswork, EPM decomposes a process into its constituent components—the information flows, human decisions, system actions and organizational touchpoints that make up the actual work—and then attributes economic effects to each across five dimensions: revenue contribution, cost and friction, risk exposure, option value and information value.

The component level matters because the economically significant finding often sits buried within the process as a whole. Two steps that look roughly equivalent on a process diagram may carry very different economic profiles once attribution is applied. A seemingly minor validation step, for example, may generate information that reduces downstream risk, while a more conspicuous approval step may be largely vestigial. A cost-only review can easily misread the two because it sees effort more readily than consequence.

Pagel describes the capability she wants in similarly practical terms: the ability to see processes at an organizational level, understand what they cost in aggregate and then break those economics down step by step. That level of resolution helps because process transformation decisions are rarely made at the level of an abstract end-to-end flow; they are made by automating, eliminating, combining, outsourcing or redesigning individual components. Consequently, finance needs an economic view at the same level where the design decision is actually being made.

The oft-ignored value of information itself

Information value is where this analytical oversight may be most acute, particularly because most processes today generate data as a byproduct of execution. For example, a credit review produces repayment-behavior signals, a claims intake creates fraud indicators, and a procurement approval accumulates supplier-performance evidence. Those outputs may have future utility well beyond the transaction or process that generated them, even though conventional cost accounting typically has no place to represent them.

Most organizations, however, still treat much of this data primarily as documentation, exhaust or a compliance burden rather than as a potentially monetizable asset. A process redesign can therefore appear efficient while externalizing, degrading or destroying information whose economic contribution was absent from the business case. Infonomics, the discipline of treating information as an economic asset with attributable value, provides the grounding for this dimension of EPM and helps expose value that can otherwise disappear during an ostensibly sensible transformation.

From cost review to capital allocation

EPM extends cost accounting by adding an economic perspective the ledger wasn’t designed for. Sure, cost remains an indispensable computation. However, the business case becomes materially more complete when the components proposed for automation or elimination are also evaluated for revenue contribution, risk absorption, optionality, information yield and the friction they create or remove.

This can change the quality of the capital-allocation discussion. A step that costs $500,000 annually certainly may be a strong automation candidate, yet the savings figure is incomplete if the same step prevents $2 million in avoidable losses, preserves a customer relationship, generates valuable information or creates an option the business may need later. Conversely, a relatively inexpensive step can still be economically destructive if it adds delay, rework or customer attrition. The point is not to manufacture spurious precision around every benefit; rather, it is to make the relevant sources of value visible, estimable and subject to the same scrutiny as cost.

Which brings us back to Pagel and the question she hears whenever she proposes a change: “How much money are we looking to save here?” Savings are only one side of the economic case. The more consequential question may be what each affected component contributes today, what value may disappear if it is changed, and what new value the redesigned process could create.

Indeed, a transformation can look compelling when the savings are visible and the value at risk remains invisible. Economic process modeling gives finance a way to juxtapose both in the analysis, so that a proposed change can be judged not merely by what the organization expects to spend less, but by what the work itself is actually worth.

Why we need technology economists

The advent of AI is precisely why organizations need technology economists, not just IT finance professionals.

IT finance is primarily concerned with budgeting, accounting, cost allocation, depreciation, chargebacks and financial reporting. These disciplines remain important, but they assume a relatively stable relationship between technology spending and business outcomes. AI breaks that assumption.

Technology economics asks a fundamentally different question: How do technology investments create, destroy, shift or delay economic value?

AI introduces a set of economic dynamics that traditional IT finance was never designed to evaluate.

AI creates non-linear economics

In traditional IT, spending $10 million typically produced a somewhat predictable capacity increase or operational improvement.

With AI, a $10 million investment might generate $100 million in value. It might generate no value at all. It could increase costs while appearing successful. It could also create strategic advantages that do not show up in financial statements for years.

A technology economist studies the relationship between technology inputs, organizational capability, productivity outcomes and economic value creation.

IT finance largely records the spending.

AI changes the economics of labor

AI is not merely another technology platform. It acts as a form of digital labor.

Organizations now face questions such as:

  • Should work be done by humans, AI, automation or a combination?
  • What is the marginal cost of an AI-generated transaction versus a human-generated one?
  • How does AI affect productivity elasticity?
  • When does AI create labor substitution versus labor augmentation?

These are economic questions, not accounting questions.

AI simultaneously creates technology inflation and deflation

A fascinating paradox is emerging: AI can reduce costs in some areas while dramatically increasing costs elsewhere.

For example, fewer coding hours. More GPU costs. Lower service desk costs. Higher cybersecurity costs. Reduced consulting expenses. Increased data management expenses.

Technology economists study entire economic systems and value chains.

IT finance often sees only line items.

AI requires measuring economic outcomes, not technology outputs

Historically, organizations measured projects delivered, systems implemented, budgets achieved and uptime percentages.

The AI era requires measuring:

  • Revenue generated
  • Margin improvement
  • Risk reduction
  • Productivity gains
  • Decision quality improvement
  • Time-to-market acceleration
  • Innovation capacity

Technology economists focus on these outcome measures.

This is one reason why AI performance measurement frameworks, including AI-focused balanced scorecard approaches, are becoming increasingly important.

AI introduces massive opportunity costs

One of the largest AI risks is not technological failure.

It is investing in the wrong AI initiatives.

A bank might spend $50 million building an AI solution that saves $5 million annually while ignoring another opportunity that could have generated $500 million in new revenue.

Technology economics focuses on capital allocation efficiency, opportunity cost, marginal returns and portfolio optimization.

These concepts sit outside traditional IT finance.

AI makes technology a strategic production function

Historically, technology supported the business.

Increasingly, technology is the business.

In many industries, AI determines customer experience, operating efficiency, innovation speed and competitive advantage.

Technology is becoming a primary production factor alongside labor, capital and natural resources.

Organizations therefore need experts who understand the economics of technology as a production asset.

AI creates new forms of technical and economic debt

Many organizations are deploying AI rapidly without understanding:

  • Long-term infrastructure costs
  • Model maintenance costs
  • Data quality costs
  • Governance costs
  • Security costs
  • Regulatory costs

A technology economist examines the total lifecycle economics.

The cheapest AI solution today may become the most expensive solution over the next decade.

Why this matters

The central challenge of the AI era is no longer “Can we build it?”

The challenge is, “Should we build it, where should we deploy it, what value will it create, what risks will it introduce and what is the optimal economic allocation of technology capital?”

Those are technology economics questions.

IT finance professionals are essential for controlling and reporting technology spending.

Technology economists are essential for determining whether that spending creates sustainable economic value.

As AI becomes embedded into every business process, the organizations that outperform will not necessarily be those with the biggest AI budgets. They will be those that best understand the economics of technology itself — how AI, data, infrastructure, labor, risk and innovation combine to create measurable business value. That is the domain of technology economics.

Why a cheaper model won’t lower your AI bill

Somebody in finance has already forwarded you the pricing comparison. A downloadable frontier-class model at a fraction of what you’re paying now, with the obvious question attached: why are we still on flagship rates? The honest answer comes in two halves. Open weights are the best thing to happen to enterprise AI buyers since the category existed, and switching to one still won’t lower your bill this year. Both of those are true, and the gap between them is where the useful work sits.

How we got here

In mid-July, a 2.8-trillion-parameter open-weight model shipped with performance close to the commercial frontier, and the full weights followed ten days later under a custom license. Markets moved before Washington did. Semiconductors were hit hardest that session and one widely held chip ETF finished the week almost 9% lower (CNBC). Washington began weighing restrictions on open-weight models soon after, and the industry answered inside a fortnight.

On July 24, twenty-five companies published a letter asking policymakers to leave downloadable model weights alone (Tom’s Hardware). Not a single founding signatory sold access to a closed-frontier model. Three major labs were absent at launch, two signed within 72 hours (TheNextWeb) and the roster passed 270 organizations inside ten days (Forbes). The sole holdout published its position days later, agreeing with much of the letter while disputing two safety claims.

What open weights hand you

Start with what genuinely changed, because it’s larger than the coverage suggests. A downloadable model at frontier-class capability puts a permanent public floor under what that capability can be sold for. But no supplier prices at whatever the market will bear once a comparable input is obtainable elsewhere, and that shift doesn’t reverse.

Stakeholders already know how to think about this.  They just haven’t been filing AI under the right heading, which is concentration risk. A single provider holding a load-bearing production input, controlling both pricing and release schedule, would sit on the risk register in any other procurement category. The only reason AI was able to bypass this was that there was no alternative worth naming. Now there is one.  The leverage shows up at renewal whether or not you ever deploy an open model, since the negotiating position changes the moment the alternative becomes credible.

It changes what you can responsibly commit to, as well. Until now, a multi-year AI investment has meant betting the program on one supplier’s pricing decisions and deprecation schedule, and that’s a hard paper to take into an investment committee. Commitments get easier when the input underneath them has a substitute. Workloads governed by data residency rules come back into scope too, and for some companies that means markets they’d written off.

Investors’ point of view is a little different in this scenario, and probably more accurate. Valuations built on sustained pricing power at the model layer assume something the capability data no longer supports. As models converge, the primary durable margin moves toward distribution, proprietary data and internal workflows that the customers cannot rip out. This happens to be the ground that the coalition’s founding signatories already hold.

None of that requires a single enterprise to switch models. So, the case against restrictions is a real one, whatever mix of principle and self-interest sits behind it. And note one of its own asks: public funding for shared evaluation frameworks, which the signatories evidently agree don’t exist yet.

The monopoly is breaking, just not on the scoreboard everyone watches

Stanford’s 2026 AI Index puts the leading closed model ahead of the leading open model by 3.3% as of March 2026, having been 0.5% ahead in August 2024 (Stanford HAI). The same chapter records six labs clustered inside 25 Arena Elo points at the top and reads that convergence as pushing competition toward cost and reliability. For most enterprise work, a 3.3% capability gap is not a reason to pay a multiple.

Market share tells a different story. Menlo Ventures, surveying 495 US enterprise AI decision-makers, puts three vendors at 88% of the enterprise LLM API market between them, on 40%, 27% and 21% (Menlo Ventures). The same research found enterprises tend to stay with whichever vendor they picked, upgrading within that provider even where switching costs are low.

Both are true and reconciling them is the point. Suppliers price differently when they know you can leave, and that holds whether or not you ever. The alternative never has to be used to change what you pay. So, the pricing monopoly is gone while market share sits exactly where it was. Pricing power was the monopoly that mattered to buyers, and open weights broke it.

The headline price is not the cost

This part is arithmetic. On published rates one recent open model looks roughly a third the price of a leading commercial system. Cost per completed task tells a different story, and the firm that measures it states the mechanism plainly: because cost tracks real token usage, models producing longer answers or more reasoning bill more per task even at identical per-token prices (Artificial Analysis). One current frontier model cost about twice its predecessor per task on that measure, driven entirely by token consumption and not by any price change.

It’s not possible to move a production workload to a budget-friendly mode without having a per-tasking definition of good enough, and literally no one has that handy. Ask what accuracy a workflow requires and you’ll hear crickets, or a number invented on the spot. Ask what it currently achieves and you’ll get the same silence. Same shrug, different meetings. Until both questions have answers, a pricing table is just somebody else’s workload dressed up as your business case.

I work on AI infrastructure, and the primary hurdle that I keep hitting isn’t a technical one. Writing down what good enough means requires somebody to put their name on a number they’ll be held to later. That’s an organizational decision, and not an engineering one, which is exactly why these documents don’t exist in most companies. Teams spend multiple quarters comparing models but barely spend a week agreeing what exactly they’re comparing them for. The related thing I’d say from that seat is that most groups believe they evaluated a model when what they did was try it. Someone ran twenty prompts, liked what came back, and the decision got made in the room. That’s a demo. Demos flatter every model about equally, which is why they can’t tell you whether the cheap one is costing you anything.

Self-hosting won’t rescue the math for most buyers either. A frontier-scale checkpoint runs well past a terabyte, so outside the regulated cases above, the win shows up as hosted providers competing for your workload.

What will actually move your bill

Two things will, and neither of them is a model release. First is the eval infrastructure, since it turns any price difference into a decision that you can defend. Scrape a few hundred real queries from the peak-load and freeze them as your golden test set. Get the people who own the business outcome to write down what a good response looks like, in specifics instead of adjectives. Test the existing model first and generate the scores. Most teams underestimate this step, and it’s the one that makes every future comparison possible.

The second is your own compliance position, which the August decision did nothing to simplify. Every proposal in circulation points at documentation and audit trails, and the holdout lab’s own position points the same direction from the opposite side of the debate. What reaches the buyer either way is a demand for evidence about what your systems can do and what they did. Most 2027 budgets don’t carry that line.

And on the stock market question

Expect volatility on release days and don’t mistake it for repricing. A capable open model lands and chip stocks sell off within hours, one such session costing a single chipmaker close to $600 billion (CNBC). They recover over the following weeks.

My read is that markets keep filing these as demand shocks when they are supply-side price events. Cheaper capability has driven adoption and compute consumption up together every time, which is the opposite of what a selloff assumes. So, keep the two conversations apart. A chip selloff tells you about supplier margins and nothing about your own AI spend, and boards that conflate them freeze budgets during a dip or wave them through during a rally. What would genuinely reprice this sector is a regulatory outcome raising the cost of shipping capability, or an adoption curve that flattens.

Scaling enterprise AI without breaking the bank: A CIO’s guide to AI unit economics

Uber’s experience highlights a new enterprise AI challenge: adoption can scale faster than an organization’s ability to measure economic value. As companies move from AI pilots to widespread deployment, the question is no longer whether employees will use AI — it is whether every AI investment can justify its cost.

Generative AI is changing the economics of enterprise technology. Every inference request, AI agent execution and model interaction can create recurring costs, while cloud infrastructure, GPUs, data, security, integration and governance add to the total cost of delivering AI. The economics that made an AI pilot look compelling can look very different at enterprise scale.

The next phase of enterprise AI will not be defined by the number of models deployed or pilots launched. It will be defined by sustainable business value. For CIOs, CFOs and business leaders, success depends on maximizing business outcomes while controlling the cost of delivering AI.

AI success is an economics problem, not just a technology problem.

AI unit economics: The new measure of AI success

Manufacturers measure cost per unit produced. Banks track cost per transaction. Enterprise AI requires a similar discipline — not measuring how many models are deployed, but how much business value is generated for every dollar invested.

Traditional software investments typically involve predictable costs. AI introduces a dynamic cost structure where every interaction creates ongoing expenses, including compute, inference, storage, data retrieval, monitoring, integration and governance.

A simple framework for evaluating AI investments is:

AI unit economics = (business impact × adoption × reusability) ÷ total cost of delivering AI

Consider an illustrative AI-enabled invoice-processing workflow. If AI reduces processing time, increases straight-through processing and the same capability can be reused across accounts payable, procurement and supplier onboarding, its economics improve not simply because the model is cheaper — but because the value and reuse increase faster than the cost.

This equation reflects a simple principle: AI investments create the most value when they solve high-impact problems, achieve broad adoption and create reusable capabilities while keeping operating costs under control.

Business value may include productivity improvements, faster decisions, improved customer experiences, revenue growth, cost reduction or reduced operational risk.

The objective is not to minimize AI spending. It is to maximize the value generated from every AI dollar.

Understanding the cost drivers and metrics that matter

AI unit economics depends on understanding both consumption drivers and business value drivers. Infrastructure, GPU compute, inference usage, data management, security, compliance and governance all contribute to AI costs.

CIOs should move beyond tracking total AI spend and monitor metrics such as cost per inference, token consumption, GPU utilization, model usage, latency, adoption rates, productivity improvements, automation levels and business impact.

CIOs should treat AI consumption as a portfolio allocation problem — not simply an infrastructure problem.

The winners will not be the organizations that deploy the most AI. They will be the organizations that know where every AI dollar creates measurable business value.

Optimizing AI unit economics: Practical strategies for CIOs

Improving AI unit economics requires more than reducing costs. It demands thoughtful architectural and operational decisions that maximize business value while minimizing unnecessary AI expenditure. The following strategies can help CIOs achieve that balance.

1. Use the right technology for the right problem

Not every business problem requires a large language model. Many structured prediction challenges — such as demand forecasting, fraud detection, predictive maintenance, churn prediction and pricing optimization — are often better solved using traditional predictive machine learning models.

These models typically require fewer computational resources and can deliver comparable or superior performance for well-defined prediction problems.

Large language models create the greatest value for language-intensive tasks such as enterprise search, document analysis, conversational assistants, software development and content generation.

The right question is not, “Where can we use generative AI?” It is, “What is the simplest technology capable of delivering the required business outcome?”

2. Manage AI as a portfolio, not a collection of projects

Many enterprises still evaluate AI initiatives individually. Leading organizations manage AI as a strategic portfolio.

Every AI investment should have clear business objectives, success metrics, ownership and exit criteria. Experiments should either demonstrate measurable value and scale or be discontinued.

A portfolio approach helps eliminate duplicate investments, increase reuse of AI capabilities and shift funding toward initiatives with the strongest business impact.

3. Optimize AI architecture and model selection

AI infrastructure decisions are now financial decisions. Unlike traditional applications, AI workloads create continuous demand for compute resources, making inference costs a major operational expense as adoption grows.

Organizations are increasingly adopting hybrid AI architectures that combine public cloud flexibility with private infrastructure for high-volume, sensitive or regulated workloads. This approach can improve resource utilization, reduce data movement costs, strengthen data sovereignty and create more predictable operating expenses.

However, infrastructure optimization alone is not enough. Enterprises must also ensure that each workload runs on the right model. Not every interaction requires the most advanced — and most expensive — foundation model.

CIOs should adopt intelligent model routing strategies that match workloads with the right models based on complexity, performance and cost. Smaller language models, open-source models and domain-specific models can handle routine tasks such as classification, extraction and summarization at significantly lower cost.

Premium foundation models should be reserved for complex reasoning, advanced analysis and high-value decision support where their additional capabilities justify the expense.

The goal is not to maximize model size or infrastructure investment — it is to optimize AI consumption for measurable business outcomes.

4. Redesign business processes — Don’t just add AI

Adding AI to inefficient processes rarely creates transformational value. The biggest improvements come from redesigning workflows around AI capabilities.

For example:

Traditional workflow:
Employee → AI Assistant → Invoice

AI-enabled workflow:
Invoice → AI Agent → Human Exception Review

In this model, AI handles routine tasks while employees focus on complex decisions.

As organizations transition from basic copilot tools to autonomous agentic AI architectures capable of independent execution, the greatest value will come from designing workflows where AI agents handle multi-step operational tasks while humans focus on exception handling, complex judgment and strategic goals.

5. Measure outcomes and strengthen AI foundations

Providing employees with AI licenses does not automatically create productivity gains. Without clear use cases, adoption strategies and outcome measurement, organizations can increase AI spending without achieving proportional business value.

Leading enterprises focus on value realization by measuring outcomes such as hours saved, productivity improvements, automation rates, customer experience improvements, revenue impact and cost reductions.

However, productivity measurement alone is insufficient. Sustainable AI economics also depends on the foundations that make AI reliable, scalable and trusted. High-quality data and strong governance act as value multipliers by reducing errors, improving adoption and enabling responsible scaling.

Weak foundations can quickly erode AI economics. Poor data increases operational costs by creating inaccurate outputs, more human review, lower employee trust and repeated model execution.

Similarly, governance should not be viewed only as a compliance requirement. As IT leaders navigate the operational costs and requirements of AI governance, strong responsible AI practices — including security controls, explainability, regulatory oversight and human oversight — reduce operational risk while increasing confidence in AI-driven decisions.

Clean data improves model performance, while effective governance ensures AI systems are reliable, secure and scalable. Together, they improve AI unit economics by reducing waste, increasing adoption and maximizing the business value generated from every AI investment.

Measuring AI economics is only useful if organizations build the operating discipline to manage it continuously.

6. AI FinOps: Operationalizing AI unit economics

Cloud computing created FinOps to bring financial accountability to infrastructure consumption. As explored in CIO.com’s breakdown of FinOps expanding beyond traditional cloud costs, managing variable enterprise technology costs requires unified collaboration between engineering, finance and business leaders.  AI requires the same discipline, but with a more direct connection between technical consumption, financial accountability and measurable business outcomes.

The key is connecting technical consumption metrics with financial and business outcomes:

Consumption MetricsBusiness Impact Metrics
Inference cost per transactionRevenue impact
Token consumptionProductivity improvement
GPU utilizationHours saved
Model utilizationAutomation rate & cost savings achieved

AI spending should become as transparent, measurable and accountable as any other strategic operating expense.

Financial discipline is no longer optional; it is essential for scaling AI responsibly.

From AI adoption to AI advantage

The organizations that lead the next phase of enterprise AI won’t necessarily deploy the largest models or spend the biggest budgets. They will make better AI investment decisions.

They will choose the right technology instead of the newest technology. They will redesign business processes instead of simply automating existing ones. They will build reusable enterprise capabilities rather than isolated pilots.

Most importantly, they will manage AI as an economic asset — not merely a technological one.

The future of enterprise AI belongs to organizations that maximize AI unit economics: scaling adoption, reusing capabilities across functions and maintaining disciplined control over infrastructure, inference, operations and governance costs while delivering measurable outcomes.

The future winners will not be those who deploy AI everywhere. They will be those who know where AI creates economic leverage — and where it does not.

Why AI TCO is so tricky — and how to start calculating it

Achieving return on investment is impossible without knowing the total cost of ownership (TCO) of an initiative — and when it comes to AI, CIOs are finding cost calculations anything but straightforward.

Subscription and token costs are a big part of the calculus, but several other factors go into the cost of AI projects, says Ben Schein, chief AI and analytics officer at AI data platform provider Domo. Chief among those are cloud infrastructure costs and the human time involved in guiding or correcting AI outputs, he notes.

In addition, many organizations have multiple divisions using different AI tools for vastly different purposes.

“There’s not like a single ledger,” Schein says. “Right now, and maybe for the foreseeable future, there’s sort of like a multiple ledger approach to how all this works.”

A shifting paradigm

While token costs have dropped significantly in the past two years, costs vary wildly between models and AI providers, and the price drops are often offset by increased usage. And AI providers have also explored other kinds of consumption-based pricing, including API calls, compute time, or documents processed.

All this makes it difficult to measure TCO, Schein says.

“You have sort of these subscriptions, you have the consumption and the tokenization, you have some of the infrastructure you might be paying for,” he says. “There’s also a human tax that introduces new time for verification and review, and if the AI is sloppy or creating slop, you might be inadvertently adding to your costs without knowing it.”

It’s difficult to measure TCO because AI doesn’t have a single cost center, agrees Shane Cronin, head of FinOps and ITAM services at systems integrator SHI.

“By the time you’re looking at the bill, you’re dealing with token consumption, cloud infrastructure, multiple AI models, governance tooling, integration work and, increasingly, autonomous agents making decisions across systems,” he says.

IT leaders at many organizations still define AI success through narrow technical metrics instead of prioritizing business outcomes, Cronin adds.

“Calculating token costs is relatively straightforward,” he adds. “Calculating whether those tokens actually created measurable business value is much harder. That’s where most CIOs are today.”

Unpredictability and hidden costs

Michael Moran, chief technology and information officer at contact center outsourcing provider NQX, sees several other factors leading to further unpredictability over AI costs.

For example, data center costs are rising, AI vendors are starting to shift from subsidized pricing to profitability, and organizations have increasingly complex AI use cases, he says.

“IT leaders should temper expectations that AI inherently reduces costs,” he adds. “Instead, it’s important to understand that full automation is likely to be prohibitively expensive for most enterprises, and that brands will need to balance AI investment with human engagement strategies that improve long-term value rather than cut costs in the short term.”

If AI implementations work exactly as expected right out of the grate, TCO should be relatively easy to calculate, he says. But agentic AI implementations often require much more human training and intervention than expected.

“These are the hidden costs that are often underestimated or ignored altogether when initially calculating TCO,” Moran adds.

Visibility is the first step

Chris Cagnazzi, chief innovation officer at IT solutions provider Presidio, is one IT leaders seeking to get a handle on the complexity of calculating AI TCO by applying playbooks from the cloud migration era.

Cagnazzi has adapted Presidio’s cloud cost optimization platform, PRISM, to track AI costs internally and to help customers do the same.

The first step toward tracking AI costs is visibility, he says. IT leaders should know every model running across their organizations, the cost per user per month, and what kinds of prompts each user is writing, he explains, adding that Presidio is using real telemetry to track internal AI use, as well as internal tools to direct prompts to cost-efficient AI models.

The second, more difficult, step is turning visibility into action, he adds. “You have to think about mapping the usage back to the owners, whether it’s users or groups,” he says. “Then you look at, what are some of the anomalies? And if you’re looking at those anomalies, do you have governance in place around overspend?”

What Presidio has found is that the bill for AI services represents only about 30% of the total cost, Cagnazzi notes.

“The other costs really lie in areas around the hidden AI stack,” he says. “Those things around orchestration or retrieval, observability of the guardrails, or the rereads and the redos. There’s a lot of cost that people are missing.”

While traditional IT costs can be fairly predictable, AI costs are driven by usage and can increase because employees are repeatedly using inefficient prompts, Cagnazzi says.

“The spend is hard to forecast; it’s hard to see the true hidden costs behind the bill,” he adds. “If that prompt is less efficient, it might produce a bill that’s 100% higher than what it should be.”

The good news, says Domo’s Schein, is that IT leaders have a lot of variables to play with to control AI spending. They can encourage users to use more cost-efficient AI models, they can track employee usage of AI, and they can test different prompts and other interactions for cost effectiveness, he says.

“The price spread on the different models is crazy,” he says. “You could say, ‘I have no ROI on this investment; if I could get the same outcome with a model that costs one-30th as much, I may have ROI.’”

Google adds pay-as-you-go Gemini pricing as enterprises seek control over AI spending

AI agents could make software development and other enterprise tasks more productive, but they are also making technology spending harder to predict. Unlike traditional software licenses, the cost of running an agent can vary depending on the models it uses, the number of tokens it consumes, and how long it runs.

Google on Wednesday added new pricing options, discounts and cost-management tools to Gemini Enterprise that it says are aimed at helping enterprises reduce the cost of certain AI workloads while giving enterprises better visibility into where their AI budgets are going.

As part of the new pricing options, the hyperscaler introduced a pay-as-you-go model and Flexible Savings Plans (FSPs).

While the pay-as-you-go model allows enterprises to pay for the compute and tokens they consume instead of committing to a base subscription, which in turn avoids paying for empty seats or unused capacity, the FSPs offer discounts of 10% for one-year commitments and 20% for three-year commitments on Gemini Enterprise spending.

Flexible pricing lowers barriers, but adds new trade-offs

For enterprise teams and their CIOs, the pay-as-you-go model lowers the barrier to adoption and is better suited to experimentation, temporary projects, and agent workload bursts, said Stephanie Walter, practice lead of AI stack at HyperFRAME Research.

“While Per-seat pricing forces you to buy capacity before you know if an idea is worth it, the pay-as-you-go lets you spin up an agent experiment on a Friday afternoon and only pay for what it actually burns,” echoed Manoj Chandra Jha, principal analyst at Nord-IQ Research.

That means the newer pricing model also removes procurement friction from the experimentation loop, Paul Chada, cofounder of agentic AI startup Doozer AI, pointed out.

“When a pilot requires a license commitment, every experiment needs a business case. When it’s metered, an engineer can run the pilot on Tuesday and show finance a real bill on Friday. That shortens the distance between idea and evidence, which is where most enterprise agent programs die,” Chada said.

However, these advantages come with their own set of trade-offs, especially predictability.

“One user request can trigger an opaque chain of model calls, reasoning steps, and tool invocations, so consumption can grow much faster than employee headcount with the possibility of surprise bills at the end of a billing period,” Walter said.

That unpredictability also means the new pricing model does not automatically translate into cost savings, echoed Jha.

“It’s mostly a shift, not a discount. But matched to the right workload, it can save real money: bursty, unpredictable agent usage no longer subsidizes idle seats, while steady, high-volume usage may still be better suited to a committed plan. The savings come from matching each workload to the right pricing model, not from pay-as-you-go being cheaper by default,” Jha added.

However, the FSPs have their own caveats, especially the three-year plan.

While the FSPs can offer meaningful savings for enterprises with steady or growing AI usage, the three-year commitment is harder to justify in the wake of models, prices, and application architectures changing so quickly, Walter said, adding that the commitment is not just financial but also about choosing a platform as well.

Further, the analyst cautioned that the FSPs are less suitable for enterprises that have yet to establish a reliable baseline for consumption, as committing too early could turn an unpredictable operating expense into a predictable overcommitment.

Currently, FSPs are available for self-serve customers and customers already on enterprise agreements.

The pay-as-you-go model, though, remains only available to select customers with the hyperscaler planning a broader rollout “soon”.

Deferred execution trades speed for lower inference costs.

In addition, the hyperscaler is introducing a third cost-cutting option that is based less on how much an enterprise consumes than on how quickly it needs the result.

The option, named deferred execution pricing, will allow enterprises to mark eligible agent workloads for execution during off-peak capacity windows, with Google offering discounts of up to 50% on inference costs in return, the hyperscaler said in a statement.

Deferred execution pricing, it added, is aimed at workloads that can tolerate delays, potentially giving enterprises a way to lower costs for background tasks and other agent workloads where an immediate response is not essential.

That up to 50% discount in inference costs, Walter pointed out, can be material for CIOs at scale.

However, they must decide which work can safely wait and whether delayed tasks still meet the business requirement, Walter cautioned, adding that Deferred execution fits tasks such as evaluations, document processing, indexing, batch summarization, code analysis, and other background tasks.

Further, the analyst warned that CIOs also need to consider a “development tax” when considering deferred execution: “If agents need to be redesigned to accommodate real-time vs. asynchronous execution, the engineering effort required to build and maintain those different workflows can offset some of the savings.”

But even for workloads that can tolerate those trade-offs, the option will not be immediately available, with Google initially limiting deferred execution pricing to select workloads only. Details of which workloads are eligible were not immediately available.

New FinOps tools target AI spending visibility

Separately, Google is also adding an AI spend anomaly detection capability with root-cause analysis and pairing centralized billing reports with a FinOps agent that can generate natural-language summaries of where AI budgets are being spent.

While the anomaly detection capability is designed to flag projects where AI spending is trending higher than normal and identify the top three SKUs driving the increase, the FinOps agent is intended to make it easier for CIOs to understand where their AI spending is going.

Anomaly detection as a feature, according to Walter, can be valuable for agentic workloads because their consumption can increase through loops, retries, or unexpectedly long execution chains that users never see.

“Identifying what is driving a spike shortens the investigation for CIOs and enterprise teams,” Walter added.

The FinOps agent, meanwhile, Jha said, could help CIOs and other business leaders ask questions about AI spending that traditional dashboards were not designed to answer. However, its usefulness will depend on the quality of the underlying cost attribution, as the agent can only explain spending it can accurately associate with particular teams, projects, or workloads, Jha added.

Nobody knows where their AI budget is going

The early stages of AI adoption focused primarily on getting organizations to adopt AI-enabled solutions such as copilots, agents, assistants and AI-powered workflows. The priority was demonstrating that AI had the potential to deliver value, not scrutinizing every dollar spent along the way.

At the time, this made sense. There were relatively few people using AI, budgets were often funded through innovation initiatives and the cost of experimentation was low compared to the potential upside.

Organizations treated AI spending as a learning experience, so finance teams had little reason to scrutinize every model call or workflow because the priority was learning what AI could do, not optimizing what it cost.

When AI moves from pilot to production

After AI goes from pilot programs to production environments, the financial implications of using AI become very different. Every time a user enters a prompt, calls a model, retries an action, takes an action with an AI agent or executes an AI workflow, the total amount of money spent on AI goes up.

The amount at stake is rising quickly. Gartner expects worldwide spending on AI to reach $2.59 trillion in 2026, an increase of 47 percent from 2025. As companies put AI into more products and daily tasks, even small inefficiencies will repeat across millions of requests and add up to substantial costs.

The movement toward autonomous AI agents will only accelerate this trend. An agent may take much longer to perform its task than a single chatbot interaction. It may also use several different models, interact with other systems and continue to operate autonomously until it completes its task. Since organizations are likely to deploy many agents, AI costs will increase due to two factors: more people using AI and AI performing more work.

These concerns are already affecting which projects survive. Gartner predicts that more than 40 percent of projects involving AI agents will be canceled by the end of 2027 because of rising costs, unclear value or weak controls. A company may approve an agent because it works during a pilot, then reconsider it once thousands of people begin using it and every task triggers several paid requests.

A small number of experimental interactions will become thousands or millions taking place throughout an organization each day. Organizations that were previously asking how to expand their use of AI will now be asking: Do the benefits derived from each AI initiative outweigh the long-term operational costs associated with it?

This change in perspective is a good thing. Organizations are starting to look at AI as an operational expense instead of a shiny new toy. They are now evaluating AI initiatives based upon whether the initiative’s benefits exceed its long-term operational costs, rather than evaluating them based upon enthusiasm for the new technology.

Organizations can see the total AI bill but not where it comes from

The biggest issue for organizations is understanding where those costs originate.

Model providers typically send invoices detailing total usage for a given period. Companies know how much they spent on AI, but they often cannot explain which workflows generated those costs, whether that spending created meaningful business value, or which teams should ultimately own it.

Research from the FinOps Foundation shows that practitioners rank controlling the cost and use of tokens in software delivered as a service as their top concern related to AI. The reasons include bills that reveal little about what caused the expense and systems that provide no built-in way to trace costs back to the people or work responsible for them. A total on an invoice cannot tell a company whether one useful product feature caused the expense or whether an agent repeatedly called a model without improving the result.

That lack of visibility represents a significant blind spot. Imagine receiving your monthly cloud infrastructure bill without knowing which applications consumed your compute capacity, or receiving your monthly utility bill without knowing which buildings consumed your electricity. Most organizations would never tolerate that level of uncertainty elsewhere in their technology stack, yet many are currently managing AI spending in precisely this way.

Without attribution, organizations cannot determine which AI systems are delivering measurable business value, which workflows are inefficient or where unnecessary costs are accumulating.

Optimizing AI spending requires visibility into workflows

Much of today’s discussion about optimizing AI spending revolves around selecting a lower-priced model or negotiating better pricing with model providers.

Those discussions are valid because pricing is one of the few variables organizations can easily measure. Lower model prices alone will not solve the problem. As AI becomes embedded in more workflows and autonomous agents perform more work, organizations often consume far more tokens than they save through lower pricing. The greater opportunity to reduce AI costs often lies in workflow design.

OpenAI provides a clear example. Developers often send the same instructions or previous conversation back to a model each time they make a request. OpenAI introduced Prompt Caching in 2024 so developers could reuse material the model had recently processed and receive a 50 percent discount on those input tokens. The company later increased the discount to 75 percent for repeated material sent to its GPT-4.1 models. The savings come from changing how an application sends information to the model, which means a company can lower its bill without choosing a less capable model or negotiating a new contract.

Optimizing AI spending therefore requires organizations to develop visibility into how work flows through their AI systems. Organizations need to understand how agents interact with one another, how workflows execute, where redundant processing occurs and which steps provide the greatest value relative to their cost.

They also need to know when a system retrieves information it never uses, repeats a failed request, sends the same material several times or calls an expensive model for work that a cheaper one can complete. Each decision may add only a fraction of a cent to one task, but the same mistake repeated across millions of tasks can erase the financial benefit the system was supposed to produce.

Once organizations gain this level of understanding, they can optimize intelligently instead of simply selecting the least expensive model.

AI spending will drive organizations toward greater selectivity

Over the last several years, the AI industry has focused on determining whether AI belongs everywhere. The next stage will focus on identifying where AI creates the greatest value.

Some workflows will produce enough business value to justify substantial AI investment, while others simply will not. Long-term success will depend on understanding where AI creates meaningful value, where the costs outweigh the benefits and which AI systems actually justify their ongoing expense.

Companies should begin by recording which team, product, customer and task caused each paid request. They should compare that expense with the result the system produced, set limits that stop agents from retrying work indefinitely, and alert the people responsible when the cost of a task rises unexpectedly. Engineers can then inspect the costly work, remove repeated steps, reduce the amount of information sent with each request or choose a less expensive model when the quality remains acceptable.

A monthly invoice arrives too late and says too little. Companies need to trace spending while the work takes place, assign responsibility for it and decide whether the result earned its cost. Those practices will help leaders determine where AI deserves more investment, where the system needs repair and where it should be turned off.

The GPU bill is the new AWS bill

The call usually opens with praise. The AI feature shipped on time, users love it and engagement charts are pointing the right way. Then finance closes the quarter, and the feature everyone celebrates loses money on every single request. That is the part the CTO called about. I get some version of this call every week. I work in developer relations at a GPU cloud provider in Silicon Valley, putting me in the room, or at least on the video call, when engineering teams decide how to buy and run AI infrastructure. The longer I do this work, the more familiar the pattern becomes. I watched companies learn cloud-cost discipline in the 2010s, usually after an end-of-month bill delivered a nasty surprise. GPU spending is the same lesson with two important changes: the hardware costs roughly ten times more per hour, and mistakes pile up faster. We’ve seen this movie.

We know the ending

Before moving into AI infrastructure, I spent years in data analytics at an automotive software company. One part of that job was cleaning up a decade of accumulated cloud enthusiasm, which sounds harmless until you inherit the bill. After we consolidated three overlapping analytics platforms into one, we cut about 220,000 dollars a year while keeping every capability intact. That money accumulated through reasonable-sounding subscriptions, one after another, because nobody owned the basic question: what did it cost to produce those numbers? The industry still hasn’t solved it. Flexera’s annual State of the Cloud research has for years found that organizations estimate more than a quarter of their cloud spend is wasted. An entire discipline, backed by the FinOps Foundation, grew around squeezing that waste back into a manageable shape. It took most companies years to learn those habits.

What bothers me is simpler: I keep seeing solid engineering teams drop that discipline the moment the purchase order says GPU. AI spend gets treated like a bold bet instead of an operating cost, and then the ordinary scrutiny disappears. That is where the trouble starts. The waste patterns of 2015 come back wearing 2026 pricing. The invoice tells you what you paid, separate from what you earned.

The invoice tells you what you spent, separate from what you got

GPU capacity is priced by the hour, so teams naturally budget and report by the hour. It feels neat. It lines up with the bill. And it hides the problem that really crushes margins. The number that decides whether an AI feature survives is cost per request: everything you spend on inference infrastructure divided by the requests you serve. Those two metrics line up only when your hardware stays busy. For user-facing AI, that stays rare. Traffic moves with human attention, so it flares for a few hours and then drops off a cliff.

One team I worked with had reserved a cluster built for a peak that showed up for about two hours a day. On the invoice, the hourly rate looked almost cheap. Once we divided it by served requests, it was ugly, and the team had honestly seen it for the first time when we ran the numbers together on a call.

That division is the most useful exercise I can offer a reader of this column. Take last month’s total inference spend. Divide it by the number of requests you served. If the answer makes someone in the room go quiet, you have found money and you found it with arithmetic a spreadsheet has been waiting to do for you.

Workload shape, rather than vendor choice, decides the right pricing model

When the number looks ugly, the instinct is to push for a better rate or go hunting for another provider. I sell GPU capacity for a living, so I’ll say it plainly: the rate is rarely the issue. The unit price of AI compute keeps falling; Stanford’s AI Index has documented inference prices dropping by orders of magnitude in just a few years. That still leaves a team paying for capacity it barely touches. Waste eats the discount whole.

The fix lasts longer when you match the buying model to the workload itself, which is a point Andreessen Horowitz made well in its guide to the cost of AI compute: access to compute matters less than the shape of the commitment you sign for it. AI workloads usually split into two very different cases, and they want opposite deals. Sustained work, such as training runs, fine-tuning and batch processing, keeps hardware busy around the clock. This is what reserved or dedicated capacity is for. The economics are simply better when the machines stay hot. Reserved or dedicated capacity is made for that, and the per-unit economics pay you back for the commitment. Spiky work, which covers almost everything with a person on the other end, is the reverse. Usage-based pricing earns its markup there, because you only pay when you serve. The per-unit price goes up and the total bill drops. Finance teams resist that sentence until the numbers hit their own sheet.

The best production setups I see are hybrids. A team keeps a modest baseline, sized to the floor of traffic, the level demand almost never sinks below, and lets usage-based capacity soak up the rest. Teams under roughly ten million tokens a month often skip infrastructure entirely and stay on a model-as-a-service API until volume justifies the switch. The reserved slice stays busy. The bursts stay covered. The architecture quietly records a choice the team meant to make, which is rarer than it should be.

3 questions are cheaper than a contract

When a team asks me to review a GPU commitment, I keep coming back to the same three questions, and I would rather they ask them before the signature than after it.

  1. What does our measured utilization curve look like? Skip the neat projection in the deck. Instrument a week of production traffic before signing anything at all. Teams almost never predict their own curve correctly, and that surprise costs nothing before the contract, then plenty afterward.
  2. What is our cost per request at ten times today’s volume? Scale can change the answer, sometimes in our favor. Spiky demand may smooth as users spread across time zones, shifting the calculation toward reserved capacity later. If nobody in the room can answer, the organization is buying a snapshot, not a strategy.
  3. What would switching cost us? Open-weight models let us rerun the analysis with any provider, then act on what the numbers say. Proprietary endpoints tie your costs to another company’s pricing whims. Either route can make sense, yet flexibility has a dollar value and deserves space beside the hourly rate.

The discipline is the differentiator

Outside my day job, I’ve judged more than eight AI hackathons this past year, at Microsoft offices in Chicago and Mountain View, plus events with OpenAI and Google Developers Group. Even there, surrounded by teams building through a weekend, I can see the production problem waiting ahead: brilliant models, minimal thought about what serving them will cost. Nobody wins a hackathon with a unit economics slide. Plenty of companies quietly fail without one.

Years in data analytics left me with a conviction I repeat to every team willing to listen. A dashboard nobody costs out is a liability; an AI feature carries the same risk. The companies that survive the next pricing cycle will be the teams able to name their cost per request from memory and explain their infrastructure in one sentence, with a week of traffic data behind it, rather than those squeezing the lowest hourly rate from a vendor. Ten years ago, cloud bills taught engineering leaders to ask what their systems cost. Now the GPU bill is asking again, at ten times the stakes. The lesson lands harsher now: guessing survives only until the next ugly bill arrives at the worst time. The teams that move first will claim the margin everyone else is still chasing.

CIOs earn AI reprieve, but ROI pressure is surging

2026 arrived as the year AI ROI would need to get real. After years of experiments and pilots that largely failed to scale, CEOs’ No. 1 priority for CIOs was to achieve demonstrable benefits from AI investments.

Under that pressure, CIOs started to feel the heat, with 71% of IT leaders in February saying they believed they had until midyear to prove AI value or face budget or job fallout, according to a survey published by AI platform provider Dataiku. For many, experience may have informed that anxiety, as three-quarters of CIOs surveyed then also said they had remorse over at least one major AI vendor or platform selection made in the past 18 months..

Six months later, and passed that midyear mark, CIOs who have set their course for AI ROI are finding that destination remains elusive. Still, there hasn’t been a spate of CIO firings, and AI spending continues to grow. About 71% of organizations plan to increase AI spending this year, but only 27% expect near-term ROI, according to recent research by IT solutions provider TEKsystems.

That metric aligns with findings from CIO.com’s State of the CIO survey from earlier this year, when 40% of IT leaders said some AI initiatives (between 30% and 70%) were meeting ROI goals. While progress remains the same, the drumbeat to prove value goes on.

“CIOs are feeling pressure to demonstrate that AI is delivering measurable business value, not just experimentation,” says Jed Dougherty, SVP of AI and platform at Dataiku. That’s because, while CIOs’ worst fears haven’t been realized, organizations are putting greater scrutiny on AI investments, he adds.

“The CIOs who are succeeding aren’t deploying AI everywhere,” he says. “They’re building the governance, data, and operational foundation that lets the business scale AI responsibly and demonstrate real outcomes.”

A pronounced focus on AI spending and costs

Bob Hutchins, CEO at AI advisory firm Human Voice Media, believes CIOs’ early year anxiety wasn’t likely based on any formal deadlines.

And if any CIOs have been fired since February because they missed AI targets, those changes are likely hidden from public scrutiny either in reorganization efforts or other leadership changes, he says.

“Midyear came and went and there was no apparent bloodletting,” he adds. “I haven’t seen credible evidence of mass firings of CIOs due solely to missing return on investment targets for artificial intelligence.”

But organizations do seem more focused on spending their AI budgets wisely, Hutchins notes. Nearly half of all organizations surveyed recently by KPMG have delayed, stopped, or scaled back AI projects due to budgetary constraints, he says.

“Companies are stopping poorly performing projects, scaling back pilot programs, decreasing the number of vendors they use, creating cheaper models of products and services, and giving more control over AI approval to the financial department,” he adds.

Ryan Ries, chief AI and data scientist at AI and cloud consulting firm Mission Cloud, also sees IT leaders still under pressure to improve AI results.

While firing a CIO midyear looks bad on an earnings call, IT leaders now face budget triage efforts related to AI, he says.

“Money still flows to projects with a hard number attached,” he adds. “Pilots without one get quietly starved but not killed outright. The AI landscape is constantly changing, and companies are trying to figure out all the new tools like coworking and coding solutions.”

Moreover, facing increased uncertainty over AI pricing, CIOs are re-examining AI adoption metrics and becoming more aware of the hidden costs of AI.

Cost control is also receiving greater emphasis as some early agentic forays have shown how an AI agent can cost more than an employee without limits in place.

CIOs still on the hot seat

CIOs are also finding that it is taking their organizations more time to figure out how to use AI tools to their full advantage, and that they are constantly reacting to errors, Ries notes.

As a result, IT leaders appear to be putting in more effort to find AI value than they were earlier this year, he says. “Fewer are succeeding than leadership wants to admit,” he adds.

While most organizations were experimenting with the so-called “art of the possible,” IT leaders seeing success are focused on attaching a metric to every AI project before launch, not after, he says.

Many IT leaders still aren’t taking that approach, however. “The CIOs still stuck are the ones running pilots that never graduate to production, usually because nobody can explain what the model is doing under the hood, and teams are trying to answer the wrong questions with AI,” he says.

It’s possible that CIOs have gotten a reprieve due to the ongoing complexity of the AI ROI mandate, Ries adds.

“Boards don’t fire on a spreadsheet’s calendar, but the underlying pressure was real, and it hasn’t eased,” he says. “Instead of a hard cutoff, CIOs now face constant reporting. Monthly board briefings on AI performance are becoming standard, not optional.”

CIOs should remain on their toes and focus on driving AI value, Ries says. “The fear of a July guillotine was overstated,” he adds. “The fear of ongoing, permanent scrutiny was not, if anything it has increased, due to how quickly cost overruns can happen.”

Budget pressure is real

Like other observers, Mridul Nagpal, CTO and co-founder of AI software development company Krazimo, sees a growing focus on AI budgets and untargeted spending.

“The pressure is real, but it’s reshaping spend more than cutting it,” he says. “What’s actually at risk is the undifferentiated AI budget — the ‘we’re doing AI’ line item with no outcome attached.”

Many CIOs are now starting to show AI value, but by narrowing their approaches, not expanding them, he adds.

“CIOs who funded broad experimentation are the ones sweating; CIOs who tied spend to a specific, measured workflow are defending, and often growing, their budgets,” he says. “The fallout is landing on unaccountable AI spend, not AI spend per se.”

The CIOs showing returns have quietly killed sprawling AI pilot portfolios and doubled down on a handful of use cases that reached production, Nagpal adds.

Like Ries, Nagpal believes that earlier CIO fears were a bit overblown and, at the same time, they’ve gotten more time to prove AI value.

“Boards softened the ‘or else’ because the whole market discovered the pilot-to-production gap is real and hard, so the deadline quietly moved,” he says. “But the underlying expectation didn’t disappear — it matured from ‘show me AI’ to ‘show me AI that pays for itself.’”

Your AI bill just came due. Nobody warned you it would look like this

For most of the last two years, I priced AI the way I’ve priced every piece of software in thirty years of building technology: a number per seat, budgeted once, revisited once a year. It’s the only pricing model most IT leaders have ever had to plan around.

Then I started building AI systems that ran at real usage instead of pilot usage, and that instinct didn’t survive contact with an actual bill. A workflow that ran a handful of times a day during testing runs hundreds of times a day once a team adopts it. The per-token price never moved. The bill did, and it climbed in multiples, not percentages.

I’d made the mistake almost everyone makes. Software has a seat price. You buy access, use it as much or as little as you want, and the invoice barely notices. AI doesn’t work that way. Every query, retrieval and step an agent takes to finish a task consumes tokens, and tokens are metered like electricity, not sold like a subscription.

I learned how AI economics works by watching the shift from the inside. I was building a system that combines retrieval, workflow automation and human review to support a marketing team’s daily work. The finance conversations that came with moving from pilot to daily use are what made it click.

Here’s the disconnect that blindsides budgets. Most IT leaders know roughly what they’re paying per API call or per seat. Fewer know what they’re actually spending, because those two numbers move independently.

The mental model that fails first

Seat-based pricing trained a generation of IT leaders to treat software cost as fixed. Add ten users, the license line moves in ten predictable increments. Budget it once a year and move on.

AI breaks that model. Two employees on the same seat can generate wildly different bills depending on what they ask the system to do. Someone summarizing a short email uses a fraction of the tokens someone running a multi-step research task across several documents does. It’s the same license, but completely different cost.

I’ve watched this happen more times than I can count. The deeper a task goes, the more tokens it burns, and that’s true whether the task is trivial or genuinely valuable. Cost tracks the depth of the work, not a judgment about who’s using the tool well.

Nobody explains that part upfront. Adoption is supposed to be the win. In token economics, adoption is also the thing that drives the bill up.

And here’s the twist that makes it counterintuitive: per-token prices have fallen fast. Stanford’s 2025 AI Index Report found that inference cost for a system performing at GPT-3.5’s level dropped more than 280-fold between November 2022 and October 2024. Look only at the price sheet and you’d swear AI got cheaper. Then look at what teams are actually doing about it. The FinOps Foundation’s State of FinOps 2026 report found that 98 percent of FinOps teams now manage AI spend, up from 31 percent two years earlier, and they named it their top forward-looking priority for the year. That kind of urgency doesn’t gather around technology that’s getting easier to forecast.

Where the surprises hit

The surprises show up in three specific places. I’ve had a direct hand in all three while scaling systems from pilot to production.

The first is agentic workflows. A simple prompt and response might use a few thousand tokens. An agent that plans a task, retrieves documents, calls a tool, checks its own output, and retries when something looks off can burn ten times that for a single request. Every one of those steps gets billed. When I moved a workflow from single-shot generation to a multi-step process with retrieval and review built in, token use per task jumped in a way the original cost model never accounted for. Nothing was broken. The system was just doing more, and more is exactly what usage-based pricing charges you for.

The second is context growth. Retrieval-augmented generation pulls source material into every query so the model has something accurate to work from. The more documents you retrieve, the more context you feed in, and every token of that context gets billed on top of the question itself. A well-tuned RAG system retrieves exactly what’s needed. A loose one retrieves everything that might be relevant. The difference between those two shows up on the invoice.

The third is the success penalty, and it’s the one that catches leaders off guard, because it looks like good news until the invoice says otherwise. A pilot gets approved on light usage and a small budget. It works. People like it. Usage climbs faster than anyone modeled, because that’s what adoption looks like when a tool is genuinely useful. Most cost models don’t leave room for growth that fast while the project is still labeled a pilot.

What cost governance actually looks like

None of this means AI adoption should slow down. It means cost has to become an architecture decision instead of a financial afterthought that surfaces once the system is already in production. Four disciplines carry most of the weight.

Route models by task complexity. Not every task needs a frontier model. Routine summarization, formatting, and basic classification run fine on smaller, cheaper models with almost no quality loss, which frees the frontier model for the work that genuinely needs that level of reasoning. This single change has bent my cost curve more than anything else I’ve tried.

Monitor at the workflow level, not the account level. Your total monthly AI spend tells you almost nothing. Knowing that one specific workflow accounts for sixty percent of it tells you exactly where to look. Granular tracking built in early is the difference between explaining a cost increase to leadership in one sentence and spending a week on forensics.

Treat agentic pipelines like any production system that can run away from you. That means retry limits, timeout logic and a hard stop when a task loops longer than expected. A recent piece on why most agentic AI projects stall before they scale makes a related point: governance, not the model, becomes the real constraint once these systems move from demo to production. Cost control is a direct extension of that same discipline. Skip the circuit breaker and the risk isn’t just a bad output. It’s an uncapped bill.

Plan for price change instead of betting against it. Current pricing is still partly subsidized by vendors competing for market share, and that won’t hold. GitHub’s move to usage-based billing for Copilot this year previews where the rest of the market is going. Build your cost model on the assumption that per-unit pricing normalizes upward, not that it stays this generous.

These surprises don’t come from a vendor overcharging you or a system malfunctioning. They come from applying a twenty-year-old cost model to a technology that was never built to work that way.

You don’t need to fear usage-based AI pricing. You need to stop budgeting for it like a subscription. Getting ahead of this doesn’t mean spending less on AI. It means knowing what each dollar bought before finance asks, and fixing the architecture instead of the budget.

The number at the bottom of the invoice matters less than whether you can explain every line above it.

Beware of the AI pilot trap

For many organizations, AI is proving easy to pilot but difficult to scale. Pilots often look inexpensive because they run on narrow datasets with a handful of users, explains Ben Schein, chief AI and analytics officer at cloud software company Domo. “But the cost lives in deployment, the moment you connect that capability to real workflows and the systems of record behind them,” he says. “That’s when the real bill appears.” So CIOs must always budget for the gap between when it works in a demo and when it produces governed and durable value.

width="1240" height="828" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Ben Schein, chief AI and analytics officer, Domo

Domo

Organizations can easily get caught out because they run pilots as a technology experiment instead of a business initiative, he adds. “The interesting question is never whether AI can do the thing in a demo,” he says. “It’s whether it should run in this process, and whether it survives contact with production.”

There’s also a lot of pressure on IT teams to be doing something with AI simply because everyone else is, says Naren Gangavarapu, chief transformation and AI officer at Australian Cruise Group.

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Naren Gangavarapu, chief transformation and AI officer, Australian Cruise Group

Australian Cruise Group

He calls it AI theater because there’s a big show around AI even though there aren’t that many successful applications of the technology in production environments.

AI costs out of control

According to John D’Emic, CTO at AI observability platform Revenium, one of the big traps when running a pilot is failing to anticipate how quickly consumption can spiral as adoption grows. “As an example from our own engineering org, back in May, a developer opened an AI coding session on his laptop, and it stayed open for four days,” he says. “By the time it closed, it had run 4,819 calls and cost us $3,762. We didn’t budget for this, and no alert fired. But that one session cost more than a lot of teams spend on their entire monthly AI tooling.”

width="1240" height="828" sizes="auto, (max-width: 1240px) 100vw, 1240px">

John D’Emic, CTO, Revenium

Revenium

While this showcases how a developer can make a costly error, Dmitriy Anderson, CIO and digital and social commerce leader at home and gardening retailer Leroy Merlin South Africa, believes the pilot trap frequently happens when employees with little or no software development experience vibe code applications. “It doesn’t matter if you can create something in 15 or 20 minutes if the result is AI slop,” he says. “Think dirty code, no consideration for safety, security, and possible data exposure.” In most cases, these pilots are developed with one of the frontier apps, and someone probably used their personal AI subscription, so the costs are negligible, he adds. But if you have a company of several thousand people, and you now want to roll this tool out more broadly, that’s where costs can get out of control.

This scenario is only exacerbated by the introduction of agentic AI, D’Emic adds. “Agents don’t spend money at human speed,” he says. “In the old cloud days, an engineer could spin up infrastructure in minutes and finance might not see the bill for a month, which was painful but recoverable. Agents, though, call APIs around the clock without waiting on anyone’s approval.”

Mind the trap

While cost is a big factor in the AI pilot trap, it should be treated as a symptom of a bigger problem, says Schein. The underlying issue is governance and observability. “An autonomous workflow can fan out into more queries, API calls, and model invocations than anyone scoped,” he says. “So if you can’t see what it’s doing, and spend compounds quietly, you only find out once the invoice arrives.”

In a recent LinkedIn post, Anderson outlined how in just six weeks he built a platform for a fraction of the sticker cost using three AI models orchestrated together. The traditional estimate to build the same tool would have required 2,472 engineering hours from a team, and was expected to take around nine months. “I went through the proper engineering steps and planning, and made sure the application passed a series of cybersecurity frameworks,” he says. “The purpose of this exercise was to showcase that AI can still speed up the process even if you take the time to work through the necessary steps. You can build with AI rigorously and securely.”

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Dmitriy Anderson, CIO and digital and social commerce leader, Leroy Merlin, SA

LMSA


So to turn AI experiments into enterprise value, every AI interaction must be attributable: who triggered it, against what data, on which model, and at what cost, Schein says. For each workload, be sure to ask how often it runs, which model tier the job actually needs, and what triggers it, human or automatic. “A frontier model on an automatic trigger and a small model called on demand are completely different cost curves for the same task,” Schein adds.

For Anderson, it’s helpful to use AI to highlight potential gaps, assumptions, or blind spots in your ideas early on. “When you start building an idea, ask the agent to interview you,” he says. “It will go through every phase and ask questions about the important facets of the process, from scalability and budget to deployment options. You can even make AI write a prompt for itself, because it knows its capabilities and quirks better than you ever will. It’s called meta prompting.”

Anil Inamdar, global head of data services for the Instaclustr BU at NetApp, suggests CIOs cost out the whole program, not just the demo. “Generally, the model itself is the cheapest part of the program,” he says. For him, it’s important to have security and governance people in the scoping meeting, not the launch meeting.

width="1240" height="827" sizes="auto, (max-width: 1240px) 100vw, 1240px">

Anil Inamdar, global head of data services. Instaclustr BU. NetApp

NetApp

He believes the pilot trap is also, or perhaps mostly, a sequencing trap. “A lot of teams are wired to build first and ask permission later, only to discover months down the line they can’t pass a security review or data privacy audit without a painful and costly rebuild. It’s also valuable to define what failure looks like before you define success.

“Pilots tend to die because of no result, which isn’t the same as a bad result,” Inamdar says. “Emphasize to the deployment team on day one that if a target result by a certain month isn’t seen, we shut it down. Otherwise, you’re funding a zombie pilot because everyone’s invested and no one wants to be the one to call it out.”

The case and model for real-time AI cost visibility at the infrastructure layer

I spend most of my time inside other companies’ engineering teams, building systems to track and optimize AI spend. The conversation almost always starts the same way. Someone pulls up a dashboard, points at a number bigger than it should be, and says some version of “we know it went up; we just can’t tell you why.”

I use an analogy for it: AI-era CIOs are like city planners optimizing a busy intersection. They can measure the volume and hear the pleas to fix congestion, but can’t tell whether a vehicle is a truck or a bike, or why it’s on the road. Without that, they can’t design the right fix, so they build a highway at great expense when the data would show all it needed was a bike lane.

Every model choice and budget conversation happens against that blurry picture, and the traffic gets heavier every quarter. Gartner expects worldwide AI spending to grow 47% this year, with agentic AI software up roughly 141%. By 2028, it projects an average Fortune 500 enterprise will run over 150,000 agents, up from fewer than 15 in 2025.

Why the cloud playbook can’t answer the AI question

AI presents a fundamentally different problem than cloud cost management, where we answered, “whose spend is this?” largely by tagging the resource. A VM has an owner, a bucket belongs to a team and FinOps optimizes from there. Billing was slow but acceptable, because spend moved inside predictable bands, a human provisioned each resource before it cost anything, and governance capped how fast costs grew. Surprises were unpleasant, rarely existential.

AI took that model, put it to the test and laughed it out of the room. A token call isn’t a resource the way a VM is, so tags have nothing to attach to. To make them affordable, hyperscalers run large models on shared, multi-tenant infrastructure, dropping the per-token price but stripping out the granularity needed to track and control spend.

One API key can carry a dozen workflows across three teams, and the consumer is often an autonomous agent, not a person. The failure modes are also new. An agent can drift off task, loop on a retrieval endpoint hunting for an answer it can’t find, and restart from scratch when it comes up empty. At a few dollars per million tokens, it seems trivial, until that loop runs across thousands of parallel sessions and clears five figures in a day. None of it trips a provisioning gate or maps to a taggable resource. The control points that cloud governance leaned on don’t exist for AI.

The consequences can be severe. Uber spent its entire 2026 AI budget in four months after Claude Code usage ran far ahead of its projection, and they’re far from alone. I’ve worked with plenty of large enterprises hitting the same thing without the headlines. In almost every case, the teams driving the spend weren’t accountable for the budget, and no one saw the scale until it hit an invoice. It’s a credibility destroyer, as these overruns erode margins and stakeholder confidence for the leaders on whose watch it happens, even when their only fault was doing their best with the tools they had.

The data backs this up. Recent DoiT research into enterprise finance leaders found that 89% of the organizations that rate themselves most mature at FinOps still overspent on AI last year, and by the widest margin in the study. Again, the teams with the best cost discipline overspent the most. They built excellent governance for human-gated, taggable resources, then tried to retrofit it for workloads that have neither. They see the overrun, just too late to stop it, and without the insight to know where to focus.

The instrumentation era got us closer, at a real cost

Our best answer was to instrument the application itself: wrap every LLM call in an OpenTelemetry span, inject cost-allocation metadata and propagate that context across a multi-agent workflow so token counts roll up to a budget owner. I’ve done it dozens of times, and it works: accurate per-request attribution, far better than splitting the monthly bill by best guess.

But it was always a workaround. You can only measure what you thought to instrument, and even when you do that, the answer often arrives too late.

One customer I worked with recently had built about ten agents that worked together to automate security and quality evaluation of their software, most of them running constantly. Their bill kept climbing with no obvious cause, and it took a long investigation to find why: one agent, running 24/7, was burning 20 times the tokens of any other, caught in an infinite loop that respawned a never-ending process every time it spun up. The signal was right there in the data, but it only surfaced after weeks of digging and real money out the door.

There’s an irony too. AI is supposed to buy back engineering time, and instrumentation spends it right back. You put senior engineers on measurement plumbing to control the cost of the thing meant to make them more productive. Plus, it’s fragile; every new model, SDK and agent framework is another integration to keep alive. “Instrument everything you might ever build” isn’t realistic when tens of thousands of employees and countless agents launch new workflows daily.

What changes when attribution moves to the infrastructure layer

The approach now emerging, and the one changing how I run these engagements, moves attribution from the application down to the infrastructure. Instead of teams describing their spend through code they wrote, you observe what’s running underneath them.

The mechanism is a kernel-level sensor, the same eBPF technology that security and observability tools use to watch system calls without touching the applications above them. It maps each unit of GPU, CPU, memory and network back to the process, container and request that caused it, then joins every outbound model call with provider cost data so token spend on Anthropic, OpenAI, Gemini or Bedrock is matched to the workload and agent that drove it.

This delivers the same answer OpenTelemetry can give you, at equal or better accuracy, but in real time and without the instrumentation work. Nothing has to be planned in advance, and nothing gets missed. For the first time, the data to manage AI spend is continuously available at an actionable grain. That runaway agent my customer dealt with would have been identified and addressed as soon as it began to snowball.

Governance stops being a top-down audit

Changes in governance is the part I find most interesting, because with continuous, accurate attribution, you can design governance purpose-built for AI rather than retrofitting a system built for an era with different physics.

Strategy still belongs at the top, with the CIO and finance leaders setting the direction, the budgets and the priorities. But daily responsibility for staying inside those lines shifts toward practitioners. Cloud governance was top-down because provisioning ran through the few people with the full business context to weigh it. With AI, spending decisions happen everywhere at once, across shared keys, agents and dozens of teams, so adherence has to sit with the people making them.

Real-time attribution makes that possible. An engineer can see the cost of what they’re building as they build it. A frontier-model call when a smaller model would do, a retrieval loop fanning out across a fragmented data store, an agent retrying a failed tool call in a tight loop — those stop being mysteries at the quarterly review; they’re caught while the work is still warm. And because the same numbers reach leadership, they can weigh spend against value and steer the portfolio rather than litigate a bill nobody can explain. Back at the intersection, the planner can finally tell the trucks from the bikes, and so can every engineer on the road.

What visibility is actually for

So, the biggest foundational roadblock (pun intended) that tripped up even the most disciplined teams is solvable now in a way it wasn’t a year ago. But solving it was never the real goal. Visibility is a means to an end; it just had to be solved first.

Real-time, granular attribution lets companies start treating AI as a managed investment instead of a pay-and-pray experiment. When you connect spend to the work, the work to an outcome and the outcome to the business case that justified it, you can govern AI with intention and put money where it earns its keep.

The last two years rewarded productivity to whoever shipped faster with AI. The next phase gets won on efficiency and impact, by whoever gets the most value per dollar of inference and can prove it. That’s a matter of strategy, and for the first time the data exists to compete on it. What you’re really adding for the customer, and what it costs, stop being a guess.

The vendor consolidation trap: When one throat to choke costs more than it saves

Vendor consolidation is sold as discipline. Fewer vendors, simpler architecture, better pricing through volume, one throat to choke when something breaks. Every one of those benefits is real on paper. The problem is that the biggest cost of consolidation rarely appears on the slide the procurement team uses to sell it internally, and it does not show up on the savings tracker until the first renewal cycle after the ink is dry.

Within CIO Mastermind’s topic-specific cohorts, which I sometimes facilitate, I hear a version of the same story often enough to recognize the pattern early. A consolidation program gets pitched against a strong multi-year savings target. The first year or two look good. Then a renewal arrives, the remaining vendor prices to the switching cost the company just built for itself, and a meaningful share of the projected savings quietly erodes. The company still ends up with fewer vendors. It does not always end up with the leverage the original business case promised.

What consolidation actually removes

What consolidation actually removes is competitive pressure on the vendor you keep.

That is the part most business cases leave out. Going from a dozen vendors in a category down to three or four feels like simplification, and it is. It is also a message to the vendors you kept about how expensive it would be for you to leave. The fewer live alternatives you maintain, the more accurately a vendor can price to your captivity rather than to the open market. A consolidation deck typically shows how many vendors are being reduced. It rarely shows how many of the remaining vendors could credibly be replaced inside a reasonable switching window. That second slide is the one worth building before the program starts. That capability is often missing.

Procurement teams are not being dishonest when they leave that slide out. Their incentive is to close the program and book the savings target, and the pain of a diminished market shows up two or three years later, on someone else’s dashboard. By the time the first hard renewal arrives, the people who built the original business case have often moved to a different project entirely, and the CIO who is still in the seat is the one negotiating from the position the program created.

The condition that decides the outcome

Consolidation programs succeed or fail on one question, and it has to be answered honestly before the program starts.

Can you walk from this vendor at renewal?

Not in theory. Not with twelve months of migration work. At renewal, inside the window the contract gives you, with a credible alternative that has been exercised recently enough to be real. If the answer is yes, the vendor will price to keep you. If the answer is no, the vendor will price to what you can absorb. Consolidation that leaves you unable to walk is a long-dated price increase with a celebratory kickoff meeting, dressed up as a savings program.

Contract language deserves particular scrutiny here, because it is where a lot of the false confidence comes from. Multi-year agreements often include price increase caps that look protective at signing. Those caps are usually written around the product as it exists at signing. Vendors may repackage functionality into new or higher-priced tiers, leaving the contractual cap covering less of what the company actually needs. A cap that looked airtight in the negotiation can end up covering a shrinking share of what the company actually pays for at renewal.

CIO.com has covered the leverage problem for years, including a piece on how to increase your renegotiation leverage with vendors that frames the handcuff problem directly. The advice in articles like that one is sound. The hard part is applying it in the middle of a consolidation program, when the procurement team is telling you that keeping alternatives warm is wasteful and the CFO is asking why the savings number is dropping.

The CIOs who hold their leverage tend to do one thing differently. They keep one credible alternative warm in every major category they consolidate, even after the primary vendor is chosen. Warm means more than a name on a shortlist. It means a live relationship with the alternative’s account team, some recent proof of concept, and at least one internal team that has actually touched the alternative’s platform. That readiness carries a real cost. Maintaining it may cost far less than an uncontested renewal can quietly take away.

The number that actually matters to the CFO

Most consolidation programs get measured against a single number: the savings projected in year one of the business case. That number rewards aggressive consolidation and quietly punishes the CIO who keeps an alternative warm, because the carrying cost of that alternative shows up immediately while the protection it buys only shows up at the next renewal, two or three years later. Judged against a one-year number, the cautious approach always looks worse.

The number worth tracking instead is the savings figure three years out, measured against what the business case originally promised. That is the number an aggressive consolidation program tends to miss once a full renewal cycle has run its course, and it is a fairer test of whether the program actually worked. It also reframes the conversation with the CFO. A carrying cost presented as insurance against a specific, quantifiable renewal risk is a different ask than a carrying cost presented as overhead, and it tends to get a different answer.

What to do if you inherited the problem

Most of the CIOs I talk to are not starting a consolidation program. They inherited one. They sit down in a seat where the leverage is already gone and the next renewal cliff is six or nine months out.

If that is where you are, the fastest way back to a real negotiating position is not to rebuild leverage everywhere at once. That approach takes years and asks the CFO to fund carrying costs across the entire portfolio before there is any evidence it will pay off. Pick one category instead, ideally not the largest one but the one where a credible alternative can be stood up fastest, and rebuild it inside twelve months. Speed matters more than scale here. A live proof of concept in a smaller category, exercised recently enough to be real, does more for your negotiating position than a partially built case in a larger one.

One proof that you can still move part of the portfolio changes the conversation at every other renewal table. A vendor who knows you have already done it once treats the next renewal differently than a vendor who has only heard you claim you could.

The second you cannot walk, the price stops being yours to negotiate. Most consolidation programs remove your ability to walk as their first move, and most CIOs do not realize they have given it up until the next renewal arrives and reminds them.

OpenAI targets heavy users with premium ChatGPT Business seats

OpenAI is introducing a higher-priced “Premium” tier for its ChatGPT Business offering, allowing enterprises to assign higher-capacity access to select users alongside standard licences – a move analysts said is about enterprise AI vendors redesigning pricing to capture more value from high-intensity workloads.

The company said the new tier provides “5x more usage than Standard” and “removes the five-hour usage limit,” enabling users to “take on larger projects and work with fewer interruptions.”

“Premium seats cost $125 per user per month, or $100 per user per month when billed annually,” OpenAI said in a statement. “Standard seats remain $25 per user per month, or $20 per user per month when billed annually.”

OpenAI said enterprises can “mix Standard and Premium seats across the same team” and “upgrade or reassign seats as business needs change,” with administrators able to “monitor usage across the workspace” and “manage billing, usage, and spend limits in one place.”

Vendors converge on seat-plus-usage pricing

Analysts said the introduction of a higher-capacity tier reflects a broader shift toward hybrid pricing models.

“Read this as vendors converging on a two-layer bill rather than abandoning flat pricing,” said Bhupendra Chopra, chief revenue officer at Kanerika. “There’s a predictable per-seat charge for everyday chat, and a separate metered charge for heavy agentic work.”

Chopra said vendors are packaging this differently. “Google bundles the first into Workspace and meters the second. Microsoft folds Copilot into its bundles and sells credit packs for agent runs. OpenAI has no productivity suite to hide the seat cost inside, so it built a higher seat tier instead.”

That model is reflected in current offerings. Microsoft 365 Copilot pricing positions Copilot as a per-user add-on to Microsoft 365, while Google Gemini enterprise pricing shows model and agent usage billed separately from Workspace plans. Anthropic’s Claude pricing similarly outlines subscription tiers alongside usage-based model access.

For CIOs, this changes how AI spending is managed, Chopra said. “Your AI budget now has a fixed component and a variable one, and they need different owners. Seat count is a procurement problem. Metered spend is a FinOps problem.”

Premium seats target high-usage workloads

OpenAI said Premium seats are designed for “your most active teammates,” citing use cases such as “organizing inventory,” “building marketing campaigns,” and “analyzing business performance.”

The company said Premium users can “take on bigger projects and keep work moving,” with “predictable weekly usage resets” and the option to add “shared workspace credits” if limits are reached.

Analysts said this reflects a shift in how enterprise AI usage is being monetized.

“This is not vendors admitting that flat seats failed. It is vendors admitting that one seat no longer describes one economic profile,” said Sanchit Vir Gogia, chief analyst at Greyhound Research. “The seat is the engine an enterprise licenses. The fuel is now billed separately.”

Gogia added that vendors are structuring pricing differently around that model. “OpenAI has kept a fixed seat and meters what sits above it. Microsoft has kept a $30 Copilot licence carrying no entitlement to its agentic layer at all. Google, meanwhile, has left its seat editions alone while switching agent meters on by published date.”

Greyhound Research pointed to the treatment of usage caps as a signal of how capacity is being priced.

“Most will read this as segmentation. Greyhound Research reads it as rationing, and the five-hour limit is the tell,” Gogia said. “On July 12, OpenAI removed that limit free of charge. At the end of July, it was restored. On August 10, it became a paid entitlement.”

CIOs face allocation and cost decisions

The higher-priced tier raises questions about how enterprises assign access, analysts said.

“Don’t start with people. Start with your billing data,” Chopra said. “If someone is regularly drawing down shared workspace credits, a fixed higher seat may cost less and forecast better than open-ended credit consumption.”

He said allocation based on hierarchy can lead to inefficiencies. “You end up paying premium rates for executives who open it twice a week while the analyst-blocked mid-close stays throttled.”

Gogia echoed that view. “Power user should describe a workload pattern, not a job title,” he said. “The right question is whose work becomes materially more valuable when Standard stops being enough.”

Credits positioned to drive metered usage

According to the statement, OpenAI is offering incentives, stating that eligible customers can receive “$100 worth of workspace credits (2,500 credits) for each Premium seat they add, up to 5 seats.”

Analysts said such credits are tied to usage-based pricing. “Worth being clear that this isn’t a discount. Credits are currency inside OpenAI’s metered layer,” Chopra said. “Free credits get a workspace comfortable for consuming metered features,” Gogia said the offer should be treated as a pilot incentive. “The promotion should fund the pilot. It should not shorten the runway,” Gogia said, adding that billing controls, including pooled credits and auto-recharge settings, require close oversight.

Why AI infrastructure needs a new operating model

The next AI infrastructure crisis may come from unmanaged inference capacity. For the past several years, the AI infrastructure conversation centered on one question: how do we get more compute?

That made sense. Enterprises needed GPUs, cloud capacity, foundation models and room to experiment. Compute became shorthand for AI readiness.

Production AI changes the operating discussion. Utilization, routing, latency, throughput, cost control, policy, privacy and governance now need to be managed together. A GPU that sits idle creates no business value. A model endpoint with unpredictable latency frustrates users. An inference stack that cannot be measured end-to-end becomes difficult to defend when usage grows and finance asks where the money is going.

CIOs need governed capacity.

Governed capacity means operating AI infrastructure as a production system rather than a collection of disconnected resources. They need to know how much useful output their infrastructure produces, where that output runs, why it runs there, what it costs, how it performs, what policy applies and whether the system can be controlled as demand changes.

Enterprises buy AI infrastructure to deliver answers, summaries, recommendations, software code, customer interactions, analysis, automation and agent workflows. Those outputs need to be reliable, measurable and affordable enough to keep running.

The pilot-era stack is reaching its limit

The first wave of enterprise AI rewarded speed. Teams bought GPUs, reserved cloud capacity, tested APIs, adopted open-source models and assembled whatever stack helped them move.

Infrastructure inefficiency then becomes a business issue.

The symptoms are familiar: more systems to manage, more vendors to coordinate, more integration work and less visibility into what drives cost and performance.

That creates friction across the organization. IT teams support AI workloads that behave differently from traditional enterprise applications. AI teams need speed, but often lack the infrastructure control to tune cost, latency, utilization and performance together. Finance teams want predictable unit economics, but the stack was assembled under pressure and is hard to measure end to end.

Most teams can now get access to models and compute. Fewer can show how each workload is performing, where it runs and what it costs.

Capacity needs control

Extra capacity can still leave teams with idle infrastructure, uneven latency and unclear unit costs.

The useful questions are operational. Can the organization see utilization across teams, tenants, models and infrastructure pools? Can it route workloads based on cost, latency, privacy, availability and service objectives? Can it measure cost per token, cost per inference, cost per user interaction or cost per business workflow?

Inference behavior changes constantly. Demand fluctuates. Longer contexts increase cost. Model choice affects latency and output quality. Utilization varies across workloads. A customer-facing assistant may prioritize response time. A batch workflow may prioritize throughput and cost.

A procurement-led AI strategy cannot manage that complexity on its own. CIOs need an operating model for production inference.

Enterprise Linux offers a useful analogy. Linux gave companies flexibility and attractive economics, but enterprises needed a trusted operating layer and support model before using it for business-critical systems. AI infrastructure is reaching a similar stage. The models, hardware and software components already exist. Many organizations now need a way to operate them consistently and economically in production.

Token economics is becoming a management discipline

The useful output of many AI systems is delivered through tokens. That makes token economics a practical operating metric.

Token volume needs context. A token that helps complete a task, answer a question or resolve a customer issue creates value. A token generated through poor routing, excess latency or an unnecessarily expensive model adds cost without improving the outcome.

How much useful output are we getting per dollar? How much per watt? How much per GPU? How much per workload? How much per unit of latency? How much per business outcome?

Manufacturing leaders do not only ask how many machines they own. They ask what those machines produce, how often they sit idle, how much waste they create, how much energy they consume and how efficiently raw materials become finished goods.

AI infrastructure needs the same operating discipline: utilization, throughput, reliability, cost control and visibility into what the infrastructure is producing.

Enterprises need usability and control

Serverless AI APIs are fast to start and easy for developers. They work well for many use cases. As usage grows, economics can become harder to control and visibility into infrastructure behavior is limited.

Self-managed infrastructure gives teams more control and can improve long-term economics for persistent workloads. It also adds operational burden. Teams have to manage deployment, scaling, routing, model serving, monitoring, reliability, performance tuning, security, isolation and utilization.

Enterprises want the simplicity of managed services without giving up visibility and control. Developers should be able to access AI services without managing the underlying stack. Infrastructure, security and finance teams still need to see placement, cost, latency, utilization, tenant policy, service levels and risk.

That is the role of an inference operating layer: turning fragmented infrastructure into governed, measurable capacity that teams can manage as demand changes.

Beyond procurement

The more successful an AI application becomes, the more inference it consumes. As inference grows, cost, latency, utilization and governance determine whether the application can scale.

AI can repeat the cloud-cost pattern many CIOs already know. A service begins as an innovation accelerator, usage expands across teams and the bill grows faster than governance. By the time the organization tries to regain control, the architecture, workflows and vendor dependencies are difficult to unwind.

GPUs remain essential. Models remain essential. Data remains essential. Production AI also needs an operating layer around those assets.

The next generation of AI leaders will ask a harder question:

How much useful intelligence can we produce from our infrastructure, at what cost, with what reliability, under what policy and under whose control?

The answer will determine whether AI becomes a controlled production capability or another expensive system the business struggles to explain.

This article is published as part of the Foundry Expert Contributor Network.
Want to join?

Data center backlash could slow CIOs’ AI plans

A growing backlash against building new data centers in the US may have huge cost implications for CIOs planning to expand their organizations’ AI initiatives.

Protests against building new data centers were organized in 42 states in mid-July, with participants concerned about new facilities driving up electricity and water costs and using large swaths of land.

As of mid-July, 10 states, including Florida, Georgia, and Virginia, had active data center construction moratoriums in place, and eight other states had pending legislation, according to datacenterbans.com.

In addition, as of May, 23 states had approved large-load tariffs that require data centers to pay the full infrastructure cost for their facilities, says Arif Gasilov, a partner in the natural resources and built environment division of sustainability advisory firm Gasilov Group.

IT leaders need to calculate the backlash into their planning for the compute and other IT infrastructure needs that new data centers would meet, he says.

“What this means for CIOs is that power cost assumptions built in 2023 are wrong in close to half the country,” Gasilov says. “A CIO planning an AI deployment that depends on colocation or cloud capacity in any of these states should be asking their provider what the rate structure looks like under the new tariffs and recalculating economics.”

In some cases, it may be possible to go smaller to avoid the moratoriums or tariffs on large data centers, but some state regulations target facilities close to each other as opposed to individual data centers, he notes.

Deployment challenges

If the backlash continues, IT leaders may need to rethink the way they deploy AI, says Chuck Girt, CTO at fiber-optic network provider FiberLight.

With fewer options for AI compute power, organizations would have less flexibility in where they deploy AI workloads, he suggests.

“I don’t think the rate of data center construction changes the direction AI is headed, but it could influence how organizations deploy and access AI at scale,” he says. “Most enterprises aren’t going to build this infrastructure themselves; they’re going to rely on cloud and data center environments to provide the compute AI requires.”

A lack of data center options could put many organizations in a bind, says Kevin Surace, CEO of biometric security vendor TokenCore.

“Compute capacity is becoming as strategically important as electricity, semiconductors, and network connectivity,” he says. “Fewer data centers mean less available capacity, reduced geographic redundancy, longer provisioning times, and greater dependence on a small number of cloud providers and locations.”

Organizations that have not secured capacity could find that their AI strategy is technically sound but physically impossible to execute on schedule, he suggests.

Surace, also an AI and green energy expert, is concerned that generalized fear about older data center designs is turning into blanket opposition to new construction. Modern facilities have cut down on the massive water use of older data centers, he notes, and some are using renewable energy generation. Nuclear power will become an electricity option soon, he adds.

Cost pressures rising

In the meantime, IT leaders should expect higher costs for compute and other IT infrastructure provided through data centers, Surace says.

“Demand for AI compute is accelerating, so constraining the supply of facilities, electricity and high-density capacity will place upward pressure on cloud pricing, colocation, accelerator access, and long-term capacity contracts,” he adds.

Organizations that have the capacity will should be able to protect themselves through multiyear agreements and dedicated infrastructure, he suggests. Smaller organizations, startups, and universities could face the greatest percentage increases and may simply be priced out of leading-edge AI capabilities, he adds.

Therefore, Surace advises CIOs to treat compute and energy as strategic supply-chain risks. Organizations should secure capacity as soon as they can, avoid dependence on one cloud or one geographic region, and use smaller and more efficient AI models where appropriate, he recommends.

He also suggests that CIOs ask data center providers several hard questions:

  • Where does the water come from?
  • Is the cooling loop closed?
  • Who pays for new grid infrastructure?
  • What percentage of power is generated onsite?
  • What environmental monitoring is publicly reported?

Data centers can mitigate some of the community concerns, he says. “Transparency and early community engagement are far less expensive than lawsuits, project cancellations, and moratoriums,” he adds.

Backlash against inefficiency

While protests are likely to continue, some don’t see the concerns about data centers as a condemnation of AI. Instead, the problem is with inefficient AI deployments, says Anurag Gurtu, cofounder and CEO of agentic AI platform provider Airrived.

“Enterprises don’t actually want more data centers; they want more intelligence per watt, per GPU, and per dollar,” he says. “The winners won’t be those with the biggest infrastructure footprint, but those extracting the most value from every unit of compute.”

Limitations on data centers will impact companies only if their AI strategies depend on nearly unlimited infrastructure, he adds.

“The next generation of AI will be constrained by compute, power, and economics,” Gurtu says. “Organizations that optimize models, deploy domain-specific AI, and leverage hybrid architectures will continue to innovate, while those relying solely on scaling hardware will face diminishing returns.”

While limited compute options could lead to higher prices, the solution is to focus on efficiency, he adds.

“Rising infrastructure costs also accelerate innovation in model optimization, inference efficiency, and intelligent orchestration,” Gurtu says. “History shows constraints often become the catalyst for the next wave of breakthroughs.”

❌