Visualização normal

Antes de ontemSecurity | CIO
  • ✇Security | CIO
  • AI inference is getting cheaper, but your agents are getting more expensive
    The good news is that large language model (LLM) token costs are coming down. The conundrum: The overall cost of AI workloads is going up. Gartner research predicts that, while token costs will fall by 95% by 2030, inference costs for agentic workflows will increase more than fivefold over the next two years. This is because AI app builders are using more, and often more expensive, tokens as LLMs get ever more complex. It is what Gartner calls the “inference paradox.”
     

AI inference is getting cheaper, but your agents are getting more expensive

17 de Agosto de 2026, 22:40

The good news is that large language model (LLM) token costs are coming down. The conundrum: The overall cost of AI workloads is going up.

Gartner research predicts that, while token costs will fall by 95% by 2030, inference costs for agentic workflows will increase more than fivefold over the next two years. This is because AI app builders are using more, and often more expensive, tokens as LLMs get ever more complex. It is what Gartner calls the “inference paradox.”

In other words, “the rate of innovation is outpacing the cost curve,” Gartner analysts Will Sommer and Sabine Zimmerhansl noted in their report. “The market is captured by a token-deflation illusion.” Buyers dangerously assume that as AI providers improve token economics, these savings will be reflected in their roadmaps. But, simply put, “they will not.”

‘Swarms’ of agents eating up tokens

There is no doubt that AI delivers massive value, and is often better, and faster, than people at many routine tasks, the analysts pointed out. Agents can also more quickly identify patterns across siloed systems. For instance, Gartner has seen customer success agents reduce response times by 99%.

But as AI evolves, token usage increases, and token value is variable, Sommer and Zimmerhansl noted. Advanced AI agents that can reason already cost up to 150x more on a single task than basic AI chatbots.

A simple chatbot must read and interpret a request and quickly deliver a “probabilistically reasonable” answer, but agents, as they become more sophisticated, must think, question, and adapt when something goes wrong, and they increasingly run continuously, and often invisibly, in the background.

“They need to be able to validate their results for accuracy without necessarily having a human in the loop,” the analysts wrote. “They need to talk to other agents.”

All of this demands more resources, and token consumption increases exponentially as “swarms” of increasingly autonomous agents trigger and call each other. This incurs a “massive inference tax” before a user even gets their result, they noted.

The hardware costs to train medium-sized agentic models with advanced reasoning capabilities is 2.5x greater than training simple, similarly-sized chatbots, they reported. Further, agent inference costs are 5x greater, and agents require 5x to 30x more tokens than a chatbot to handle equivalent tasks.

“Now consider how costs will balloon when running hundreds of agents that can perform dozens or hundreds of tasks each hour,” the analysts said. Costs continue to skyrocket as agents break many problems into several small tasks, call higher-order models, and require multimodal data.

“The volume of compute required for these capabilities is mind-bending,” the analysts noted.

To analyze the impacts of agentic systems, Gartner built a Tokenomics Model based on various scenarios of training and inference. These included various designs (number of layers, LLM-as-a-judge or as mixture-of-experts), technology improvements, hardware specifications, and various other cost considerations (data, infrastructure, energy, labor).

The firm ran 12 types of AI model with various capabilities, and found that basic workflows cost around $0.05 per inference token; summarization and knowledge retrieval cost roughly $0.10; more complex workflows cost around $0.30; and planning and learning cost roughly $0.40 per token. This means provider cost per token for planning and learning tasks is 8x to 10x that of basic workflows.

“Costs will inevitably escalate, and as they do, there is no guarantee that value will grow commensurately,” Sommer and Zimmerhansl contended. “ROI from each new generation of technology will be hard-earned.”

Being economical in the age of AI

Enterprises can be diligent and keep token costs in check by developing and maintaining complex multimodal systems, Gartner said. They will also need ways to measure ROI and improvement across completed workflows.

Orchestration will be the differentiator, and “inference tiering” will improve cost and performance, so enterprises should develop systems that route queries to the most cost‐efficient model and block agents from invoking frontier models by default for simpler tasks, Gartner advised. Adopting usage-based pricing is another important step; move from flat compute fees to tiered plans that scale based on need.

Enterprises can consider mandating continuous refresh cycles, Sommer and Zimmerhansl added. “Treat each model release like a ‘new car’ losing value on day one,” they wrote, and build in data fine-tuning and self-learning feedback loops.

Builders should also set minimum standards for AI execution: Define success thresholds and risk mitigation and compliance overhead up front. “Refuse to greenlight deployments until scenarios are stress-tested against token-price swings and compliance expenses,” the analysts emphasized.

Further, Gartner also advises embedding value-per-outcome into product planning. This could mean requiring every AI feature to forecast and track its spend against a “clear outcome metric,” such as tasks automated or cases successfully closed. This can help identify the low-performing workflows that require improvement or deprecation.

Still, ROI is “eminently possible,” but it requires significant effort across complex workflows, the analysts noted; enterprises can’t simply rely on traditional systems and workflows. “Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems,” they pointed out.

This article originally appeared on Computerworld.

  • ✇Security | CIO
  • Server prices to rise by up to 87% at OVHcloud
    OVH is increasing the prices of its servers, some by as much as 87%, for both new and existing customers, blaming AI’s insatiable demand driving the rising cost of the RAM and storage it uses in its data centers. The European cloud operator specializes in low-cost bare metal and public cloud offerings. CIOs will be familiar with the balancing act OVH has had to perform over the last year. In a Monday post explaining the upcoming increases, OVH chairman Octave Klaba w
     

Server prices to rise by up to 87% at OVHcloud

11 de Agosto de 2026, 14:47

OVH is increasing the prices of its servers, some by as much as 87%, for both new and existing customers, blaming AI’s insatiable demand driving the rising cost of the RAM and storage it uses in its data centers.

The European cloud operator specializes in low-cost bare metal and public cloud offerings.

CIOs will be familiar with the balancing act OVH has had to perform over the last year. In a Monday post explaining the upcoming increases, OVH chairman Octave Klaba wrote on X,  “We have to place the right volume of orders, month by month, over 12 months, with no guarantee of the purchase price and without knowing what will be the real demand from our customers.”

Still, he added, “even though our prices are increasing, we remain the cheapest on the market for bare metal and public cloud; where before we could be 3x cheaper, we will be 2x cheaper (if our competitors don’t increase their prices).”

The increases will hurt hard-core gamers hardest, with the cost of the company’s most recent gaming servers rising 87%. (Older gaming instances are unaffected.)

High Grade, high price

But enterprises will also feel the pain from climbing component costs: OVH’s latest High Grade bare metal servers, with up to 2 x 96 cores of AMD Epyc 9005 series processors, 36 hard disks per server, and high-density cooling systems, will go up in price by 59%; older models built to the 2024 spec will go up 26%.

Lower-performance servers will also see increases of 40%-49% for the most recent models, and 26%-37% for older models.

The new prices take effect from Sept. 1 for new orders, and from Oct. 1 for renewals.

It’s not just baseline server prices that are increasing; optional additional memory and storage are going up in price too. OVH already increased the cost of these extras for new server orders as of July 1, with RAM prices rising 127% and disks 89%. From Oct. 1, renewals will be affected too, with the price of additional RAM in the latest servers rising by 40%, and that of larger disks by 15%. For servers built to 2024 specs, the increases will be 20% and 10% respectively.

Existing customers can lock in current prices for servers already in production for up to four years if they pay in advance by Oct. 1, Klaba wrote. Existing commitments will not be affected by the increases until they are due for renewal.

Small instances, big increases

The price rises are more nuanced when it comes to public cloud systems. In future, OVH will break out storage and IP address rental costs separately, and will allow customers to mix and match storage capacity and compute.

“In appearance, hourly compute cost won’t change,” Klaba wrote. “On the other hand, low-latency Block Storage and IPv4 addresses, previously included in our Gen3 instances (B3, C3, R3) will appear as two separately billed line items on Oct. 1.”

The result is price increases of as little as 1.4% for the most powerful instances, or as much as 21.9% for smaller instances, he said.

OVH will continue to offer a 15% discount for a commitment of one year, or 30% for three years, he said, but will no longer offer discounts for shorter terms.

This article originally appeared on NetworkWorld.

❌
❌