Visualização de leitura

AI inference is getting cheaper, but your agents are getting more expensive

The good news is that large language model (LLM) token costs are coming down. The conundrum: The overall cost of AI workloads is going up.

Gartner research predicts that, while token costs will fall by 95% by 2030, inference costs for agentic workflows will increase more than fivefold over the next two years. This is because AI app builders are using more, and often more expensive, tokens as LLMs get ever more complex. It is what Gartner calls the “inference paradox.”

In other words, “the rate of innovation is outpacing the cost curve,” Gartner analysts Will Sommer and Sabine Zimmerhansl noted in their report. “The market is captured by a token-deflation illusion.” Buyers dangerously assume that as AI providers improve token economics, these savings will be reflected in their roadmaps. But, simply put, “they will not.”

‘Swarms’ of agents eating up tokens

There is no doubt that AI delivers massive value, and is often better, and faster, than people at many routine tasks, the analysts pointed out. Agents can also more quickly identify patterns across siloed systems. For instance, Gartner has seen customer success agents reduce response times by 99%.

But as AI evolves, token usage increases, and token value is variable, Sommer and Zimmerhansl noted. Advanced AI agents that can reason already cost up to 150x more on a single task than basic AI chatbots.

A simple chatbot must read and interpret a request and quickly deliver a “probabilistically reasonable” answer, but agents, as they become more sophisticated, must think, question, and adapt when something goes wrong, and they increasingly run continuously, and often invisibly, in the background.

“They need to be able to validate their results for accuracy without necessarily having a human in the loop,” the analysts wrote. “They need to talk to other agents.”

All of this demands more resources, and token consumption increases exponentially as “swarms” of increasingly autonomous agents trigger and call each other. This incurs a “massive inference tax” before a user even gets their result, they noted.

The hardware costs to train medium-sized agentic models with advanced reasoning capabilities is 2.5x greater than training simple, similarly-sized chatbots, they reported. Further, agent inference costs are 5x greater, and agents require 5x to 30x more tokens than a chatbot to handle equivalent tasks.

“Now consider how costs will balloon when running hundreds of agents that can perform dozens or hundreds of tasks each hour,” the analysts said. Costs continue to skyrocket as agents break many problems into several small tasks, call higher-order models, and require multimodal data.

“The volume of compute required for these capabilities is mind-bending,” the analysts noted.

To analyze the impacts of agentic systems, Gartner built a Tokenomics Model based on various scenarios of training and inference. These included various designs (number of layers, LLM-as-a-judge or as mixture-of-experts), technology improvements, hardware specifications, and various other cost considerations (data, infrastructure, energy, labor).

The firm ran 12 types of AI model with various capabilities, and found that basic workflows cost around $0.05 per inference token; summarization and knowledge retrieval cost roughly $0.10; more complex workflows cost around $0.30; and planning and learning cost roughly $0.40 per token. This means provider cost per token for planning and learning tasks is 8x to 10x that of basic workflows.

“Costs will inevitably escalate, and as they do, there is no guarantee that value will grow commensurately,” Sommer and Zimmerhansl contended. “ROI from each new generation of technology will be hard-earned.”

Being economical in the age of AI

Enterprises can be diligent and keep token costs in check by developing and maintaining complex multimodal systems, Gartner said. They will also need ways to measure ROI and improvement across completed workflows.

Orchestration will be the differentiator, and “inference tiering” will improve cost and performance, so enterprises should develop systems that route queries to the most cost‐efficient model and block agents from invoking frontier models by default for simpler tasks, Gartner advised. Adopting usage-based pricing is another important step; move from flat compute fees to tiered plans that scale based on need.

Enterprises can consider mandating continuous refresh cycles, Sommer and Zimmerhansl added. “Treat each model release like a ‘new car’ losing value on day one,” they wrote, and build in data fine-tuning and self-learning feedback loops.

Builders should also set minimum standards for AI execution: Define success thresholds and risk mitigation and compliance overhead up front. “Refuse to greenlight deployments until scenarios are stress-tested against token-price swings and compliance expenses,” the analysts emphasized.

Further, Gartner also advises embedding value-per-outcome into product planning. This could mean requiring every AI feature to forecast and track its spend against a “clear outcome metric,” such as tasks automated or cases successfully closed. This can help identify the low-performing workflows that require improvement or deprecation.

Still, ROI is “eminently possible,” but it requires significant effort across complex workflows, the analysts noted; enterprises can’t simply rely on traditional systems and workflows. “Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems,” they pointed out.

This article originally appeared on Computerworld.

Server prices to rise by up to 87% at OVHcloud

OVH is increasing the prices of its servers, some by as much as 87%, for both new and existing customers, blaming AI’s insatiable demand driving the rising cost of the RAM and storage it uses in its data centers.

The European cloud operator specializes in low-cost bare metal and public cloud offerings.

CIOs will be familiar with the balancing act OVH has had to perform over the last year. In a Monday post explaining the upcoming increases, OVH chairman Octave Klaba wrote on X,  “We have to place the right volume of orders, month by month, over 12 months, with no guarantee of the purchase price and without knowing what will be the real demand from our customers.”

Still, he added, “even though our prices are increasing, we remain the cheapest on the market for bare metal and public cloud; where before we could be 3x cheaper, we will be 2x cheaper (if our competitors don’t increase their prices).”

The increases will hurt hard-core gamers hardest, with the cost of the company’s most recent gaming servers rising 87%. (Older gaming instances are unaffected.)

High Grade, high price

But enterprises will also feel the pain from climbing component costs: OVH’s latest High Grade bare metal servers, with up to 2 x 96 cores of AMD Epyc 9005 series processors, 36 hard disks per server, and high-density cooling systems, will go up in price by 59%; older models built to the 2024 spec will go up 26%.

Lower-performance servers will also see increases of 40%-49% for the most recent models, and 26%-37% for older models.

The new prices take effect from Sept. 1 for new orders, and from Oct. 1 for renewals.

It’s not just baseline server prices that are increasing; optional additional memory and storage are going up in price too. OVH already increased the cost of these extras for new server orders as of July 1, with RAM prices rising 127% and disks 89%. From Oct. 1, renewals will be affected too, with the price of additional RAM in the latest servers rising by 40%, and that of larger disks by 15%. For servers built to 2024 specs, the increases will be 20% and 10% respectively.

Existing customers can lock in current prices for servers already in production for up to four years if they pay in advance by Oct. 1, Klaba wrote. Existing commitments will not be affected by the increases until they are due for renewal.

Small instances, big increases

The price rises are more nuanced when it comes to public cloud systems. In future, OVH will break out storage and IP address rental costs separately, and will allow customers to mix and match storage capacity and compute.

“In appearance, hourly compute cost won’t change,” Klaba wrote. “On the other hand, low-latency Block Storage and IPv4 addresses, previously included in our Gen3 instances (B3, C3, R3) will appear as two separately billed line items on Oct. 1.”

The result is price increases of as little as 1.4% for the most powerful instances, or as much as 21.9% for smaller instances, he said.

OVH will continue to offer a 15% discount for a commitment of one year, or 30% for three years, he said, but will no longer offer discounts for shorter terms.

This article originally appeared on NetworkWorld.

How to fake a data trail (and maybe lower prices) (Lock and Code S07E16)

It may sound entirely bizarre but the prices you once paid for hotels, educational classes, or staplers could have all been higher because you used a Mac computer, lived in a certain zip code, or lacked an Office Depot in your neighborhood.

No, really.

In 2012, The Wall Street Journal reported that the travel booking site Orbitz showed Mac users pricier hotel options than PC users, because the company had determined that Mac users spend, on average, 30% more a night on hotels. That same year, The Wall Street Journal (once again) reported that Staples.com showed higher prices to visitors who lived farther away from a competitor like Office Depot. And in 2015, the reporting outfit ProPublica revealed that customers in certain zip codes were shown higher prices for college test prep courses offered by The Princeton Review.

As that investigation found, if customers:

“type some zip codes into the company’s website, they are offered The Princeton Review’s premier course for as little as $6,600. For other zip codes, the same course cost as much as $8,400. One unexpected effect of the company’s geographic approach to pricing is that Asians are almost twice as likely to be offered a higher price than non-Asians.”

This is surveillance pricing put into action.

Under surveillance pricing, companies collect as much data as possible about consumers so that they can alter the literal prices those consumers pay for the exact same goods as everyone else. It is reportedly what caused some customers to see higher prices for televisions in the Target app when those customers were physically located in a Target parking lot. It is also allegedly why Home Depot customers in wealthy neighborhoods oddly paid less. And it is what Delta Airlines walked away from after public backlash.

The near-omnipresence of surveillance pricing is also why so many videos can be found online today that claim that minor alterations to a person’s data trail—like changing an IP address using a VPN or shopping for airline tickets on a public library’s computer—can lead to lower prices online.

The proof behind these claims, however, is harder to test.

Thankfully, one person has already tried.

Video journalist Chris Parr, known on YouTube as Chris the Producer, ran a wild experiment into whether he could “stress-test” surveillance pricing. Far beyond changing his IP address or making online purchases from different locations, Parr started from scratch. By first registering an LLC in the state of Wyoming, Parr granted that LLC both a credit card and a phone, effectively creating a brand new consumer persona to be tracked. But creating a realistic data trail for his LLC would require a little extra help—help that Parr received from an actor he hired for the part.

Today, on the Lock and Code podcast with host David Ruiz, we speak with Parr about his experiment into surveillance pricing, including a high-wire drone act to purchase a White Castle Crave Case in the air space above his home state’s wealthiest neighborhood:

“To the data collectors, they don’t know that this phone is floating in the air, like 200 feet in the air. They just see a geolocation on it.”

Tune in today to listen to the full conversation.

Show notes and credits:

Intro Music: “Spellbound” by Kevin MacLeod (incompetech.com)
Licensed under Creative Commons: By Attribution 4.0 License
http://creativecommons.org/licenses/by/4.0/
Outro Music: “Good God” by Wowa (unminus.com)


Listen up—Malwarebytes doesn’t just talk cybersecurity, we provide it.

Protect yourself from online attacks that threaten your identity, your files, your system, and your financial well-being with our exclusive offer for Malwarebytes Premium for Lock and Code listeners.

❌