2 Sources
[1]
Agentic AI costs set to balloon fivefold by 2028
The cost of agentic AI workflows is forecast to increase more than fivefold by the end of 2028 as users adopt more complex applications of the technology. As Nvidia and other tech giants push inference and agentic AI as the next stage of the AI wave, Gartner warns that the cost of implementing these systems will rise even as foundation models become cheaper. The analyst firm has cast its eye over the nascent world of AI agents - systems designed to act independently in pursuit of a goal - and sees multiple challenges ahead. Leaving aside the substantial security concerns, these software agents are considerably more complex than chatbots. Gartner believes falling model prices are tempting users to build more complex workflows, whose greater token consumption can outweigh those savings and drive up overall inference costs. In other words, tokens are becoming more cost-efficient, but those savings are not keeping pace with the rising cost of more advanced AI capabilities. The rate of innovation is outpacing the cost curve, Gartner claims. "The harsh economics of the inference paradox are exemplified by the differences between a simple chatbot and an AI agent," says Gartner senior director analyst Will Sommer. "Where a simple chatbot must read and interpret a query and quickly respond with a probabilistic reasonable answer, an AI agent must constantly reason, negotiate, and question itself," he explains. Those processes add up: routing a task to an agentic reasoning model increases inference costs at least fivefold, and potentially by much more as the task becomes more complex. Securing a return on investment from such advanced AI tools therefore demands either much greater returns than basic models provide or better optimization of inference, routing, and orchestration, Gartner warns. This could mean assigning each task to the most cost-efficient model capable of handling it. The move by some AI providers from flat-rate subscriptions to usage-based billing hasn't helped, as The Register reported last month. Token-heavy workflows can produce runaway costs under the new model. Perhaps it is no wonder Gartner predicted earlier this year that 40 percent of organizations would demote or decommission AI agents because of problems with the heavily hyped technology. The analyst biz also cheerily forecast last year that at least half of all generative AI projects would blow their budgets because of poor architectural choices and a lack of expertise, while most attempts to build custom models would be abandoned. ®
[2]
AI Inference Costs per Agentic Workflow to Surge Over 5X by 2028, Gartner Predicts
Product Leaders Are Facing the Inference Paradox: Better Unit Economics Is Escalating the Overall Cost of AI Without Providing a Clear Pathway to Commensurate and Predictable Value AI inference costs per agentic workflow will increase more than fivefold through 2028 according to Gartner, Inc., a business and technology insights company. As AI products evolve from assistive features to multistep execution, product leaders face a new margin challenge with falling model prices subsidizing more complex workflows and escalating total AI costs. As a result, inference cost management has become a top priority for product leaders. "Product leaders cannot rely on more efficient token economics to rationalize AI costs," said Will Sommer, Sr. Director Analyst. "Each successive generation of AI capability will necessitate more, and often more expensive, tokens. There is no reliable, economical one-size-fits-all model on the horizon. Producing competitive AI products will require developing and maintaining complex multimodel ecosystems." Gartner has identified three fundamental trends driving token economics: * Foundational model cost economics are rapidly improving. * Improved AI efficiency is unlocking the deployment of more powerful, and more expensive, models to enable higher-value, more sophisticated AI applications. * More sophisticated AI workflows use far more tokens than simple chatbot interactions, driving higher overall inference costs. These dynamics mean that tokens are becoming more cost-efficient, but not as quickly as AI capabilities and the costs associated with those capabilities are increasing (see Figure 1). The rate of innovation is outpacing the cost curve. Figure 1: AI Complexity Outpaces Falling Token Costs Source: Gartner (August 2026) This is the Inference Paradox, defined as better unit economics escalating the overall cost of AI without providing a clear pathway to commensurate and predictable value. "The harsh economics of the Inference Paradox are exemplified by the differences between a simple chatbot and an AI agent," said Sommer. "Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate, and question itself." All of these responsibilities add up. Compared to a basic chatbot interaction, routing a task to an agentic reasoning model increases provider inference costs by at least five times, and often much more as task complexity grows. Ensuring ROI from advanced AI, like reasoning agents, demands exponentially higher returns relative to basic models, or highly optimized inference-tiering, routing and orchestration to calibrate complex tasks relative to more cost-efficient intelligence. Both of these outcomes are eminently possible but will require significant effort across complex workflows. "Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems," said Sommer.
Share
Copy Link
Gartner forecasts AI inference costs per agentic workflow will surge more than fivefold through 2028. While foundation model prices fall, the complexity of agentic AI systems and their massive token consumption are driving overall costs upward, creating what analysts call the Inference Paradox.
The cost of implementing agentic AI systems is set to explode, with Gartner predicting AI inference costs per agentic workflow will increase more than fivefold by the end of 2028
1
2
. This alarming projection comes as tech giants like Nvidia push inference and agentic AI as the next frontier, but the economics tell a different story. While foundation model prices continue to fall, the complexity of multistep agentic workflows is driving token consumption to unprecedented levels, creating what Gartner calls the Inference Paradox—better unit economics that paradoxically escalate overall AI costs without guaranteeing proportional value2
.Will Sommer, Gartner Senior Director Analyst, explains the core challenge facing organizations: "Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens"
2
. The Inference Paradox emerges from three fundamental trends reshaping the landscape. First, foundational model cost economics are rapidly improving. Second, this improved efficiency unlocks deployment of more powerful and expensive models for sophisticated applications. Third, these advanced workflows consume far more tokens than simple chatbots, driving higher overall inference costs2
. The rate of innovation is outpacing the cost curve, meaning tokens become more cost-efficient but not quickly enough to offset the escalating demands of advanced AI capabilities1
.
Source: CXOToday
The stark difference between simple chatbots and AI agents illustrates why costs are ballooning. Sommer notes, "Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate, and question itself"
1
2
. These processes add up quickly. Routing a task to an agentic reasoning model increases provider inference costs by at least five times compared to basic chatbot interactions, and potentially much more as task complexity grows1
2
. AI agents operate independently in pursuit of goals, requiring substantially more computational resources than their simpler predecessors1
.The shift from flat-rate subscriptions to usage-based billing models by AI providers has intensified cost concerns
1
. Token-heavy workflows can produce runaway costs under these new pricing structures, making budget predictability a major challenge for organizations deploying agentic AI systems. Gartner warns that falling model prices are tempting users to build more complex workflows, but the greater token consumption outweighs those savings and drives up overall costs1
.Related Stories
Securing ROI from advanced AI tools demands either exponentially higher returns than basic models provide or highly optimized inference-tiering, routing, and orchestration strategies
2
. Organizations must calibrate complex tasks relative to more cost-efficient intelligence by assigning each task to the most cost-efficient model capable of handling it1
. Sommer warns that "defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems"2
. Producing competitive AI products will require developing and maintaining complex multimodel ecosystems rather than relying on a single one-size-fits-all approach2
.Gartner's forecast carries significant implications for the future of AI adoption. The firm predicted earlier this year that 40 percent of organizations would demote or decommission AI agents due to problems with the heavily hyped technology
1
. Additionally, Gartner forecasted that at least half of all generative AI projects would blow their budgets because of poor architectural choices and lack of expertise, while most attempts to build custom models would be abandoned1
. Beyond cost concerns, agentic AI systems present substantial security challenges that organizations must address1
. Watch for organizations to prioritize inference cost management as a critical capability, focusing on sophisticated orchestration strategies that balance capability with efficiency.
Source: The Register
Summarized by
Navi
[1]
17 Jun 2026•Business and Economy

26 May 2026•Business and Economy

28 Jul 2026•Business and Economy

1
Technology

2
Technology

3
Policy and Regulation
