6 Sources
[1]
We're asking the wrong question about the cost of enterprise AI
Enterprise AI's real cost battle: tokens vs infrastructure economics Enterprise AI is reaching an important economic turning point. For the past two years, organizations have largely evaluated AI through the lens of token pricing and model capability. As AI moves from experimentation into business-critical operations, that approach is becoming increasingly incomplete. The question is no longer simply what each token costs, but what it costs to deliver AI capability that is affordable, sustainable and commercially predictable at enterprise scale. Every AI interaction ultimately depends upon physical IT infrastructure, consuming compute, memory, networking, electricity and cooling regardless of how those costs are presented to the customer. Understanding the economics of that infrastructure is becoming just as important as understanding the capabilities of the models themselves. Organizations that focus only on the price of a token risk overlooking the factors that will ultimately determine the long-term cost, resilience and sustainability of enterprise AI. Why the price on the invoice isn't the whole story Organizations naturally focus on the invoice because it represents the visible cost of AI. Consumption-based pricing appears straightforward, transparent and easy to compare. Yet every token reflects far more than access to a language model. Behind every AI interaction sits physical infrastructure consuming compute, memory, networking, electricity and cooling. Those infrastructure costs remain largely invisible to the customer despite having a direct influence on the economics of enterprise AI. As AI moves from isolated pilots into production workloads, infrastructure decisions are repeated across millions of inference requests, making their long-term commercial impact increasingly significant. Understanding enterprise AI therefore requires organizations to look beyond the invoice. The efficiency, resilience and operating characteristics of the infrastructure delivering AI capability increasingly determine what AI will cost over its operational lifetime. Infrastructure efficiency is becoming a competitive advantage The architecture underpinning an AI platform has a significant influence on its economics over time. Purpose-built inference infrastructure can significantly improve energy efficiency compared with architectures optimized primarily for AI training workloads, often without requiring complex liquid cooling. Those efficiency gains have important commercial consequences. Lower energy demand reduces cooling requirements, simplifies facility design and lowers operating costs throughout the lifetime of the infrastructure. At enterprise scale, even relatively small efficiency improvements become commercially significant when repeated across millions of inference requests. Consumption-based, token-metered AI remains an appropriate deployment model for organizations with variable or exploratory workloads. As AI becomes embedded within everyday business operations, however, many organizations are finding that continually increasing token consumption creates an operational cost model that becomes progressively harder to forecast and control. Token volumes measure the level of AI activity, but they do not measure the business value created by that activity. Enterprise leaders are therefore becoming increasingly focused not simply on the cost of consuming AI, but on the long-term economics of delivering AI capability in a commercially sustainable way. Dedicated inference infrastructure represents a different economic model. Rather than paying for every interaction, organizations invest in AI capability with predictable operating costs, greater control over performance, data location and operational resilience. The discussion therefore shifts from purchasing tokens to building sustainable AI capability. Where that infrastructure is combined with on-site renewable generation and long- duration energy storage, organizations can further improve cost predictability while strengthening operational resilience and supporting long-term sustainability objectives. A broader conversation about enterprise AI economics Enterprise AI is entering a more mature phase of adoption. The discussion is no longer centered solely on model capability or the cost of individual tokens. Increasingly, organizations are asking how AI can create measurable business value while remaining commercially sustainable over the long term. That changes the conversation in the boardroom. Success is no longer measured simply by the volume of AI consumed, but by the outcomes it delivers. Token usage may indicate the level of AI activity, but it does not measure the value created for the organization. The focus therefore shifts towards deploying AI capability that delivers predictable commercial returns, operational resilience and strategic advantage. As organizations become increasingly dependent upon AI, they must also recognize that the underlying models are not static. Foundation models continue to evolve through incremental updates and refinements, many of which may be difficult for users to detect but can nonetheless influence behavior and outputs. Enterprise AI therefore requires governance that extends beyond monitoring consumption. Organizations need confidence that they understand not only what AI costs to operate, but also how the capability itself is changing over time and what those changes mean for business performance, compliance and risk. The ability to govern both the economics and the evolution of AI is becoming a strategic capability in its own right. Infrastructure remains the foundation that enables those outcomes, but it is no longer the destination of the discussion. The real objective is to create AI capability that delivers measurable business value, remains commercially sustainable and can be governed with confidence as technologies continue to evolve. Organizations that understand the relationship between infrastructure, operating economics, governance and business value will be better placed to realize the long-term benefits of enterprise AI. We've reviewed, rated, and ranked the best productivity tools. This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today. The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit
[2]
'Once the right balance between cloud and local AI is found, organisations will find the sweet spot between cost, performance and security': The future of AI strategy and how businesses can get the best results
Businesses of all shapes and sizes are adopting AI to improve productivity and efficiency, but where they should be seeing improvements from this strategy, instead they're seeing rising token costs, struggles with integrating tools, and more security risks. As with all new technologies, adoption comes first and the procedures and governance are a few steps behind. Experiments are taking place in almost every industry, and the lessons learned will help guide other businesses into successful adoption and AI maturity. But employees fear replacement and sometimes spend more time questioning the results of their AI assisted work. In the worst circumstances more work is created in trying to ensure employees trust the technology, and solving the new, unseen challenges that come with adapting an AI strategy. AI costs and challenges in the road ahead Rising token costs are one of the biggest challenges businesses face. Without clear ways to measure how token costs reflect performance, it's very difficult to assess if AI spend is actually offering any performance benefits. This is especially true when new, more powerful models are being released - and staying ahead of the competition means the accompanying, ever-increasing costs is the price of doing business. But while employees may have just finished their training, or setting up a new workflow for one AI model, introducing the next can increase complexity and harm any new productivity gains. There is therefore a balance to be struck between AI integration, its associated costs, and the productivity gains employees see. Rampant spending and reckless adoption can turn an AI strategy from a business-boosting asset into a stress-inducing, trust-eroding liability. Ruth Patterson, Managing Director, HP, UK & Ireland says that the businesses seeing the most success during this technological revolution aren't necessarily integrating it at every turn. Instead, they're "applying it to practical workflows in a way that protects data and delivers real results." I spoke to Patterson to understand the challenges businesses face in delivering an AI strategy that shows real results, and how organisations can tackle the challenges that come with adoption AI. * How are enterprise attitudes towards AI token use changing in 2026? Businesses are waking up to a simple but overlooked truth about AI: the more they use, the more it costs. As AI has become part of everyday work, the financial cost of millions of interactions, charged at a token level, has become much more visible - and not just to the IT department. Finance leaders are watching token usage fill a sizeable chunk of their balance sheets, rightly prompting much closer scrutiny of where workloads are processed and whether every task needs to be sent to the cloud. As the conversation moves from AI experimentation to value, leaders must now focus on building an AI strategy that's both commercially sustainable and operationally efficient. That means moving away from AI for AI's sake to a more tailored approach that balances cost, security and performance. * What are businesses prioritising as they move beyond initial AI experimentation? Where have businesses seen the greatest gains? The biggest gains so far have come from AI taking repetitive tasks off people's desks. You will have heard this said a lot, but the benefits are real. Whether it's summarising meetings, drafting documents, searching internal knowledge or helping employees find information more quickly, AI really is giving people back time to focus on work that requires judgement and creativity. Employees increasingly recognise that the future of the workplace is AI-enabled, whereby their own skills augmented by digital solutions. To succeed, organisations must offer access to the right technology and create environments where people feel empowered to experiment with AI. However, as more organisations begin to move beyond this experimentation phase, they are starting to become much more disciplined. They're asking whether AI is secure, whether employees are using approved tools and whether the technology is genuinely improving productivity rather than simply adding another application to the estate. The businesses making the fastest progress in this environment aren't necessarily using the most AI. They're the ones applying it to practical workflows in a way that protects data and delivers real results. * What are the biggest challenges organisations face when deploying AI at scale? Organisations must juggle several competing priorities as they scale deployment: how to protect sensitive data, manage operational costs, maintain performance and ensure employees trust the technology they're using. The trust question is particularly important as AI moves further into everyday use. Employees need confidence that the tools are reliable and approved, while organisations need confidence that data is protected and whether the cost model is right. Getting IT infrastructure tuned correctly is key to making this work at scale. One of the most important decisions that organisations need to make is which AI workloads belong in the cloud, and which make more sense running on the device. That is ultimately a business decision - as opposed to purely a technology or IT decision - because it affects performance, cost and security. * How does on-device AI slot into an organisation's overall AI strategy? I don't see cloud AI and on-device AI as competing approaches - I see them as complementary. Cloud services will remain essential for large-scale models and complex reasoning. But not every AI task needs to leave the device. For everyday activities like summarisation, transcription or content creation, running AI locally can reduce dependence on cloud infrastructure and give organisations greater control over sensitive information. AI only creates value when it becomes part of everyday workflows. By adopting on-device AI-solutions, employees can reduce latency and increase productivity with the peace of mind that their data is secure. This offers an easy route-in for employees beginning to implement AI into their day-to-day workflows. The right approach to enterprise AI strategy is pretty simple: place workloads where they make the most sense. Because once the right balance between cloud and local AI is found, organisations will find the sweet spot between cost, performance and security. * How are business leaders evaluating the success of AI investments around productivity? The conversation has moved beyond adoption metrics. Now leaders want to understand whether AI is creating measurable improvements in business performance. They're looking at time saved, faster decision-making, improvements to workflow efficiency and whether employees can spend more time on higher-value work. Technology leaders are also beginning to examine the broader economics of AI. Productivity gains need to be considered alongside operational expenditure, infrastructure requirements and long-term scalability. Ultimately, AI should reduce friction. If employees can complete work more efficiently while organisations maintain control over cost and governance, that's where the real return on investment begins to emerge. AI-enabled digital solutions now offer benefits such as persona-based device optimisation or integrated sentiment analysis to measure employee satisfaction with digital tools. These innovations truly redefine experience management and facilitate higher employee satisfaction and productivity. It also enables leaders to stay closer to their employees and gauge whether AI is improving their experience or simply adding another layer of complexity. * How do you expect enterprise AI infrastructure to evolve over the next few years? We will continue to see significant investment in AI infrastructure. But the conversation needs to go beyond capacity, because capacity doesn't create business value alone. The next phase will be about ensuring infrastructure enables AI workloads to run where they are most effective. That means treating the endpoint as part of the AI infrastructure and using it to process selected tasks locally, alongside cloud services and edge computing. At the same time, we can expect to see greater emphasis put on AI infrastructure that supports governance, security and efficiency alongside raw compute capacity. Together, these infrastructure shifts will start to deliver the practical value that employees and businesses expect from AI but aren't currently seeing. * What trends are you watching most closely in enterprise AI right now? Three stand out for me: First, the move towards hybrid AI architectures. Organisations are becoming much more deliberate about deciding which workloads should run in the cloud and which belong on the device. Second, the growing focus on the economics of AI. As adoption scales, businesses are paying much closer attention to operational expenditure, infrastructure efficiency and the total cost of AI deployment. Finally, we're seeing the role of the endpoint evolve. For years, the PC has been framed as an access device only. Today, it has become an intelligent computing platform, capable of running AI workloads securely and efficiently. Businesses don't need a big AI budget to get all three right. They just need to spend time working out, task by task, where intelligence really belongs. Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
[3]
Hidden costs are enterprise AI's next challenge
The Fast Company Executive Board is a private, fee-based network of influential leaders, experts, executives, and entrepreneurs who share their insights with our audience. Actual usage dictates today's AI budgets. What starts with a handful of a company's users experimenting with a new tool can quickly evolve into dozens of workflows, hundreds of employees, and thousands of prompts every day. And as AI usage grows, the bill grows with it. More visibility gives finance, revenue, and technology leaders a much better understanding of AI resource use and fast-growing costs. Visibility also introduces tough questions: Which workflows are creating value? Which are just consuming compute? Where is AI activity misaligned with business impact? As the conversation shifts to economics, leaders want to understand what they're getting from AI credits consumed and whether investments produce meaningful business outcomes. While AI should increase business value, if rising token consumption outpaces business impact, you may no longer be able to justify the investment. Why AI costs scale faster than teams expect You ask a question, and a well-reasoned answer appears a few seconds later. Even better, an agent has an answer before you prompt it. On the surface, the interaction feels simple. Underneath, a great deal happens before that answer reaches you. The agent loads instructions, retrieves information, and assembles context. Agents may check multiple sources, rerun tasks, or process the same information more than once. Each step consumes tokens, or the small units of text an AI model reads and writes. Tokens are what you actually pay for, so the more steps, the higher the bill. Since January, mentions of "token costs" in corporate documents such as earnings transcripts have nearly tripled, according to our data. Many organizations are learning this the hard way, with nearly 8 in 10 IT leaders reporting they were surprised by charges due to AI models or how much the models were used. Most of an AI's work happens out of sight, so costs are only reconciled when the invoice arrives. Most AI platforms require organizations to piece together the underlying information layer themselves. This relies on different data sources, business applications, and access permissions across multiple tools. As a result, companies must ensure the AI pulls the right information, preserves the right context, and produces answers that can be verified if challenged. In high-stakes environments, every additional connection introduces more complexity, more opportunities for error, and ultimately more cost. The cost of accuracy Many organizations also encounter unexpected costs. An AI platform may seem inexpensive when evaluated on subscription price alone. A forecast built around a handful of use cases changes once AI becomes part of daily workflows. As a result, consumption and cost can accelerate in ways that are difficult to predict. In complex environments, inefficient architecture can lead the model to undertake extra work just to produce a reliable answer, when the real question is "Can I trust and defend what it tells me?" A workflow that appears inexpensive during a pilot can look very different as teams add more data sources, load more context into prompts, and introduce additional validation steps. A reliable answer requires extra tokens with every additional search, trip between systems, and step needed to refine the answer. Why quality becomes an economic issue Enterprise AI's deeply embedded assumption is that if you connect a model to enough information, intelligence will follow. In practice, that assumption breaks down more often than many realize. A system can connect to dozens of sources and still miss the most important piece of evidence. It can retrieve information without preserving context, and surface facts without understanding why those facts matter. The answer may look polished, cited, and complete, while still missing the critical signal necessary to make a confident decision. The challenge isn't always visible. Weak answers often look strong. They read well and may even include citations. A prompt 20% less efficient can drive costs 200% higher as models spend more tokens compensating for weak retrieval and noisy context. These small inefficiencies add up quickly. Organizational growth becomes harder to sustain when AI consumes more resources without creating proportionally more business value. That's when rising token costs stop being a technology issue and start becoming a business liability. Token efficiency starts with architecture In less than a year, a technical discussion about tokens has become a boardroom priority, showing up in budget reviews, operating strategies, and conversations about return on investment. Every architectural decision has downstream consequences. Content quality, retrieval design, context preservation, and how information moves through the system all influence what AI must do before producing an answer. Those decisions shape how accurately chief information officers forecast AI demand and how confidently chief financial officers allocate capital behind it. Bringing together trusted information and context into a single intelligence layer reduces unnecessary work throughout the system before an answer ever reaches the user. The result is a system that scales more predictably, creates much better token efficiency, and gives teams answers they can trust without introducing unnecessary complexity and manual reviews as AI usage grows. Build discipline around AI spend As AI usage grows exponentially, executives and IT leaders must ensure business value grows alongside it. Enterprise AI's next phase will demand greater operating discipline from leaders. Managing token costs must become an ongoing review of how teams use AI, where the system introduces unnecessary complexity, and whether the architecture helps teams reach trusted answers efficiently. As chief information officers optimize architectures, chief financial officers will have greater confidence that AI investments translate into measurable business value. Organizations building that discipline early will better handle AI spending and enable systems where efficiency improves as usage grows. This creates an advantage that scales effectively over time, rather than creating unpredictable, runaway costs that compound. Kiva Kolstein is president and chief revenue officer at AlphaSense.
[4]
Tokenmaxxing: Why AI consumption needs control
AI spend made headlines again recently with the Claude Fable 5 model from Anthropic. Before security concerns led to the model being suspended, there were also cost concerns. Anthropic says Fable costs $10 or approximately €9 per million input tokens and $50 per million output tokens. This is double the price of the company's previously most expensive model, Claude Opus 4.8. Posts soon began to pop up on LinkedIn, showing just how quickly teams were going through their tokens and, as a result, their budget. There are some caveats here. Namely, that Fable 5 is an advanced model and, for most businesses, won't need to run non-stop or be used for every task. But therein lies a key issue: AI use is accelerating and models are evolving. But the level of control and visibility businesses have over how much is being spent, by who and for what is lagging behind. How AI consumption became a finance problem There is a massive shift within the UK software market toward AI and specifically Anthropic's ecosystem. Proprietary data from Pleo looking at the top tech merchants based on number of spending customers, shows that Anthropic (Claude) surged from 12th place in Q4 2025 to 7th in Q1 2026. Meanwhile, the average spend per customer increased +43.0% in this time. This rapid climb signals that Anthropic has reached enterprise maturity in the UK market with businesses moving beyond the experimentation phase. But while this reflects growing confidence in AI adoption, it also presents some financial challenges. On the whole, AI has redefined how the workplace runs, but it is not a free trial. The cost of tokens has gone up, and new models that can achieve what was seemingly unthinkable a few years ago come with a price tag to match. The new challenge for business leaders is to leverage these technologies but also limit rampant spending. This is why many organizations are turning to their finance teams. Finance has the visibility to dig into the details and map AI use across the organization, whether it quietly shows up as a subscription renewal or a new budget request. But more than that, they can be instrumental in ensuring teams embrace open conversations, not just OpenAI. AI activity does not translate to AI value Just about every organization will have developed transformational ways of using AI tools. But, whether they know it or not, there will be wasteful ones too. When it comes to inefficient use, some of the major culprits include asking AI agents open-ended questions, model mismatch where tokens are burned unnecessarily; and duplicate tools, resulting from shadow AI and overlapping subscriptions. These prevent businesses from seeing the full picture; one that is, in all probability, very expensive. User literacy can improve this. But for finance teams they must start with the grey area of AI consumption. Two teams might show as active AI users, but one that's using an LLM to produce more content faster is doing something fundamentally different to one that's using it for peripheral productivity tasks. In fact, only 29% of European SMEs using Gen AI are doing so in core business activities. To improve the control they have over AI, organizations must start by elevating their visibility from who is using AI, to who is using it to become smarter, faster and more productive. How to regain control over AI use A complete view of AI spend is essential, regardless of whether costs are rising. Breaking spend down by department, team and budget helps identify both disproportionate usage and areas where adoption may be lagging. These should be combined with performance metrics such as the time-to-first-draft on marketing content; code review cycle times in engineering; support ticket resolution time in customer support; and so on. This combination of spend and performance can reveal whether AI investment is translating into measurable productivity gains and not just higher software costs. Visibility should also extend to model-level usage. As mentioned before, the cost difference between frontier reasoning models and lighter alternatives can be tenfold. Monitoring model and vendor usage alongside token consumption helps organizations route routine tasks to lower-cost options, maximize ROI and reduce unnecessary spend. Finance teams should therefore expand reporting and budgeting frameworks to include AI-specific metrics. A key question at month-end is whether AI-enabled teams are increasing output and capacity without increasing headcount. This provides a clear headline for AI's impact, can justify investment and distinguish between high-value and low-value AI usage. Ultimately, effective control over AI is not about costs alone. It is about understanding where AI is creating value and ensuring investment is aligned with business outcomes. AI control is at your fingertips The good news is that none of these metrics require a sophisticated AI analytics stack. Finance teams should already have the tools for real-time visibility into what's being spent and where. All that's needed now is to fold AI into the mix and collaborate with other departments to measure and improve its ROI. The outcome is that organizations control AI use through oversight, without restricting spend, adoption or innovation through lengthy procurement processes. Spend policies, category controls and clear approval thresholds control what is spent, and everything is measured. But crucially, teams don't slow down as a result. The only difference is that AI is optimized for impact. We've featured the best AI chatbot for business. This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today. The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit
[5]
AI's trillion dollar token reckoning
Enterprise AI strategy spent two years chasing a single objective: reach the frontier before competitors do. The default path was a public cloud account, an API key from OpenAI or Anthropic, and a willingness to absorb cost in exchange for speed. That reality is now running out of road. The numbers tell the story. Gartner forecasts worldwide AI spending reaching $2.52 trillion in 2026, up 44% year on year, with $1.37 trillion of that flowing into AI infrastructure alone. In fact, in mid-2025, they claimed that procurement of AI had entered a "Trough of Disillusionment," where scaling depends on predictable ROI, rather than visionary pilots. The pressure has now shifted from how fast enterprises can pilot AI to whether they can sustain, govern, and defend it in production. The Race to the Front is Over - Now Comes the Bill We are now past AI 1.0, where simple access to cutting-edge AI was the differentiator. Now it's AI 2.0's turn, where inference economics, data gravity, latency and control decide the outcomes. Token prices have fallen almost tenfold annually since 2021, but AI spend overall by organizations has increased. That's because more capable models have enabled greater ambition. Anthropic, OpenAI, and Mistral are now stratifying offerings between flagship reasoners and lower-cost workhorses precisely because customers refuse to pay flagship prices for every task. McKinsey's 2025 State of AI survey confirms the pattern - adoption is increasing, but impact at scale remains elusive for most organizations. Now CIOs have stopped asking which model, but where each workload needs to run and how much it's going to cost. Inference Cost Inflation Banks delivering the next best action are a good example: the in-app, in-branch, or call-center recommendation served in milliseconds against a customer's live context. The best banks prove that personalization at this layer can lift revenue by 5-15%. To give a firsthand example, a global bank we work with launched an AI assistant that has already resolved more than 1.5 million customer inquiries in its first year, driving huge efficiencies. But the inference economics are unforgiving at this scale. A single agentic decision can chain five to twenty model calls, each carrying its own context window. The cost gap between £0.50 and £3 per million input tokens seems trivial in a single-turn demo. Spread across hundreds of millions of customer events, it becomes the difference between a money-making feature and a money-burning one. This isn't a hypothetical either. Uber's 5,000-strong engineering team's use of Claude Code burned through the company's entire annual AI budget in the first four months of this year. And AI companies are responding to this market shift. Decagon, after re-architecting onto an open-source multi-model stack on NVIDIA Blackwell, dropped cost per voice query by sixfold. Next best action isn't a marketing decision anymore; it's an economic decision. Organizations making the structural shift now will outcompete those treating model selection as an afterthought. Complexity Doesn't Disappear, It Just Moves The hardest lesson of the past 18 months is that model commoditization does not reduce enterprise complexity but relocates it. Open weights from Mistral or DeepSeek cut experimentation cost, but orchestration, governance, evaluation, and integration burdens move up the stack and sit with the buyer. Enterprise leaders should be measuring unit economics per useful task, operational burden per deployed agent, and the ratio of inference spent on the governance scaffolding around it. That ratio is typically 1:5 or worse. A second architectural shift is arriving: sub-quadratic attention. Approaches from DeepSeek, Google, and Cartesia are collapsing the cost of long-context reasoning by orders of magnitude, with recent benchmarks showing 100x to 300x cost reductions at comparable accuracy. Large banks will now be able to run whole-portfolio risk modelling, multi-decade fraud detection and cross-jurisdiction Know-Your-Customer (KYC) as single-pass operations - no more chunked retrieval workarounds. Telcos can make network operations, predictive maintenance and multi-year customer journey reasoning more economically viable at scale. And manufacturers can move full-plant simulation and supply-chain disruption forecasting from periodic batch jobs to continuous reasoning. The architecture that wins will not be the one with the cheapest token. It will be the one that places compute closest to the data, under the right jurisdiction, with governance that holds. Sustainable, sovereign, controlled - that's the new triad. The enterprises that build for it now will define the next decade. We've reviewed, rated, and ranked the best business plan software. This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today. The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit
[6]
Beware the token trap: Why saving on inference might put your ADLC at risk
Token use can create unexpected, sizeable costs for organizations Agentic AI's prolific use of tokens can create sizeable, unexpected costs for organizations. But saving on token costs without factoring in risk can be a fatal step. As upfront prices for flagship artificial intelligence models continue to shrink, organizations have begun to wise up to the hidden costs they encounter with agentic AI models. Specifically, the costs of tokens, which may look tiny when viewed as individual charges, can add up exponentially as AI agents become more active, leaving organizations with hefty AI expenditures they may not have anticipated. This is putting CISOs in something of a bind. If they seek to save money on inference costs, primarily driven by token generation incurred by agentic AI, they may increase their security risk and accumulate hidden technical debt that puts their Agentic Development Lifecycle (ADLC) in jeopardy. It's a problem that many CISOs may not have factored into their security budgets, but it cannot be left unaddressed. The effectiveness of automated security processes is being impeded by fragmented pricing across the AI landscape, whether we're talking about hyper-optimized nano models (essentially lightweight, yet powerful models built for a specific use, like Google's Nano Banana 2 image generator) or premium reasoning engines, like Salesforce Atlas or OpenAI o3. Organizations do have to keep a close eye on token costs to prevent them from spiraling, but CISOs also need to examine how agentic AI is affecting their security. The hidden costs of AI agents Erratic pricing has been a trademark of generative AI pretty much from the beginning. About two years after OpenAI released ChatGPT, the Chinese company DeepSeek shook up the AI market with the release of a powerful, open-weighted large language model whose training parameters were publicly available, allowing users to customize the model and build on the cheap compared with other generative AI models. ChatGPT-maker OpenAI and other AI companies started doing the same, and suddenly, the costs of using GenAI systems dropped off a cliff. In fact, prices fell faster for GenAI than for any other technology in history. The emergence of agentic AI has introduced some stealth costs into the equation, however. The costs of agentic software range from free for open-source models to enterprise agents, with prices that vary from one-time fees (roughly $15,000 for basic models to more than $1 million for global enterprise models) to monthly subscriptions (which can range from a few thousand to $13,000 or more). But those costs are fixed. Inference costs are another story: they scale with usage and can amount to 90% of AI lifecycle costs. Tokens come into play when an AI agent requests processing from GenAI models, which charge agents for processing information. At a glance, the costs may appear inconsequential. Input tokens generally range from 15 cents to $5 per million requests. Output tokens, which require slightly more processing, cost from about 60 cents to $25 per million. They may start small, but can add up in no time, thanks to AI agents that work very quickly, autonomously, and unpredictably. They are designed to interact with systems and other agents throughout the enterprise. A single action might generate scores of LLM calls. Token use, which has grown exponentially with the use of AI agents, has already increased IT budgets by about 20% according to recent estimates. The accelerating cost of agentic AI is prompting CISOs to look for ways to save money where they can, and one way is to identify LLMs that charge the least per token. But what they may not be considering are the risk factors associated with those LLMs. If CISOs concern themselves only with the costs, they may open themselves up to security risks. But better security doesn't necessarily have to cost more. Depending on what they're using agentic AI for, they may find they don't always have to trade security for lower token costs. Getting costs (and risks) under control There are a few things organizations can do to help stop token costs from getting out of hand, including: Match Agents and LLMs to the Job at Hand. Commodity AI systems can cost little or nothing, but they lack the deep reasoning for complex security synthesis. But not every application or function within the organization requires a reasoning engine. You can set up agents to work with low-cost LLMs on low-risk projects, while preserving higher-cost LLMs for critical tasks. It's also worth being aware of which agents are likely to request more LLM calls. Factor Risk Scores in Choosing Agents and LLMs. The security implications of using AI can't be ignored. When developing a budget plan, include risk factors. Monitor Workflows. Keeping a close watch on workflows can help you track costs and performance, allowing you to better understand which tools work best in which situations. Lean on Human Oversight. Despite agentic AI's autonomy, in fact, because of agentic AI's autonomy, forgetting about the importance of the human element is risky business. Teams need thorough upskilling in secure development, with clearly defined ownership roles. And they must be given prominent oversight roles throughout the ADLC. Agentic AI is fast becoming integral to enterprise operations, and organizations must control its associated costs. But a race to the bottom on token pricing creates hidden technical debt. Instead, CISOs need to weigh security performance when choosing AI tools as part of establishing an up-to-date security maturity model and an AI governance policy that emphasizes performance, costs, and risk management. Only that approach allows agentic AI to be deployed without either breaking the budget or putting your entire organization at risk. We've featured the best AI website builder. This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today. The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit
Share
Copy Link
Enterprise AI spending is projected to reach $2.52 trillion in 2026, with $1.37 trillion allocated to infrastructure alone. Organizations are discovering that rising token costs and hidden infrastructure expenses are outpacing business value, forcing a strategic shift from model experimentation to economically sustainable AI deployment focused on unit economics and governance.
Enterprise AI is reaching a critical inflection point as organizations confront the harsh economic realities of scaling AI from experimentation to production. Gartner forecasts worldwide AI spending will hit $2.52 trillion in 2026, representing a 44% year-over-year increase, with $1.37 trillion flowing into AI infrastructure alone
5
. This explosive growth masks a troubling trend: rising token costs are outpacing measurable business value, forcing finance and technology leaders to fundamentally rethink their AI strategy3
.
Source: TechRadar
The challenge extends beyond visible invoice amounts. What starts as a handful of users experimenting with AI tools quickly evolves into dozens of workflows, hundreds of employees, and thousands of prompts daily
3
. Nearly 8 in 10 IT leaders report surprise at charges due to AI models or usage volumes, revealing a dangerous gap between consumption and cost visibility3
. The recent Claude Fable 5 model from Anthropic exemplifies this pressure, costing $10 per million input tokens and $50 per million output tokens—double the price of Claude Opus 4.84
.Behind every AI interaction sits physical IT infrastructure consuming compute, memory, networking, electricity and cooling—costs that remain largely invisible despite directly influencing enterprise AI economics
1
. Organizations focusing solely on token pricing risk overlooking factors determining long-term cost, resilience and sustainability. Purpose-built inference infrastructure can significantly improve energy efficiency compared with architectures optimized for training workloads, translating to lower cooling requirements and simplified facility design1
.The hidden costs of enterprise AI manifest in unexpected ways. A prompt 20% less efficient can drive costs 200% higher as models spend more tokens compensating for weak retrieval and noisy context
3
. An agentic decision can chain five to twenty model calls, each carrying its own context window, making the cost gap between £0.50 and £3 per million input tokens the difference between profit and loss across hundreds of millions of customer events5
. Uber's engineering team burned through their entire annual AI budget in just four months using Claude Code, demonstrating how quickly AI spending can spiral without proper governance5
.The conversation is shifting from purchasing tokens to building sustainable AI capability. Consumption-based pricing remains appropriate for variable workloads, but as AI embeds within everyday operations, continually increasing token consumption creates an operational cost model that becomes progressively harder to forecast and control
1
. Dedicated inference infrastructure represents a different economic model, offering predictable operating costs and greater control over performance, data sovereignty and operational resilience1
.
Source: TechRadar
Proprietary data shows Anthropic surged from 12th to 7th place among top tech merchants in Q1 2026, with average spend per customer increasing 43%
4
. This rapid climb signals enterprise maturity but also presents financial challenges requiring AI consumption control. Ruth Patterson, Managing Director at HP UK & Ireland, notes that businesses seeing success aren't integrating AI at every turn but "applying it to practical workflows in a way that protects data and delivers real results"2
.Related Stories
Enterprise leaders must measure unit economics per useful task rather than total AI activity. Token volumes measure AI activity levels but not business value created by that activity
1
. Organizations should track performance metrics like time-to-first-draft on marketing content, code review cycle times in engineering, and support ticket resolution times to determine whether AI investment translates into measurable productivity gains4
.Major inefficiencies include asking AI agents open-ended questions, model mismatch where tokens burn unnecessarily, and duplicate tools resulting from shadow AI and overlapping subscriptions
4
. Only 29% of European SMEs using generative AI deploy it in core business activities, suggesting widespread wasteful usage4
. The cost difference between frontier reasoning models and lighter alternatives can be tenfold, making model selection an economic decision rather than a technical one5
.Sub-quadratic attention approaches from DeepSeek, Google and Cartesia are collapsing long-context reasoning costs by orders of magnitude, with recent benchmarks showing 100x to 300x cost reductions at comparable accuracy
5
. Large banks can now run whole-portfolio risk modeling and multi-decade fraud detection as single-pass operations without chunked retrieval workarounds. Decagon re-architected onto an open-source multi-model stack on NVIDIA Blackwell, dropping cost per voice query sixfold5
.
Source: TechRadar
Gartner claims AI procurement entered a "Trough of Disillusionment" in mid-2025, where scaling depends on predictable ROI rather than visionary pilots
5
. The winning architecture will place compute closest to data under appropriate jurisdiction with robust governance—creating a new triad of sustainable, sovereign and controlled AI deployment5
. Organizations must balance cloud and local AI to find the sweet spot between cost, model performance and security2
. Finance teams should expand reporting frameworks to include AI-specific metrics, asking whether AI-enabled teams increase output without increasing headcount4
.Summarized by
Navi
[3]
[4]
[5]
06 Jul 2026•Business and Economy

28 Jul 2026•Business and Economy

17 Jun 2026•Business and Economy

1
Technology

2
Technology

3
Technology
