14 Sources
[1]
Token-maxing is an AI cost sink - how to use agents without busting your budget
Follow ZDNET: Add us as a preferred source on Google. ZDNET's key takeaways * Token-maxing to support agentic AI is unsustainable. * Business leaders must create a strategy for token use. * Give people room to explore agents within guidelines. Twelve months in AI is an eternity. Steve Lucas, CEO at integration technology specialist Boomi, was concerned last year that CIOs were rushing into gen AI initiatives without a clear sense of direction. Now, a year later, he's concerned that IT professionals are taking a similarly rushed approach with agentic technology, and the scale of token usage is only going one way: upward. "Whether you work inside or outside a company, it feels like your core hustle is to use AI and you token max the heck out of that technology for your own job," he said to ZDNET. Also: The 3 types of people who will excel in the AI agent era, according to tech leaders A token is the fundamental unit of information processed by an AI model, and research points to the soaring and unpredictable costs of agents as the technology consumes orders of magnitude more tokens. As my ZDNET colleague Steven Vaughan-Nichols discussed recently, it wasn't so long ago -- certainly during the rise of generative AI -- that everyone was excited about their token leaderboard, which showed who had the most token usage in an organization. Today, token leaderboards are obsolete because no one can afford to waste tokens. In the agentic era, token maxing is a sign of excess, and "tokenomics" -- the practice of measuring, pricing, and managing the consumption of tokens -- is a key business activity, even for a tech CEO like Lucas, who noted that tokens are being consumed at a much faster rate. "A year ago, we weren't talking about tokenomics," he said. "But last year, I personally spent at Boomi 10 times the amount on Claude that I did the previous year -- 10 times; that's not sustainable. I can't do that every year." Also: The new enterprise AI expert every company needs - and why The onus now is on businesses and professionals to create their own enterprise-ready version of tokenomics. Agentic AI is set to transform how every business operates -- and creating a cost-effective technique for exploring agents is an urgent priority, suggested Lucas. "What matters in the enterprise is ultimately the economics of AI," he said. "Most organizations will look to AI as the enterprise engine of the future. So, what matters now is, 'Can I operate AI at a return?' That is the fundamental question." Business leaders suggested the best way forward is clear: Rather than constraining how people can use agents, give them guidelines and the context to make cost-effective model decisions. Letting people shine Snowflake CEO Sridhar Ramaswamy recognized that token consumption is on the rise, even in his own organization. However, he also said it's critical to give staff the wriggle room to explore agents. "Are we worried about how much we are spending on AI inference across our different internal teams? Absolutely," he said. "But do I see that spend as a reason not to use AI? Absolutely not." While token maxing is a concern, Ramaswamy said during a media session at his firm's recent Summit 2026 event in San Francisco that the creative use of AI can help unlock new business opportunities. Also: 40% of enterprises will scrap AI agents - 3 ways to ensure yours don't fail Agents can help boost staff efficiency and productivity, with the potential to generate new services for his firm's clients faster and more effectively. "We look at agentic AI as an opportunity to optimize not just what we do, but how we can turn a process into a product that our customers can use," he said. Like Ramaswamy, Matt Luizzi, VP of analytics at wearable technology specialist Whoop, said his company is investing in agentic explorations to find a competitive advantage. "If we want people to push themselves out of their comfort zones, we're going to need to be OK with them taking risks and understanding that you can't break anything." Also: AI is causing cognitive fatigue. Here's how to work with more haste and less speed In short, Luizzi said that he's OK with staff spending money on agents, but that doesn't mean token consumption gets out of hand. "We have guardrails in place, and monitoring and observability to let people know what they're spending," he said during a panel session at Summit 2026. "But more likely than not, we just need to sit down and enable them on, 'How could you be doing this more efficiently and what are you trying to accomplish?'" Putting everything into context The key to agentic success, suggested Luizzi, is carefully managed enablement. Yes, let people push the boundaries with agents and tokens, but don't let them go overboard. "At the end of the day, we're leaning in hard because we think there's an ROI to this. We are seeing people become more efficient. We are seeing work get done faster," he said. Also: How Workday and other software providers plan to survive AI "Those positive results mean we're not trying to be super keen on how many tokens people use but also understanding that this is something that needs to scale, so not enabling people to go crazy either." Sriram Sitaraman, CIO at technology specialist Synopsys, said in the same panel session at the Snowflake conference that people must be able to explore new avenues, including agents, to discover innovative solutions to business challenges: "You can't innovate with constraints; you can only go so far." Like Luizzi, Sitaraman said guidelines can help set tight boundaries for token consumption. He suggested that the right context for projects can help establish effective constraints. "If you ask an agent, 'What should I do tomorrow?' it might consume a whole lot of data, a whole lot of tokens, and come back with something that isn't useful," he said. "If you set a context, such as creating a North America sales ops agent, giving the project specific objectives and data, then the consumption is going to be limited, and the value that it delivers to the user is higher." Also: How to beat the AI algorithm and get the job of your dreams That focus on business context was echoed by Francois-Xavier Pierrel, group chief data and ad tech officer at French TV network TF1, who told ZDNET that a common-sense approach is the best way for business and professional users to control token use. Pierrel compared the alternative approach to a sugar addiction. Once you begin adding sugar to your food, you get used to it and start to develop an unhealthy addiction. He said AI presents similar risks. "Everything was cheap, very cheap, and we got all into it," he said, looking back at the early rush to use generative AI technologies. Now, agent obsession represents a more costly proposition. "The number of tokens used can quickly become a big number. And at some point, if we look at all the companies investing crazy money into agentic AI, the bill will come back." Also: Forget productivity: Here are 5 strategic shifts that drive real AI value Pierrel said caution is the watchword in his organization, and he gave an example of how professionals are asked to consider their options when using large language models. "For use cases that are very narrow or have strong boundaries, we say, 'Can we use a small model, something that will consume fewer tokens than expected or planned, and that will still do the same job?'" As companies look to deploy more agents as part of working processes, Pierrel suggested this kind of tokenomics strategy will be crucial to value creation. Also: AI agents are your new colleagues - how to get the best results "I think we are going to be more agile on this one, trying stuff, and then adapting to the use case to avoid having crazy bills," he said. "We're not shooting a rocket to the moon. So, let's be reasonable -- don't deploy a million agents if you can do it on a smaller scale and it will still work."
[2]
Artificial intelligence: Why firms are struggling to set prices
If you have used a free version of an ChatGPT or its AI rivals, then you are obviously getting a good deal. Firms like Microsoft, Google and Anthropic have invested hundreds of billions of dollars in developing Large Language Models (LLMs) the tech behind those services. So getting, ChatGPT, Claude or Gemini to help with your speech or holiday plans is a bargain. But, naturally, those firms want to recoup their investment, so they offer paid-for versions of their AI, which have extra features for tasks like coding or billing. Meanwhile, third party firms are building and selling services based on AI agents, usually based on an LLM, which are trained to do specific tasks. But setting a price for those services is surprisingly difficult. "Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn't make any sense, honestly, because we don't know," says Simon Gooch at Saviynt, an identity management company which is incorporating agentic AI into its services. That's because of rapidly changing economics around tokens, the building blocks of LLMs and agentic AI. When a user asks an LLM, like ChatGPT or Anthropic's Claude to answer a question, generate software code, or automate a process, that prompt is broken down into mathematical chunks called tokens, which can be processed by the model. The LLM's response also comes in the form of tokens, which are converted back into text, software code, or a set of commands to automate a process. The problem is this process is not entirely predictable. Subtle variations in the prompt can produce different answers. The same prompt will not always produce the same answer. Different models will produce different answers. Meanwhile, in agentic systems, businesses use multiple AI agents together to make decisions and take actions, further increasing both token use and unpredictability. While the cost of individual tokens - or the credits used to pay for them - has plummeted in recent years, according to analysis by Goldman Sachs, the number of tokens consumed by businesses, and consumers, has skyrocketed. The bank forecasts that token consumption will increase 24 times between 2026 and 2030 to 120 quadrillion tokens a month, as companies shift from to use AI agents. But companies, and individuals, using AI systems often have a tenuous grasp on just how many tokens they are burning through - until they either run out or get their monthly bill. Even Microsoft has reportedly reined back its engineers' use of some third party coding tools, while Uber apparently tore through its AI coding token budget for a year in a matter of months earlier this year. Will Venters, Associate Professor of Digital Innovation and Information Systems at the London School of Economics, said companies can be caught out as they experiment with or implement AI internally, as staff burn through tokens. "People are finding it really hard to manage that cost... it's a non-deterministic output, so it's a non-deterministic value," he said. Companies are finding ways to work around this. Oliver King-Smith, founder of engineering software firm smartR AI, says smaller organizations can "can fly under the radar and use [flat fee] personal accounts which I am sure the big vendors don't like." But, he says, "This has to end at some point in time, because the big guys are taking a bath on those accounts." Once the big AI platforms start facing pressure from shareholders to show a profit, he predicts: "They will start clamping down." King-Smith says companies should also think more carefully about what AI models to use. Companies also needed to be much more precise with their prompts, says Rob Steele, CFO at UK accounting software firm iplicit. "You wouldn't send someone in your family out to get the weekly shop without any kind of detailed instructions as to what you expect in that shopping basket, right?" The situation can become difficult to control when companies build AI into a product that could be rolled out to thousands of users, Venters points out. AI costs could start to balloon. For example, managers may realise they need tokens not just for core software development, but for other tasks such as testing, security, or for implementing guard rails. "It's particularly hard when you're looking at agentic processes," Ventners says. Employing more AI agents can be done with the click of a button, whereas expanding the human workforce would involve careful discussions over headcount and hiring, he says. Venters points out, while token costs might be unpredictable, it might be that the company is ultimately getting more value from their token use with AI. "It's not quite the same as a calculator," he says. "The more you give it, the more expensive it is, but the better the result may be." But companies still need to pass those costs onto their own customers. "Nobody's really figured it out," says Bill Peterson, senior director of product marketing, at Sumo Logic. The software firm is previewing new security services based on agentic AI, he explains, but is in discussion with corporate customers about how to charge for them. "We're still having some fun conversations about this internally," he says drily. Options could include simply raising prices across the board, he says, paying by results, or charging for "bundles" of incidents. But whatever price structure it chooses could be upended if and when the large language model providers change their own pricing strategies. "You get into variable pricing, and it's changing every couple of months" he says. "Customers don't like that. That's not how anybody builds a budget."
[3]
What Are Companies Getting for All That A.I. Spending?
Corporate America was enthusiastic about artificial intelligence. Until it got the bill. Whiplash around spending on "tokens," the units of computing power in which A.I. is sold, is hitting engineering teams and board rooms. First there was "tokenmaxxing," as executives encouraged as much A.I. use as possible. Then there was "tokenminning," after they rapidly burned through millions of dollars in company money. The simultaneous urgency and uncertainty is raising complex questions. What is A.I. even good for? How do you know what you're buying, and measure the value of what it enables? How is it priced, and how will those prices change in the future? What does the spending on it displace? "It's a currency where you have no instinct to know what you're using, and the accounting practices aren't even there for it," said Howard Rubin, an economist who advises companies on technology spending. "The A.I. stuff is being treated as an investment right now, but it's a risky investment in case it has no return." The high stakes and scarcity of knowledge have given rise to a new field within economics devoted to figuring out how companies buy A.I. and what value they get from it. Call it tokenomics: the study of how this limited resource is created, traded and converted into things people want. Businesses are no strangers to adopting new technologies. Electricity allowed power to be generated far from where it was used. The internet radically changed how services are delivered. Cloud computing turned data processing from a clunky on-site task to something that could be rented remotely. The implications of A.I. could be as profound, but the costs and benefits are more difficult to weigh. Perhaps the closest analogy to how A.I. works as a business expense is petroleum products. A common feedstock, crude oil, is refined into commodities such as gasoline, diesel and propane. Those fuels have different purposes, but they can be broken down into a standard measure of energy -- British thermal units -- and they are generally sold wholesale through exchanges where prices fluctuate with supply and demand. The feedstock for tokens is electricity and semiconductors. The resulting "compute" is fed into models and metered out by tokens, which are generally understood to represent a few characters of English text. But the petrochemical parallel ends there. There are now thousands of A.I. models, and little transparency around who is paying for what. There is no consensus on how much energy a token requires, which model is best to perform a given task or how many tokens it should take. "Intelligence as a utility has so many different ways in which it can be deployed, so the value identification problem has become a lot harder," said Ram Bala, a professor of A.I. and analytics at Santa Clara University. Some models, for example, will keep trying to accomplish a goal even when the mission is hopeless. "On the one hand, you want to solve a difficult problem. That's the positive side," Mr. Bala said. "The negative side is, this could go on in a loop forever, and not really come to a reasonable solution at all." The Linux Foundation, a nonprofit that hosts open-source technology projects, previously developed standards for cloud computing that made it easier to compare services across providers. In June, it established the Tokenomics Foundation with similar goals: to create common parameters that A.I. providers agree to disclose. "There's a set of decisions that every organization in the world is making right now that we want to standardize the frameworks they're thinking on, so they can have better starting places for them," said J.R. Storment, the foundation's executive director. Companies are using a new crop of tools and service providers to get some benefit from A.I. without wiping out their profit margins. Revenium, for example, helps track whether token spending improves productivity enough to justify its cost. Jason Cumberland, Revenium's co-founder and chief operating officer, said he recommends clients set limits for how much of the expensive models engineers can use. After maxing out their allocated token use, engineers turn to cheaper open-source models, forcing them to think hard about what they really need. Revenium helps clients map their token spending to specific outcomes, like the number of software features shipped. "The people who come to us are experiencing spending explosions," Mr. Cumberland said. "They want to understand whether it's worth it, but also how to curtail it in a way that doesn't just stop work, which is the challenge." One big question they all face: Where does the money come from to pay for all this extra token spending? At first, Mr. Cumberland said, companies scrounged under their proverbial couch cushions to pay for subscriptions to Claude or ChatGPT. As they came to rely on those services, they started to cut software licenses and outside services, like web development and copy writing. Now, companies are weighing whether token spending should be considered part of a company's labor budget. That's the main question economists have: whether A.I. will displace or augment human work. In a thought experiment, analysts at Bain & Company, the consultancy, recently posited a medium-term scenario in which tokens make up about a quarter of corporate operating expenses. Elisity, a cybersecurity company, says it's nowhere near there, but it is laying the groundwork with a metric it calls "bionic head count." The measure totals up A.I. spending and divides it by the average cost of a worker's salary and benefits. It then divides annual revenues by the resulting number of human and virtual "employees" to measure their collective output. To make it all feel a little more realistic, human staff members even write job descriptions for each A.I. agent in order to justify "hiring" it. Each time an engineer runs a model on a new task, it must get an evaluation. "Essentially, you're adding virtual head count. Is that resulting in incremental revenue which is all that really matters, or are you just eating at your margins?" said Charlie Treadwell, Elisity's chief marketing officer. "In a growth stage company, it shifts our mind-set of where are we going to spend the capital to grow faster." That equation could change quickly, however, if token prices rise substantially. Elisity is getting by on a flat-rate team subscription for 150 users. If the company were charged by the token -- the norm for organizations with more than a few hundred users -- spending would quadruple, Mr. Treadwell said, and it would have to reconsider its use. (Elisity has no plans to shed staff, however, and is hiring.) To make matters more complicated, tokens can be priced differently across cloud providers, and it's not clear what drives prices up or down. With A.I. laboratories in an arms race to win market share before going public, what they charge may not completely correspond to what tokens cost to produce. And A.I. can help minimize its own costs. Part of that is because of the increasing use of cheaper so-called cache tokens, which the model has already processed once and can reuse at a fraction of the price. As more businesses use A.I. through software tools that can execute tasks on their own, called agents, those agents are able to delegate more of the work to cache tokens. In a working paper published this month, a team of economists found that dynamic has driven spending down relative to what it might otherwise be, even as token consumption and the posted prices for frontier A.I. models have risen. The ability of token costs to stay competitive with human salaries will bear heavily on which kind of intelligence companies lean on in the future. But there is still a shortage of people who know how to deploy A.I. in productive ways. That is keeping consultants like Jue Wang busy. "I actually think that is the constraint that people don't talk enough about in the industry," said Ms. Wang, a partner at Bain. "We often think about 'Oh, they don't have capacity.' But they also don't have enough people to help deliver the value." One of the fundamental difficulties in studying how spending on A.I. is playing out within organizations is the lack of comprehensive data that track it. There is no centralized exchange, no futures market, no government survey that reports prices and spending. It's why Aleh Tsyvinski, an economics professor at Yale who co-wrote the recent paper on token use, was excited to analyze a data set from the platform OpenRouter representing just 2 percent of total A.I. spending. It allowed him to study how financial markets react to those expenditures, but many questions remain. "It's one of the biggest challenges of our generation," Mr. Tsyvinski said. "It's good to understand how A.I. is probably going to probably change everything. Or maybe not change anything, but it's good to have the measurement."
[4]
EY built an 'AI router' to stop its own AI bills from spiralling
The Big Four firm is steering tasks to cheaper models to control token costs, part of a wider corporate reckoning with the price of AI. EY has built what it calls an "AI router," a system that steers each task to the cheapest model that can handle it, in an effort to keep its own soaring AI bills under control. The tool, reported by Business Insider, is the Big Four firm's answer to a problem now spreading across corporate IT: the cost of the tokens that AI consumes.The logic is simple arbitrage. Not every request needs the most powerful, most expensive model, so a router sends easy work to a cheap one and reserves the pricey models for the hard problems, trimming the bill without obviously trimming the output. EY has reason to watch the meter. The firm invests more than $1 billion a year in AI, runs a fleet of some 1,000 AI agents, and has seen its AI-related consulting revenue jump around 30%, a scale at which token costs stop being a rounding error, in a market where the most AI-obsessed firms spend thousands per employee a month. Its own research shows the anxiety is widespread. In EY's latest AI Pulse survey of 534 senior US business leaders, 82% said they were concerned about token-usage costs, and 98% of those using token-based tools said the costs had made them reconsider their strategy. Yet most companies are flying blind. Only 64% of the firms surveyed said they actively monitor token usage with budgetary guardrails, which means a third are spending on AI without a clear meter, a recipe for the bill shocks that have hit the sector. The mood has shifted from more to enough. EY's global AI consulting leader, Dan Diasio, put it plainly: "'AI saves time' is no longer sufficient when costs mount and remain unclear," a line that captures the turn from adoption at any price to value at a known one. The token economics are genuinely strange. The price per token has collapsed as models get cheaper, yet enterprise AI bills have tripled, because agentic tools that run many steps consume far more tokens than a single chatbot prompt ever did. That is the paradox a router is built for. If each task can be matched to the least costly model that still does the job, a company can keep using AI aggressively while stopping the total from ballooning, which is exactly what EY is trying to prove at its own scale. It is not alone in the effort. The industry spent two years urging staff to use as much AI as possible, a fashion nicknamed tokenmaxxing, and is now swinging the other way, with firms from Atlassian to Amazon imposing budgets and controls. For a consultancy, though, the router is also a product. EY sells AI advice to other companies, so a tool that visibly tames its own costs doubles as a demonstration, evidence that the firm can do for clients what it has done for itself. The survey points the same way. Some 76% of leaders told EY that off-the-shelf software no longer meets their needs, and 91% now see building AI tools in-house as critical, a shift that favours firms selling the expertise to build them. The catch is that in-house building is hard. Nearly three-quarters of the leaders EY surveyed said their own AI development was slowing progress, and a third flagged shadow-IT and governance risks, the messy reality behind the clean promise of a router. What the story really marks is a change of question. The first phase of the AI boom asked whether a tool worked; the second, which EY's router belongs to, asks what it costs, and whether the value justifies the meter. EY's answer, for now, is to build the meter itself. A firm that spends a billion dollars a year on AI has decided that the way to keep spending is to watch every token, which is less a retreat from AI than a sign of how much of it is now in use.
[5]
AI tokens could become the kilowatt-hour of the AI age
Earlier this year, OpenAI CEO Sam Altman declared, "We see a future where intelligence is a utility, like electricity or water, and people buy it from us on a meter." Many other AI companies seem to be betting on a similar future. If that vision actually comes to be, then "AI tokens" may break out of the world of nerdy tech and econ conversations and become a much more familiar part of our lives. The number of AI tokens businesses and consumers use could even become one of the defining measures of a new industrial age -- the AI equivalent of the kilowatt-hour for electricity: the standard way we measure and pay for AI usage. Think of tokens as the meter running in the background every time you use AI. Every time an AI model reads your prompt, writes an answer, or does a task, that work is measured in tokens. They are the tiny chunks of text and other data that AI models read and generate. In general, the more work a model does, the more tokens it typically processes. While AI companies still tend to offer flat-rate subscriptions to average consumers, they're increasingly charging businesses and developers based on the number of tokens they use. At the same time, companies, particularly in the tech sector, have been using AI in more and more of their work. As their AI usage has soared, many businesses have discovered just how expensive token-based pricing can become. After a period when tech workers were engaged in a kind of AI free-for-all -- which some dubbed "tokenmaxxing" -- companies like Uber and Amazon have been putting guardrails around AI use and reducing their soaring token bills (which one clever writer at The Information recently dubbed "tokenminimizing"). But tokens aren't just the way AI companies measure usage and charge many of their business customers. They also leave behind a kind of digital paper trail that a growing number of economists and other researchers are using to track AI usage and study its economic impact. In a new working paper, Nicola Borri, Aleh Tsyvinski, and Yukun Liu do basically that. Using data from 380 trillion AI tokens, these economists try to understand how the growth of AI usage is reshaping financial markets. They ask a simple question: as overall AI consumption changes over time, which companies' stock prices tend to rise with it -- and which tend to fall? To be clear, the economists aren't tracking how many AI tokens individual companies are using -- which, let's be honest, would be way cooler. Instead, they're looking at the rising tide of overall AI usage and asking which stocks rise and fall with it. Part of their analysis is classic finance theory. The idea is that stock prices tend to reflect investors' expectations about the future, so they can offer a window into which companies Wall Street sees as most positively or negatively impacted by the rise of AI. Of course, investors can be spectacularly wrong. Hello, the long history of financial bubbles. That said, if you want to know what markets expect AI to do to the economy, stock prices are a compelling place to look. Not surprisingly, the economists find that as AI usage has grown, financial markets have treated some companies very differently than others. The companies seen as the biggest AI beneficiaries have enjoyed higher stock returns -- a pattern the researchers call an "AI premium." More interestingly, they find it's not just tech stocks that appear to earn that premium. Their findings suggest investors expect AI to benefit a wide range of companies and industries across the economy. "The story of AI is no longer just a Silicon Valley story," Tsyvinski says about their paper's findings. "Financial markets already see Main Street being impacted." What's most exciting about this study is probably less its findings and more that it offers a glimpse into future research possibilities. Instead of relying on surveys, earnings calls, or company announcements, economists may be able to study the spread of AI by following the digital paper trail created by AI tokens. AI tokens could become a powerful new source of data, allowing researchers to track AI usage in almost real time and study its economic effects, with a kind of precision not possible in past technological revolutions. An exciting way to measure the AI economy Borri, Tsyvinski, and Liu's analysis relies on a relatively new source of data: OpenRouter. OpenRouter is a kind of one-stop shop for AI models. Instead of contracting separately with OpenAI, Anthropic, Google, and dozens of other AI companies, developers can access hundreds of models through a single interface. As businesses and workers have become more conscious of their mounting AI token bills, they've turned to platforms like OpenRouter to compare prices, switch to cheaper models for less complex tasks, and manage their token spending. In the process, OpenRouter has amassed rich data on usage of AI tokens (and this data is anonymized to protect users' identities). The economists analyze the use of 380 trillion AI tokens between January 2024 and April 2026. That represents around 2 percent of monthly global AI usage. They then combine weekly growth in tokens, spending, and active users into a broad measure of AI consumption, which they call the "AI Factor." Next, they estimate which companies' stock returns move most strongly with changes in that factor. The economists find that companies whose stock prices were most sensitive to increases in overall AI consumption subsequently earned significantly higher returns. The companies that Wall Street appears to view as the biggest beneficiaries of AI outperformed those viewed as the least likely beneficiaries by about 0.64 percentage points per week -- the "AI Premium" described in the paper. Of course, markets aren't always right. These AI-related stock movements could reflect the wisdom of crowds -- or simply a herd of excited investors hoping to cash in on a technological revolution that ultimately disappoints. In other words, take these findings with a grain of salt. Probably the most interesting of their findings is that the AI premium can be found well beyond the tech world. Markets seem to believe that companies in industries ranging from airlines and cruise lines to utilities, industrial manufacturers, retailers, banks, and even waste management companies could all benefit as AI reshapes the economy. They also find that this "AI premium" is strongest for companies in the United States and Europe, and is much weaker in China and other emerging markets. Perhaps relatedly, they find that stocks are the most sensitive to increased use of the most technologically advanced, "frontier" AI models. The economists even provide a list of S&P 500 companies with high AI premiums in their paper. The top five are: AppLovin, Carvana, Lumentum, Expand Energy, and Baker Hughes. They also identify companies Wall Street appears to view as AI's biggest losers. The bottom five are: Moderna, Estée Lauder Companies, ON Semiconductor, Skyworks Solutions, and Aptiv. There are a bunch of caveats to their analysis. First, this is a working paper that has yet to be peer reviewed. Second, OpenRouter's token data likely provides a skewed picture of AI usage. People who use OpenRouter are likely sophisticated, heavy users of AI who are trying to save money as they shop between competing models to get the best bang for their buck. Their data doesn't represent the sort of average consumer, who may have a ChatGPT or Claude or Gemini monthly subscription and sticks with using those. Perhaps most importantly, even if the researchers have correctly identified the companies investors expect to benefit from AI, that doesn't mean those expectations will prove right. And if markets are reasonably efficient, much of that optimism may already be reflected in today's stock prices. So buying the stocks they identify as potential AI-related winners and selling the ones that could be AI-related losers is by no means a good financial strategy. This is not financial advice! Still, the paper's biggest contribution may not be what it says about today's stock market. It may be that it points toward a new way for economists to measure the spread of AI through the economy. We welcome this era of AI measurement-maxxing. There are still many big economic questions about AI, and better data is rarely a bad place to start.
[6]
Atlassian puts its engineers on an AI budget as the cost of 'tokenmaxxing' bites
The software maker cut thousands of jobs in the name of AI. Now it is handing the engineers who remain capped 'AI wallets,' a sign the industry's token spending has become a line item to manage. Atlassian has started giving its engineers a fixed monthly allowance for artificial intelligence, a capped "AI wallet" that warns them as they near the limit and stops when the budget runs out. The move places the software company on the cost-conscious side of a growing divide over how much AI staff should be allowed to burn through. The caps are not trivial, but they are caps. Atlassian's wallets run from $500 to $2,000 a month for research-and-development staff, with the size set by role and the option to ask for more, a structure meant to keep spending visible rather than to starve it. The company frames it as generosity with guardrails. "Atlassian provides a significant budget for our builders to leverage multiple AI tools," a spokesperson said, casting the wallet as a way to fund experimentation without letting the bills run wild. There is an irony in the metering, and it is not a small one. Atlassian has recast itself as an "AI-first" company, cutting 1,600 jobs to fund the pivot, and earlier in the year it told a group of support staff, over video, that they would be largely replaced by AI. Now it is rationing that same AI for the people who kept their jobs. Having sold the technology as efficient enough to replace workers, the company is finding it is also expensive enough to need a budget, which is a harder line to put on a motivational slide. Coding agents that once cost pennies now run tasks that consume tokens by the million, and the practice of maximising that usage, nicknamed "tokenmaxxing," has turned individual engineers into meaningful cost centres. The arithmetic explains the anxiety. A token is roughly four characters of text, and at prevailing prices of several dollars per million tokens for the leading models, an agent left to iterate on its own can run up a large bill, the kind of dynamic that has already broken the economics of tools like GitHub Copilot. Atlassian is not the first to react. Amazon quietly shut down an internal leaderboard that had turned heavy AI use into a competition, after employees gamed it by consuming tokens for their own sake. Meta went further still; the company warned some 6,000 staff that its internal AI spending in 2026 could reach into the billions, and began rolling out token budgets and controls of the kind Atlassian has now adopted. If last year's fashion was tokenmaxxing, this year's is its opposite, a turn toward treating AI spend as something to manage rather than a badge of how forward-leaning a team is. The reversal is awkward for an industry that spent two years urging staff to use more AI. Executives who once measured adoption by token volume are now measuring it by cost per outcome, a soberer metric that fits a moment of tighter budgets. Not everyone is pulling back. Some firms still offer unlimited AI budgets as a recruiting perk and a bet on productivity, wagering that the output justifies the spend and that capping it would only slow their best engineers. That is the real disagreement beneath the wallets. No one doubts the tools are useful; the fight is over whether unlimited access produces enough extra value to be worth an unpredictable, fast-rising bill. Atlassian's answer is a middle path. By funding multiple tools but metering their use, it is trying to keep the productivity of agentic AI while stripping out the waste, a balance the whole sector is now hunting for. Whether that produces better engineering or merely cheaper engineering is the open question. For now, a company that thinned its workforce in the name of AI is asking the survivors to watch every token, having learned that the technology it sold as a way to do more with fewer people is not, as it turns out, free.
[7]
AI coding agents are blowing through budgets -- Replit, Kilo Code, and Symbotic explain how they're managing it
At Kilo Code, engineers are reading or writing code themselves only about 1% of the time now, according to co-founder Emilie Schario -- the rest is agents. That shift is forcing new questions onto dev teams: which systems are safe to hand over, who cleans up when models goof up, how to support multi-model architectures, and whether skyrocketing token bills mean real progress or just burned IT budget. As far as tech leads from Replit, Kilo Code, and Symbotic are concerned, it's a natural -- and welcome -- evolution as agentic AI becomes embedded into more and more enterprise workflows. "Unless something's really broken or debugging, 99% of the time engineers are not reading or writing code anymore," Emilie Schario, co-founder of Kilo Code, said at VB Transform 2026. AI good at greenfield, not so great at brownfield For Jared Go, distinguished engineer for AI and cloud at warehouse automation company Symbotic, the current moment is about directing the focus of AI. "These are my criteria," he said. "Let's look at it from the lens of security, elegance, clean, concise code, water tightness." That way, AI does most of the heavy lifting, and human code review isn't as critical. Human involvement becomes necessary further down the line, Go noted, because agents don't make strong product decisions. "Greenfield [building brand new codebases] is so easy for agents. Brownfield [writing, updating, or maintaining existing code] we all know is where the actual challenge lies." Replit takes a bit of a different tack: While the company has "gone very agentic," they've been more conservative with AI coding, explained Amol Jain, head of product engineering. An agent reviews each pull request (PR) and assigns it a risk score; low-risk PRs are self-merged by their author, while others go to human reviewers who read the code and give feedback. "The idea was human on the loop, not human in the loop," Jain said. Replit's internal tool is essentially self-driving for software engineers; devs give a task to agents, which do end to end planning, implementation, and testing. "It's a fleet of agents that run in their own cloud virtual machines (VMs) with access controls behind token proxies so they're secure," Jain said. He shared one example where an engineer couldn't repro or solve a "very gnarly bug" deep in its systems. It was sent to an AI manager agent, which told it to go to sleep. The manager agent then spun up a bunch of underlying agents that found the issue; it subsequently spun up a bunch more agents that found the fix. Six hours later, AI had a PR ready for the bug that had puzzled human engineers. Multi-model is the future AI providers are also evolving beyond the lock-in model, as customers increasingly demand multi-model choice. Kilo Code, for its part, supports 500-plus models in its gateway. "Your software that you're using to do agentic engineering should be decoupled from the model that you're using to do it," Schario said. For instance, Schario said companies often use expensive frontier-tier models to architect a project, then switch to a less expensive open-weight model for the rest of the work. It's also important to respect model provider limitations, such as when they need to work in closed or isolated environments or providers in their specific regions. "It's factoring in what's important to you, what limitations you've set, what data retention policies you've established, what keys you've brought in, what commits you might have ... into that routing decision," Schario said. Replit, similarly, tends to have a better sense of the cost versus capability spectrum than its customers, Jain contended. "We are essentially making the decisions on users' behalf of what model to use when, in what capacity, to minimize cost and maximize capability." To tokenmaxx or not to tokenmaxx Of course, an important consideration as AI adoption increases is runaway costs, which has led to some enterprises tracking and capping AI use through tokenmaxxing. Concerns come from both sides, Schario said: internally and from customers. From the latter, she's hearing, "I accidentally spent my whole AI budget for the year ... so what do I do now?" In response, Schario said Kilo Code points customers to the same workflow: use expensive models for planning, then open-weight models for affordability. Further, sharing skills, strong guidance, and Model Context Protocol (MCP) will empower models. "Realizing where you can really uplevel your team to help them get the most out of the models they're using is going to make a big difference," Schario said. Internally, meanwhile, Schario noted one particular engineer that has a "heavy foot" and is constantly at the top of the usage board. "I regularly have to nudge, 'What are you doing there?'" she said. It's easy to look at a $600 bill for daily work and react, "Wow, that's so much," but looking at the amount of work completed can sometimes justify the cost. "Cost per pull request is the metric that I'm paying attention to right now," Schario said. "It feels like the closest proximity for how I can measure value." Ultimately, AI changes how enterprises are thinking about ROI because spend is not the problem. "The spend with no return on that spend is the problem." Symbotic, for its part, has set per-month cost tiers for its employees. The company built a tool that gives managers visibility into PRs and usage trends. They can then move users up or down a tier as they see fit, Go explained. "Having a cap and seeing how many people went up in cap this month makes a big difference when you're trying to corral these costs and make things efficient," Go said. When Cursor -- which Symbotic uses heavily -- ended a legacy discount that had grandfathered the company into a flat per-request rate even for frontier models, and moved everyone to full pricing, it forced a company-wide reckoning on efficiency, Go said. "People were saying, 'You should try this model ... This works better for this C# code, this whatever,'" he said. But the cost problem is increasingly moving out of IT; Replit, for one, broadened agents beyond engineering, and eventually found that a user on the support side had "blown through an insane amount of money," Jain said. When they looked under the hood, they figured out it was because they were running an automation on GPT 5.5 Pro Max. "At least till that point, the ROI was rather clear," Jain said. "We could see engineering productivity 3X, so no one had questioned it yet." Visibility that isn't "anti-productive," model routing, and sensible defaults are critical, he emphasized. "Most tasks do not need the frontier."
[8]
Beyond Tokenmaxxing: the rising token tax on enterprise AI
During the first wave of AI adoption, much of that cost was effectively hidden from customers. AI came wrapped in subsidized pricing, generous allowances and a relentless focus on driving usage. The message was simple: use more AI. In some organizations, usage itself has become the goal. The rise of concepts like "tokenmaxxing" took this to an extreme, celebrating volume over value and falsely equating outputs with outcomes. But while tokenmaxxing may be a questionable habit that's been widely exposed, it is not the biggest problem facing enterprise AI. The bigger issue is the hidden token tax that comes with their use of gen AI and agentic AI capabilities in everyday operations. The token tax is kicking in Enterprises are already beginning to see the impact. Uber reportedly exhausted its planned 2026 AI budget by April, just four months into the year, after rapid adoption of AI coding tools across its engineering organization. Amazon reportedly shut down an internal AI-usage leaderboard after concerns that it encouraged "tokenmaxxing," with executives urging employees not to use AI merely to increase usage metrics. As AI moves from experimentation to production, organizations are discovering that the economics of scale can look very different from the economics of experimentation. Tokenmaxxing may encourage organizations to consume more AI, but the token tax is the bill that eventually arrives. And for many enterprises, that bill is proving far larger than expected. Why is this happening? On the surface, it seems counterintuitive. Per-token pricing continues to fall, and model providers regularly announce cheaper rates. Yet with enterprise AI bills continuing to rise, the reason lies in how modern agentic systems actually work. The visible output you receive is only a small part of what is happening behind the scenes. Before generating an answer, an AI agent may interpret the request, decide which tools to use, retrieve data, evaluate results, and loop, calling tools repeatedly until the task is completed or the agent determines it is stuck. Think of it as an agent having an inner monologue: interpreting the request, planning, calling tools, checking results, and deciding what to do next. Each step may trigger additional model calls and may carry forward more context, so the visible answer can represent only a fraction of the total tokens consumed. So, while the cost per token is dropping, the number of tokens required to complete a task is often increasing dramatically. Goldman Sachs Research estimates a 24-fold increase in token consumption by 2030, reaching around 120 quadrillion tokens per month as consumers and enterprises adopt agentic technology. With this, organizations may believe they are benefiting from lower pricing while their overall costs continue to rise. And the issue is not just cost; it is also predictability, and the lack of a clear link between token usage and outcomes. This is becoming one of the biggest economic challenges facing enterprise AI, and it is only beginning to be discussed. The economics of agentic AI are making it increasingly expensive to run agents at scale. What is the token tax? We position the token tax as the hidden cost organizations incur when AI systems repeatedly consume expensive run-time agentic resources to perform work that could have been designed once and reused many times. Unlike traditional business software, where the cost of execution is largely fixed, many agentic AI systems effectively rethink the same process every time they run. The more complex the workflow, the higher the tax. Many repeatable enterprise tasks do not require, or even allow for open-ended run-time reasoning. This creates a disconnect between activity and value. Organizations may be consuming millions of tokens, but that does not necessarily translate into better outcomes. In many cases, it simply means paying repeatedly for the same agentic planning process. Are we using agents in the right place? Run-time agents are valuable when solving new, unique and underspecified problems, designing workflows, or exploring options. But using expensive reasoning to repeatedly execute the same business process introduces not just cost but also unpredictability. If you ask an agent to re-reason the same question 100 times, you may get 100 different answers. To compare this to something in the real world, it's a bit like paying a five-star gourmet chef to invent a recipe every time a new order for a meal comes in. The smarter approach is to use agentic AI once to design the optimal workflow, and then use a lighter-weight AI to select the right workflow that executes consistently. In other words, use the five-star chef's creativity and expensive brain once a season to come up with a new menu and recipes. Then use skillful but much cheaper chefs to prepare the meal following the recipe that has already been designed. Shift agentic AI left from run-time to design time This is where the distinction between design-time and run-time becomes critical. Agents can be very useful at design time, where creativity, exploration, and problem-solving create lasting value. Run-time execution is different. Here, predictability, consistency, governance and cost efficiency matter most. Run-time agents should only be used selectively, workflows should be used for deterministic, repeatable and governed processes. By using agentic AI to design processes up front, organizations can dramatically reduce the token tax associated with run-time execution while improving reliability. Sustainable AI is not about eliminating agents, but about being deliberate about where agents deliver value. The next phase of enterprise AI This next chapter will be defined less by impressive demos and more by economic discipline. Organizations will need predictable outcomes, predictable costs and a strategy for reducing the token tax embedded within agentic systems. The winners will not necessarily be the organizations that use the most AI. They will be the ones that generate the most business value from every token consumed. The companies that address this challenge early by optimizing token consumption and architecting for design-time intelligence and run-time efficiency will not just save money, but will build AI systems that are easier to trust, easier to govern, and easier to scale. As AI moves from experimentation to enterprise reality, the organizations that minimize their token tax while maximizing business outcomes will have a significant competitive advantage. We've featured the best IT automation software. This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today. The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit
[9]
Atlassian tightens tracking of staff AI use as other technology firms encourage 'tokenmaxxing'
Some companies reportedly using leaderboards for employees who used the most AI in their work Software firm Atlassian has sought to tighten tracking of its staff's AI spending by introducing "wallets" with monthly caps of up $2,000 for each employee, amid an explosion in costs at other tech companies. The move by Atlassian, which recently cited AI as part of the reason behind cutting 1,600 staff, bucks the trend of others in the tech sector who encouraged employees to use as much of the technology as possible, dubbed "tokenmaxxing". Some companies have reportedly introduced leaderboards for employees who used the most AI in their work. Tokens refer to the measurement of a response AI gives to a prompt. OpenAI has said one token is about four characters, and something like the US Declaration of Independence amounts to 1,695 tokens. OpenAI's flagship model, GPT-5.6 Sol, has a charge of US$5 for every 1m tokens, while Anthropic's Claude Fable and Mythos models are US$10 for every 1m tokens. Under tokenmaxxing, the costs quickly add up. Uber reportedly blew through its AI budget in four months, and Amazon has reportedly told employees to stop using AI just for the sake of using AI. While Atlassian never encouraged tokenmaxxing or had unlimited AI budgets, the Australian company this month introduced an "AI wallet" for staff in the research and development team. According to an internal memo seen by Guardian Australia, employees have between $500 and $2,000 monthly spend in the wallet, which can be used across four AI products, including Claude Code. Employees receive notifications as they approach their limit on their wallets, and usage is paused when the money runs out. Employees can request additional funds. It's understood Atlassian has not turned down any request for additional funds so far. A company spokesperson said it was transforming into an "AI-first company" by supporting people building and experimenting with the technology. "Atlassian provides a significant budget for our builders to leverage multiple AI tools," the spokesperson said. "AI tooling budgets are set by role based on how different teams work." They said the wallet also represented a boost in the amount employees could spend. A June PureProfile survey of 500 senior Australia staff at companies using AI, conducted on behalf of search AI company Elastic, found that 80% were concerned that high usage was being mistaken for productivity gains. It found 32% had reported pausing, cancelling or winding back AI deployments, due to cost. Elastic's ANZ manager, Jeremy Pell, said a monthly cap on AI spend was "smart" and more organisations should be doing it. "Right now, only 9% of Australian organisations currently have any limits on token or API consumption when it comes to AI agents or autonomous workflows, so any organisation that adopts this practice is an outlier," he said. Arun Chandrasekaran, a distinguished vice-president analyst at research firm Gartner, said wallets were a simple way to "incentivise the right behaviour" and stop people from using AI ineffectively. He said such expenses were becoming an important issue among businesses, with the cost explosion being driven by AI agents that autonomously undertake tasks on behalf of the user. "You suddenly have these systems that are all trying to do independent tasks that are spawning smaller agents, that are creating their own prompts and initiating requests for the model," he said. "So while the AI model prices have been falling for the last three years, the volume of tokens that particularly the AI agents are starting to send to the models ... is significantly increasing." Chandrasekaran said companies are figuring out how to drive the cost down, including using less powerful models for simpler tasks, and looking at using open weight - where people can download and run the models on their own systems - or open source models.
[10]
BNY skipped the 'tokenmaxxing' craze. Here's what AI metrics it tracks instead | Fortune
Good morning. Tokenmaxxing quickly became one of the buzziest metrics in enterprise AI. Fortune's Jeremy Kahn reported that "tokenmaxxing" turned into a status symbol at some big tech companies, where engineers were urged to climb leaderboards by burning more AI tokens. He argues that the practice skewed incentives and exposed a broader gap between AI spending and actual productivity gains. I recently spoke with Dermot McDonogh, the CFO of BNY, which is making major strides with AI. While some companies track success by the volume of prompts, tokens, or agents deployed, McDonogh said that framing never took hold inside BNY. "It's not something we spend any time talking about," he told me, noting that token costs are "modest within modest" relative to the firm's broader engineering budget. Even as the topic gained traction externally, the bank's leadership prepared to address it -- but ultimately viewed it as a distraction from more meaningful measures of value. McDonogh said that BNY had an early and deliberate AI strategy. Since the emergence of ChatGPT, the bank has spent several years building an internal, LLM-agnostic platform and forging partnerships across hyperscalers and model providers. Just as important, he said, has been CEO-level commitment and a focus on cultural adoption. "There's been a demystification," McDonogh said. "People don't feel insecure about AI. That's a really important cultural point." That approach has allowed BNY to scale AI without fixating on cost per query. Internally, systems route tasks to the appropriate models, ensuring efficiency without requiring employees to optimize prompts manually. "I couldn't tell you how many prompts we did last week," he said. "I'm focused more on outcomes." Those outcomes are increasingly measurable. In the first quarter of 2026, more than 40% of BNY's code was authored by AI, rising to roughly 50% more recently. AI is also embedded across operations: about half of annual account plans are drafted with AI, 25% of client onboarding is AI-supported, and roughly 70% of restricted-party payment screening is reviewed by AI. The impact is showing up in financial metrics. Revenue per employee rose from $338,000 in 2022 to $401,000 in 2025, while pre-tax income per employee increased from $99,000 to $143,000 over the same period. McDonogh frames these gains less as cost savings and more as capacity creation. "We haven't reduced the footprint, but it's allowed us to do more with the footprint that we have," he said. To track progress, BNY measures AI impact across core workflows -- including innovating, prospecting, onboarding, transacting, and streamlining -- while continuously building out its internal "Eliza" platform. The system serves as a firm-wide context layer, improving over time as it ingests more data and use cases. Employee adoption is also structured. Staff progress through three levels of AI proficiency, culminating in a "pioneer" designation that requires formal training and testing. Access to more advanced models is gated by expertise, reinforcing both quality and accountability. Within finance specifically, AI is already reshaping core processes. McDonogh points to regulatory reporting, balance sheet analytics and predictive modeling as key use cases. The technology is also playing a growing role in earnings preparation, helping synthesize analyst expectations and anticipate investor questions. For McDonogh, the takeaway is straightforward: AI productivity is not about how much you use, but how effectively it changes what an organization can do.
[11]
The AI 'tokenmaxxing' corporate fad is fading as workplaces look to cut costs
Tech executives cast high AI usage as a badge of honor Just a few months ago, Silicon Valley executives were promoting high token consumption as a signal of high-performing employees. The stereotypical tokenmaxxer was staying up late -- perhaps ignoring their significant other -- while orchestrating an army of 24-hour AI agents performing work on their behalf. OpenAI CEO Sam Altman said in May he was "excited to see what will happen with tokenmaxxing startups, both for how they work internally and the products they can build." Nvidia CEO Jensen Huang said, "If your $500K engineer isn't burning $250K in tokens, something is wrong." Facebook parent Meta had an internal competition rewarding token usage. The trend boosted revenue for leading AI large language model developers like Anthropic and OpenAI, but it fizzled as it became apparent it wasn't necessarily the best strategy for everyone else. Microsoft CEO Satya Nadella has admitted that tokenmaxxing can be addictive but warned in a recent blog post that customers of those models are paying twice for AI, first in spending on tokens and second by feeding all their proprietary data to them. While promoting Microsoft's own approach, Nadella's comments were unusual in the way he raised doubts about the data protection assurances of leading AI providers. Palantir CEO Alex Karp went further, telling CNBC earlier this month that something had gone "completely wrong." He said he was channeling the voice of American businesses privately "livid" about paying so much for tokens that create no value. "The basic view among enterprises in this country is, 'I'm going to chillax and waste my time with tokens. I'm going to get no value and they're going to get my IP,'" Karp said. Workplaces look more for better 'routing' of their AI work Bain & Company management consultant Jue Wang said many of the big businesses her firm advises have been taking a closer look at returns on their AI investments. "The token cost for them has been doubling, almost every other month," she said. "Let's say $200 per developer per month. Multiply that by 20,000 developers, which is often what we're dealing with at these companies, and that quickly gets you to a number that is not a line item that any general manager has planned for." Sometimes that just means not using the AI equivalent of a sledgehammer to crack a nut. "Not everything needs a Claude Opus 4.6," she said of one of Anthropic's more capable models suited to software engineering or deep research. "And yet you see so many companies, so many users, default to using Opus for everything, including generating emails." That's led to a search for tools that do AI "model routing" -- in which easier queries get automatically sent to cheaper and more efficient AI systems and more complex tasks go to more powerful models. Open-source AI models built in China offer less costly alternatives Software developer Hassan El Mghari said companies' sticker shock over the "ridiculous amount of money" spent on subscriptions to AI products from leading U.S. companies has led many away from rewarding high usage. "It's better to kind of just empower employees on how to use this stuff and let them use AI when and however much they need to," said El Mghari, who leads developer experience at the startup Together AI, which supplies developers with a variety of "open-source" AI models. At the same time, those who favor racking up as many tokens as possible are having a field day with new open-source models from Chinese startups like Moonshot's Kimi or Zhipu's GLM, which nearly match the capabilities of top U.S. models at a fraction of the price. "There is some validity to the theory that this could push tokenmaxxing a little bit further," said Raffi Krikorian, the chief technology officer at Mozilla. "But if we look at the industry overall, I think it's realizing that tokenmaxxing is a dumb thing." It's similar, Krikorian said, to how software companies once considered how many lines of code a programmer wrote to be a good metric of productivity. That later fell out of favor. "I think tokenmaxxing is moving through the exact same pattern," he said. "I think this is going to be an interesting blip that we're all going to look back to laugh at in a year."
[12]
A flex in corporate America, AI 'tokenmaxxing' fades as workplaces look to cut tech spending
A corporate fad of "tokenmaxxing" on artificial intelligence technology is hitting its limits as workplaces throwing AI at everything are seeing the costs rise without a similar spike in productivity. What started as tech industry-fueled springtime hype over squeezing as much AI-generated work as possible out of products like OpenAI's ChatGPT and Anthropic's Claude has shifted to a summertime backlash. "It's very easy to create something you don't need with AI," said Vincent Gusdorf, head of AI analytics at Moody's Ratings and author of a new report that recommends a more disciplined approach. "Tokenmaxxing" refers to maximizing usage of tokens -- the building blocks of generative AI that correspond to small pieces of text that an AI system reads or writes. Each token is about three quarters of a word. And there's typically a limit to how many you can use, with pricier versions of AI products offering higher caps. "As bills started to pile in, people realized that those new tools are quite expensive and you need to use them wisely," Gusdorf said. Tech executives cast high AI usage as a badge of honor Just a few months ago, Silicon Valley executives were promoting high token consumption as a signal of high-performing employees. The stereotypical tokenmaxxer was staying up late -- perhaps ignoring their significant other -- while orchestrating an army of 24-hour AI agents performing work on their behalf. OpenAI CEO Sam Altman said in May he was "excited to see what will happen with tokenmaxxing startups, both for how they work internally and the products they can build." Nvidia CEO Jensen Huang said "if your $500K engineer isn't burning $250K in tokens, something is wrong." Facebook parent Meta had an internal competition rewarding token usage. The trend boosted revenue for leading AI large language model developers like Anthropic and OpenAI, but it fizzled as it became apparent it wasn't necessarily the best strategy for everyone else. Microsoft CEO Satya Nadella has admitted that tokenmaxxing can be addictive but warned in a recent blog post that customers of those models are paying twice for AI, first in spending on tokens and second by feeding all their proprietary data to them. While promoting Microsoft's own approach, Nadella's comments were unusual in the way he raised doubts about the data protection assurances of leading AI providers. Palantir CEO Alex Karp went further, telling CNBC earlier this month that something had gone "completely wrong." He said he was channeling the voice of American businesses privately "livid" about paying so much for tokens that create no value. "The basic view among enterprises in this country is, 'I'm going to chillax and waste my time with tokens. I'm going to get no value and they're going to get my IP," Karp said. Workplaces look more for better 'routing' of their AI work Bain & Company management consultant Jue Wang said many of the big businesses her firm advises have been taking a closer look at returns on their AI investments. "The token cost for them has been doubling, almost every other month," she said. "Let's say $200 per developer per month. Multiply that by 20,000 developers, which is often what we're dealing with at these companies, and that quickly gets you to a number that is not a line item that any general manager has planned for." Sometimes that just means not using the AI equivalent of a sledgehammer to crack a nut. "Not everything needs a Claude Opus 4.6," she said of one of Anthropic's more capable models suited to software engineering or deep research. "And yet you see so many companies, so many users, default to using Opus for everything, including generating emails." That's led to a search for tools that do AI "model routing" -- in which easier queries get automatically sent to cheaper and more efficient AI systems and more complex tasks go to more powerful models. Open-source AI models built in China offer less costly alternatives Software developer Hassan El Mghari said companies' sticker shock over the "ridiculous amount of money" spent on subscriptions to AI products from leading U.S. companies has led many away from rewarding high usage. "It's better to kind of just empower employees on how to use this stuff and let them use AI when and however much they need to," said El Mghari, who leads developer experience at the startup Together AI, which supplies developers with a variety of "open-source" AI models. At the same time, those who favor racking up as many tokens as possible are having a field day with new open-source models from Chinese startups like Moonshot's Kimi or Zhipu's GLM, which nearly match the capabilities of top U.S. models at a fraction of the price. "There is some validity to the theory that this could push tokenmaxxing a little bit further," said Raffi Krikorian, the chief technology officer at Mozilla. "But if we look at the industry overall, I think it's realizing that tokenmaxxing is a dumb thing." It's similar, Krikorian said, to how software companies once considered how many lines of code a programmer wrote to be a good metric of productivity. That later fell out of favor. "I think tokenmaxxing is moving through the exact same pattern," he said. "I think this is going to be an interesting blip that we're all going to look back to laugh at in a year."
[13]
Workplaces Look for Cheaper AI as 'Tokenmaxxing' Fades as a Corporate Fad
A corporate fad of "tokenmaxxing" on artificial intelligence technology is hitting its limits as workplaces throwing AI at everything are seeing the costs rise without a similar spike in productivity. What started as tech industry-fueled springtime hype over squeezing as much AI-generated work as possible out of products like OpenAI's ChatGPT and Anthropic's Claude has shifted to a summertime backlash. "It's very easy to create something you don't need with AI," said Vincent Gusdorf, head of AI analytics at Moody's Ratings and author of a new report that recommends a more disciplined approach. "Tokenmaxxing" refers to maximizing usage of tokens -- the building blocks of generative AI that correspond to small pieces of text that an AI system reads or writes. Each token is about three quarters of a word. And there's typically a limit to how many you can use, with pricier versions of AI products offering higher caps. "As bills started to pile in, people realized that those new tools are quite expensive and you need to use them wisely," Gusdorf said. Tech executives cast high AI usage as a badge of honor Just a few months ago, Silicon Valley executives were promoting high token consumption as a signal of high-performing employees. The stereotypical tokenmaxxer was staying up late -- perhaps ignoring their significant other -- while orchestrating an army of 24-hour AI agents performing work on their behalf. OpenAI CEO Sam Altman said in May he was "excited to see what will happen with tokenmaxxing startups, both for how they work internally and the products they can build." Nvidia CEO Jensen Huang said "if your $500K engineer isn't burning $250K in tokens, something is wrong." Facebook parent Meta had an internal competition rewarding token usage. The trend boosted revenue for leading AI large language model developers like Anthropic and OpenAI, but it fizzled as it became apparent it wasn't necessarily the best strategy for everyone else. Microsoft CEO Satya Nadella has admitted that tokenmaxxing can be addictive but warned in a recent blog post that customers of those models are paying twice for AI, first in spending on tokens and second by feeding all their proprietary data to them. While promoting Microsoft's own approach, Nadella's comments were unusual in the way he raised doubts about the data protection assurances of leading AI providers. Palantir CEO Alex Karp went further, telling CNBC earlier this month that something had gone "completely wrong." He said he was channeling the voice of American businesses privately "livid" about paying so much for tokens that create no value. "The basic view among enterprises in this country is, 'I'm going to chillax and waste my time with tokens. I'm going to get no value and they're going to get my IP," Karp said. Workplaces look more for better 'routing' of their AI work Bain & Company management consultant Jue Wang said many of the big businesses her firm advises have been taking a closer look at returns on their AI investments. "The token cost for them has been doubling, almost every other month," she said. "Let's say $200 per developer per month. Multiply that by 20,000 developers, which is often what we're dealing with at these companies, and that quickly gets you to a number that is not a line item that any general manager has planned for." Sometimes that just means not using the AI equivalent of a sledgehammer to crack a nut. "Not everything needs a Claude Opus 4.6," she said of one of Anthropic's more capable models suited to software engineering or deep research. "And yet you see so many companies, so many users, default to using Opus for everything, including generating emails." That's led to a search for tools that do AI "model routing" -- in which easier queries get automatically sent to cheaper and more efficient AI systems and more complex tasks go to more powerful models. Open-source AI models built in China offer less costly alternatives Software developer Hassan El Mghari said companies' sticker shock over the "ridiculous amount of money" spent on subscriptions to AI products from leading U.S. companies has led many away from rewarding high usage. "It's better to kind of just empower employees on how to use this stuff and let them use AI when and however much they need to," said El Mghari, who leads developer experience at the startup Together AI, which supplies developers with a variety of "open-source" AI models. At the same time, those who favor racking up as many tokens as possible are having a field day with new open-source models from Chinese startups like Moonshot's Kimi or Zhipu's GLM, which nearly match the capabilities of top U.S. models at a fraction of the price. "There is some validity to the theory that this could push tokenmaxxing a little bit further," said Raffi Krikorian, the chief technology officer at Mozilla. "But if we look at the industry overall, I think it's realizing that tokenmaxxing is a dumb thing." It's similar, Krikorian said, to how software companies once considered how many lines of code a programmer wrote to be a good metric of productivity. That later fell out of favor. "I think tokenmaxxing is moving through the exact same pattern," he said. "I think this is going to be an interesting blip that we're all going to look back to laugh at in a year."
[14]
Finance Teams Are Done Flying Blind on AI Costs | PYMNTS.com
That approach worked while AI spending was small enough to absorb without much scrutiny. It no longer is. Enterprise software once ran on annual licenses and seat-based pricing that finance teams could forecast with reasonable accuracy, PYMNTS reported. AI, priced in tokens, compute cycles and application programming interface (API) calls, has broken that model open. New tools are being launched to address the issue. Ramp launched AI Token Spend Management on July 16, giving finance teams a single dashboard to track, allocate and control AI spending across providers including OpenAI, Anthropic, Gemini and Cursor, Ramp said in its announcement. AI token spend across Ramp's own customer base increased 20.7x since June 2025, according to the announcement. Ramp developed the product using usage data from more than 1,300 businesses, and after analyzing 110 trillion tokens, identified three distinct spending profiles based on which models companies chose, how they used caching and how tightly they controlled spend, according to Ramp's product launch. Companies Are Shifting From Maximizing Usage to Maximizing Value Per Token The mechanics of the problem explain why a new software category is forming around it. Unlike seat-based software, AI token spend scales directly with usage and can grow quickly across teams without anyone approving a specific purchase, often sitting behind provider dashboards and invoices finance teams find difficult to interpret. At one Ramp customer, Sansa Services, an employee left an expensive "Fast Mode" setting active without realizing it, driving up token spend before the mistake was caught. "We flagged that somebody was using Fast Mode unnecessarily and that resulted in six times the token spend over a seven-day period. They've since turned it off. That is one story in which basically the product pays for itself," said Neusha Sayadian, founder and fractional CFO at Sansa Services, Ramp reported on its product page. As of June, Ramp found that the average business could identify potential savings equal to 12% of its monthly AI spend, and one in three businesses found a lower-cost model alternative capable of doing the same work. CloudZero is pursuing the same shift from a different entry point. Rather than simply tracking token counts, CloudZero's financial control platform connects AI spending to the specific customers, features and teams that generated it, aiming to point to what that spending actually produced, CloudZero said in its own launch announcement. "AI is becoming central to how companies build products, serve customers, and run the business," CloudZero Chief Product Officer Scott Castle said. "But most companies still manage AI with token counts and monthly invoices, even as costs rise faster than expected and ROI remains unclear." Finance Teams Want Agents That Manage AI Spending The demand for this category is showing up directly in how CFOs prioritize agentic AI itself. Dynamic budget reallocation using real-time cost data is the highest-ranked use case for agentic AI among CFOs surveyed by PYMNTS Intelligence, with 43% expecting agents that continuously scan spending patterns, flag overruns and shift funds toward higher-priority areas to have a significant impact. That figure captures the shift underway: finance teams no longer just want visibility into AI spending after the fact. They want a system that actively manages it in real time. Companies are betting that the next phase of enterprise AI adoption depends less on pushing employees toward heavier AI usage and more on building the financial infrastructure to prove that usage is worth what it costs. For all PYMNTS AI and digital transformation coverage, subscribe to the daily AI and Digital Transformation Newsletters.
Share
Copy Link
Businesses are experiencing a dramatic shift from encouraging unlimited AI token usage to implementing strict controls as bills skyrocket. Major firms report 10x increases in AI spending, with some burning through annual budgets in months. Companies like EY are deploying AI routers to manage costs while the industry develops tokenomics frameworks.
Corporate America's enthusiasm for artificial intelligence hit a wall when the bills arrived. Companies that spent the past year encouraging employees to maximize AI token usage—a practice dubbed "token-maxing"—are now scrambling to control spiraling token costs as AI spending threatens to consume entire technology budgets
1
3
. Steve Lucas, CEO at Boomi, reported personally spending 10 times more on Claude in one year compared to the previous year, calling the trajectory "not sustainable"1
. The shift from tokenmaxxing to what some now call "tokenminimizing" reflects a fundamental reckoning with AI token usage patterns across enterprises3
.
Source: NYT
The rise of agentic AI systems has transformed token costs from predictable to chaotic. When businesses deploy multiple AI agents together to make decisions and automate processes, token consumption increases dramatically and unpredictably
2
. Unlike simple chatbot interactions, agentic workflows can consume orders of magnitude more tokens as agents iterate through complex tasks. Microsoft reportedly reined in engineers' use of third-party coding tools, while Uber burned through its entire annual AI coding token budget in just months2
. Goldman Sachs forecasts that token consumption will increase 24 times between 2026 and 2030 to 120 quadrillion tokens monthly as companies shift to agentic AI adoption2
. The fundamental challenge lies in the non-deterministic nature of Large Language Models—subtle variations in prompts produce different answers, making cost prediction nearly impossible2
.EY, which invests over $1 billion annually in AI and operates approximately 1,000 AI agents, built an "AI router" to steer tasks to the cheapest model capable of handling each request
4
. The system represents sophisticated cost-conscious AI deployment, reserving expensive models for complex problems while routing simpler tasks to cheaper alternatives. This approach addresses a paradox in token economics: while the price per token has collapsed, enterprise AI spending has tripled because agentic tools running multi-step processes consume far more tokens than single chatbot prompts4
. EY's own research revealed that 82% of 534 senior US business leaders surveyed expressed concern about AI token usage costs, with 98% of those using token-based tools reconsidering their AI pricing strategy4
. Critically, only 64% of firms actively monitor token usage with budgetary guardrails, meaning a third are spending on AI without clear meters4
.
Source: PYMNTS
The complexity of AI spending has spawned "tokenomics"—the practice of measuring, pricing, and managing token consumption as a distinct business activity
1
3
. Howard Rubin, an economist advising companies on technology spending, described tokens as "a currency where you have no instinct to know what you're using, and the accounting practices aren't even there for it"3
. The Linux Foundation established the Tokenomics Foundation in June to create common parameters that AI providers agree to disclose, standardizing frameworks for organizational decision-making3
. Companies are turning to specialized service providers like Revenium to track whether token spending improves productivity enough to justify costs, mapping token consumption to specific outcomes like software features shipped3
.Related Stories
OpenAI CEO Sam Altman envisions a future where "intelligence is a utility, like electricity or water, and people buy it from us on a meter"
5
. If this vision materializes, AI tokens could become the defining measure of a new industrial age—the kilowatt-hour of the AI age for measuring and paying for AI adoption5
. Tokens represent the tiny chunks of text and data that AI models process, with more complex work typically requiring more tokens. While AI companies like OpenAI, Microsoft, and Google still offer flat-rate subscriptions to consumers, they increasingly charge businesses based on token usage5
. Researchers are already using token data to track AI's economic impact, with economists analyzing 380 trillion AI tokens to understand how AI consumption reshapes financial markets and creates an "AI premium" for certain stocks5
.
Source: NPR
Despite mounting concerns about AI spending, business leaders emphasize the importance of giving staff room to explore AI agents within structured guidelines. Snowflake CEO Sridhar Ramaswamy acknowledged concerns about AI spending at Summit 2026 but insisted the investment creates opportunities to transform processes into products for customers
1
. Matt Luizzi, VP of analytics at Whoop, described implementing guardrails and monitoring while enabling staff to take risks, stating "we're leaning in hard because we think there's an ROI to this"1
. The fundamental question facing enterprises, according to Lucas, is "Can I operate AI at a return?"1
. Companies are finding answers by cutting software licenses and outside services to fund AI spending, with some scrounging resources before making strategic reallocation decisions3
. As EY's global AI consulting leader Dan Diasio noted, "'AI saves time' is no longer sufficient when costs mount and remain unclear"4
, marking a decisive shift from adoption at any price to value at a known cost.Summarized by
Navi
[4]
17 Jun 2026•Business and Economy

24 Jun 2026•Business and Economy

29 Apr 2026•Business and Economy

1
Technology

2
Technology

3
Technology
