6 Sources
[1]
Token-maxing is an AI cost sink - how to use agents without busting your budget
Follow ZDNET: Add us as a preferred source on Google. ZDNET's key takeaways * Token-maxing to support agentic AI is unsustainable. * Business leaders must create a strategy for token use. * Give people room to explore agents within guidelines. Twelve months in AI is an eternity. Steve Lucas, CEO at integration technology specialist Boomi, was concerned last year that CIOs were rushing into gen AI initiatives without a clear sense of direction. Now, a year later, he's concerned that IT professionals are taking a similarly rushed approach with agentic technology, and the scale of token usage is only going one way: upward. "Whether you work inside or outside a company, it feels like your core hustle is to use AI and you token max the heck out of that technology for your own job," he said to ZDNET. Also: The 3 types of people who will excel in the AI agent era, according to tech leaders A token is the fundamental unit of information processed by an AI model, and research points to the soaring and unpredictable costs of agents as the technology consumes orders of magnitude more tokens. As my ZDNET colleague Steven Vaughan-Nichols discussed recently, it wasn't so long ago -- certainly during the rise of generative AI -- that everyone was excited about their token leaderboard, which showed who had the most token usage in an organization. Today, token leaderboards are obsolete because no one can afford to waste tokens. In the agentic era, token maxing is a sign of excess, and "tokenomics" -- the practice of measuring, pricing, and managing the consumption of tokens -- is a key business activity, even for a tech CEO like Lucas, who noted that tokens are being consumed at a much faster rate. "A year ago, we weren't talking about tokenomics," he said. "But last year, I personally spent at Boomi 10 times the amount on Claude that I did the previous year -- 10 times; that's not sustainable. I can't do that every year." Also: The new enterprise AI expert every company needs - and why The onus now is on businesses and professionals to create their own enterprise-ready version of tokenomics. Agentic AI is set to transform how every business operates -- and creating a cost-effective technique for exploring agents is an urgent priority, suggested Lucas. "What matters in the enterprise is ultimately the economics of AI," he said. "Most organizations will look to AI as the enterprise engine of the future. So, what matters now is, 'Can I operate AI at a return?' That is the fundamental question." Business leaders suggested the best way forward is clear: Rather than constraining how people can use agents, give them guidelines and the context to make cost-effective model decisions. Letting people shine Snowflake CEO Sridhar Ramaswamy recognized that token consumption is on the rise, even in his own organization. However, he also said it's critical to give staff the wriggle room to explore agents. "Are we worried about how much we are spending on AI inference across our different internal teams? Absolutely," he said. "But do I see that spend as a reason not to use AI? Absolutely not." While token maxing is a concern, Ramaswamy said during a media session at his firm's recent Summit 2026 event in San Francisco that the creative use of AI can help unlock new business opportunities. Also: 40% of enterprises will scrap AI agents - 3 ways to ensure yours don't fail Agents can help boost staff efficiency and productivity, with the potential to generate new services for his firm's clients faster and more effectively. "We look at agentic AI as an opportunity to optimize not just what we do, but how we can turn a process into a product that our customers can use," he said. Like Ramaswamy, Matt Luizzi, VP of analytics at wearable technology specialist Whoop, said his company is investing in agentic explorations to find a competitive advantage. "If we want people to push themselves out of their comfort zones, we're going to need to be OK with them taking risks and understanding that you can't break anything." Also: AI is causing cognitive fatigue. Here's how to work with more haste and less speed In short, Luizzi said that he's OK with staff spending money on agents, but that doesn't mean token consumption gets out of hand. "We have guardrails in place, and monitoring and observability to let people know what they're spending," he said during a panel session at Summit 2026. "But more likely than not, we just need to sit down and enable them on, 'How could you be doing this more efficiently and what are you trying to accomplish?'" Putting everything into context The key to agentic success, suggested Luizzi, is carefully managed enablement. Yes, let people push the boundaries with agents and tokens, but don't let them go overboard. "At the end of the day, we're leaning in hard because we think there's an ROI to this. We are seeing people become more efficient. We are seeing work get done faster," he said. Also: How Workday and other software providers plan to survive AI "Those positive results mean we're not trying to be super keen on how many tokens people use but also understanding that this is something that needs to scale, so not enabling people to go crazy either." Sriram Sitaraman, CIO at technology specialist Synopsys, said in the same panel session at the Snowflake conference that people must be able to explore new avenues, including agents, to discover innovative solutions to business challenges: "You can't innovate with constraints; you can only go so far." Like Luizzi, Sitaraman said guidelines can help set tight boundaries for token consumption. He suggested that the right context for projects can help establish effective constraints. "If you ask an agent, 'What should I do tomorrow?' it might consume a whole lot of data, a whole lot of tokens, and come back with something that isn't useful," he said. "If you set a context, such as creating a North America sales ops agent, giving the project specific objectives and data, then the consumption is going to be limited, and the value that it delivers to the user is higher." Also: How to beat the AI algorithm and get the job of your dreams That focus on business context was echoed by Francois-Xavier Pierrel, group chief data and ad tech officer at French TV network TF1, who told ZDNET that a common-sense approach is the best way for business and professional users to control token use. Pierrel compared the alternative approach to a sugar addiction. Once you begin adding sugar to your food, you get used to it and start to develop an unhealthy addiction. He said AI presents similar risks. "Everything was cheap, very cheap, and we got all into it," he said, looking back at the early rush to use generative AI technologies. Now, agent obsession represents a more costly proposition. "The number of tokens used can quickly become a big number. And at some point, if we look at all the companies investing crazy money into agentic AI, the bill will come back." Also: Forget productivity: Here are 5 strategic shifts that drive real AI value Pierrel said caution is the watchword in his organization, and he gave an example of how professionals are asked to consider their options when using large language models. "For use cases that are very narrow or have strong boundaries, we say, 'Can we use a small model, something that will consume fewer tokens than expected or planned, and that will still do the same job?'" As companies look to deploy more agents as part of working processes, Pierrel suggested this kind of tokenomics strategy will be crucial to value creation. Also: AI agents are your new colleagues - how to get the best results "I think we are going to be more agile on this one, trying stuff, and then adapting to the use case to avoid having crazy bills," he said. "We're not shooting a rocket to the moon. So, let's be reasonable -- don't deploy a million agents if you can do it on a smaller scale and it will still work."
[2]
AI tokens could become the kilowatt-hour of the AI age
Earlier this year, OpenAI CEO Sam Altman declared, "We see a future where intelligence is a utility, like electricity or water, and people buy it from us on a meter." Many other AI companies seem to be betting on a similar future. If that vision actually comes to be, then "AI tokens" may break out of the world of nerdy tech and econ conversations and become a much more familiar part of our lives. The number of AI tokens businesses and consumers use could even become one of the defining measures of a new industrial age -- the AI equivalent of the kilowatt-hour for electricity: the standard way we measure and pay for AI usage. Think of tokens as the meter running in the background every time you use AI. Every time an AI model reads your prompt, writes an answer, or does a task, that work is measured in tokens. They are the tiny chunks of text and other data that AI models read and generate. In general, the more work a model does, the more tokens it typically processes. While AI companies still tend to offer flat-rate subscriptions to average consumers, they're increasingly charging businesses and developers based on the number of tokens they use. At the same time, companies, particularly in the tech sector, have been using AI in more and more of their work. As their AI usage has soared, many businesses have discovered just how expensive token-based pricing can become. After a period when tech workers were engaged in a kind of AI free-for-all -- which some dubbed "tokenmaxxing" -- companies like Uber and Amazon have been putting guardrails around AI use and reducing their soaring token bills (which one clever writer at The Information recently dubbed "tokenminimizing"). But tokens aren't just the way AI companies measure usage and charge many of their business customers. They also leave behind a kind of digital paper trail that a growing number of economists and other researchers are using to track AI usage and study its economic impact. In a new working paper, Nicola Borri, Aleh Tsyvinski, and Yukun Liu do basically that. Using data from 380 trillion AI tokens, these economists try to understand how the growth of AI usage is reshaping financial markets. They ask a simple question: as overall AI consumption changes over time, which companies' stock prices tend to rise with it -- and which tend to fall? To be clear, the economists aren't tracking how many AI tokens individual companies are using -- which, let's be honest, would be way cooler. Instead, they're looking at the rising tide of overall AI usage and asking which stocks rise and fall with it. Part of their analysis is classic finance theory. The idea is that stock prices tend to reflect investors' expectations about the future, so they can offer a window into which companies Wall Street sees as most positively or negatively impacted by the rise of AI. Of course, investors can be spectacularly wrong. Hello, the long history of financial bubbles. That said, if you want to know what markets expect AI to do to the economy, stock prices are a compelling place to look. Not surprisingly, the economists find that as AI usage has grown, financial markets have treated some companies very differently than others. The companies seen as the biggest AI beneficiaries have enjoyed higher stock returns -- a pattern the researchers call an "AI premium." More interestingly, they find it's not just tech stocks that appear to earn that premium. Their findings suggest investors expect AI to benefit a wide range of companies and industries across the economy. "The story of AI is no longer just a Silicon Valley story," Tsyvinski says about their paper's findings. "Financial markets already see Main Street being impacted." What's most exciting about this study is probably less its findings and more that it offers a glimpse into future research possibilities. Instead of relying on surveys, earnings calls, or company announcements, economists may be able to study the spread of AI by following the digital paper trail created by AI tokens. AI tokens could become a powerful new source of data, allowing researchers to track AI usage in almost real time and study its economic effects, with a kind of precision not possible in past technological revolutions. An exciting way to measure the AI economy Borri, Tsyvinski, and Liu's analysis relies on a relatively new source of data: OpenRouter. OpenRouter is a kind of one-stop shop for AI models. Instead of contracting separately with OpenAI, Anthropic, Google, and dozens of other AI companies, developers can access hundreds of models through a single interface. As businesses and workers have become more conscious of their mounting AI token bills, they've turned to platforms like OpenRouter to compare prices, switch to cheaper models for less complex tasks, and manage their token spending. In the process, OpenRouter has amassed rich data on usage of AI tokens (and this data is anonymized to protect users' identities). The economists analyze the use of 380 trillion AI tokens between January 2024 and April 2026. That represents around 2 percent of monthly global AI usage. They then combine weekly growth in tokens, spending, and active users into a broad measure of AI consumption, which they call the "AI Factor." Next, they estimate which companies' stock returns move most strongly with changes in that factor. The economists find that companies whose stock prices were most sensitive to increases in overall AI consumption subsequently earned significantly higher returns. The companies that Wall Street appears to view as the biggest beneficiaries of AI outperformed those viewed as the least likely beneficiaries by about 0.64 percentage points per week -- the "AI Premium" described in the paper. Of course, markets aren't always right. These AI-related stock movements could reflect the wisdom of crowds -- or simply a herd of excited investors hoping to cash in on a technological revolution that ultimately disappoints. In other words, take these findings with a grain of salt. Probably the most interesting of their findings is that the AI premium can be found well beyond the tech world. Markets seem to believe that companies in industries ranging from airlines and cruise lines to utilities, industrial manufacturers, retailers, banks, and even waste management companies could all benefit as AI reshapes the economy. They also find that this "AI premium" is strongest for companies in the United States and Europe, and is much weaker in China and other emerging markets. Perhaps relatedly, they find that stocks are the most sensitive to increased use of the most technologically advanced, "frontier" AI models. The economists even provide a list of S&P 500 companies with high AI premiums in their paper. The top five are: AppLovin, Carvana, Lumentum, Expand Energy, and Baker Hughes. They also identify companies Wall Street appears to view as AI's biggest losers. The bottom five are: Moderna, Estée Lauder Companies, ON Semiconductor, Skyworks Solutions, and Aptiv. There are a bunch of caveats to their analysis. First, this is a working paper that has yet to be peer reviewed. Second, OpenRouter's token data likely provides a skewed picture of AI usage. People who use OpenRouter are likely sophisticated, heavy users of AI who are trying to save money as they shop between competing models to get the best bang for their buck. Their data doesn't represent the sort of average consumer, who may have a ChatGPT or Claude or Gemini monthly subscription and sticks with using those. Perhaps most importantly, even if the researchers have correctly identified the companies investors expect to benefit from AI, that doesn't mean those expectations will prove right. And if markets are reasonably efficient, much of that optimism may already be reflected in today's stock prices. So buying the stocks they identify as potential AI-related winners and selling the ones that could be AI-related losers is by no means a good financial strategy. This is not financial advice! Still, the paper's biggest contribution may not be what it says about today's stock market. It may be that it points toward a new way for economists to measure the spread of AI through the economy. We welcome this era of AI measurement-maxxing. There are still many big economic questions about AI, and better data is rarely a bad place to start.
[3]
BNY skipped the 'tokenmaxxing' craze. Here's what AI metrics it tracks instead | Fortune
Good morning. Tokenmaxxing quickly became one of the buzziest metrics in enterprise AI. Fortune's Jeremy Kahn reported that "tokenmaxxing" turned into a status symbol at some big tech companies, where engineers were urged to climb leaderboards by burning more AI tokens. He argues that the practice skewed incentives and exposed a broader gap between AI spending and actual productivity gains. I recently spoke with Dermot McDonogh, the CFO of BNY, which is making major strides with AI. While some companies track success by the volume of prompts, tokens, or agents deployed, McDonogh said that framing never took hold inside BNY. "It's not something we spend any time talking about," he told me, noting that token costs are "modest within modest" relative to the firm's broader engineering budget. Even as the topic gained traction externally, the bank's leadership prepared to address it -- but ultimately viewed it as a distraction from more meaningful measures of value. McDonogh said that BNY had an early and deliberate AI strategy. Since the emergence of ChatGPT, the bank has spent several years building an internal, LLM-agnostic platform and forging partnerships across hyperscalers and model providers. Just as important, he said, has been CEO-level commitment and a focus on cultural adoption. "There's been a demystification," McDonogh said. "People don't feel insecure about AI. That's a really important cultural point." That approach has allowed BNY to scale AI without fixating on cost per query. Internally, systems route tasks to the appropriate models, ensuring efficiency without requiring employees to optimize prompts manually. "I couldn't tell you how many prompts we did last week," he said. "I'm focused more on outcomes." Those outcomes are increasingly measurable. In the first quarter of 2026, more than 40% of BNY's code was authored by AI, rising to roughly 50% more recently. AI is also embedded across operations: about half of annual account plans are drafted with AI, 25% of client onboarding is AI-supported, and roughly 70% of restricted-party payment screening is reviewed by AI. The impact is showing up in financial metrics. Revenue per employee rose from $338,000 in 2022 to $401,000 in 2025, while pre-tax income per employee increased from $99,000 to $143,000 over the same period. McDonogh frames these gains less as cost savings and more as capacity creation. "We haven't reduced the footprint, but it's allowed us to do more with the footprint that we have," he said. To track progress, BNY measures AI impact across core workflows -- including innovating, prospecting, onboarding, transacting, and streamlining -- while continuously building out its internal "Eliza" platform. The system serves as a firm-wide context layer, improving over time as it ingests more data and use cases. Employee adoption is also structured. Staff progress through three levels of AI proficiency, culminating in a "pioneer" designation that requires formal training and testing. Access to more advanced models is gated by expertise, reinforcing both quality and accountability. Within finance specifically, AI is already reshaping core processes. McDonogh points to regulatory reporting, balance sheet analytics and predictive modeling as key use cases. The technology is also playing a growing role in earnings preparation, helping synthesize analyst expectations and anticipate investor questions. For McDonogh, the takeaway is straightforward: AI productivity is not about how much you use, but how effectively it changes what an organization can do.
[4]
The AI 'tokenmaxxing' corporate fad is fading as workplaces look to cut costs
Tech executives cast high AI usage as a badge of honor Just a few months ago, Silicon Valley executives were promoting high token consumption as a signal of high-performing employees. The stereotypical tokenmaxxer was staying up late -- perhaps ignoring their significant other -- while orchestrating an army of 24-hour AI agents performing work on their behalf. OpenAI CEO Sam Altman said in May he was "excited to see what will happen with tokenmaxxing startups, both for how they work internally and the products they can build." Nvidia CEO Jensen Huang said, "If your $500K engineer isn't burning $250K in tokens, something is wrong." Facebook parent Meta had an internal competition rewarding token usage. The trend boosted revenue for leading AI large language model developers like Anthropic and OpenAI, but it fizzled as it became apparent it wasn't necessarily the best strategy for everyone else. Microsoft CEO Satya Nadella has admitted that tokenmaxxing can be addictive but warned in a recent blog post that customers of those models are paying twice for AI, first in spending on tokens and second by feeding all their proprietary data to them. While promoting Microsoft's own approach, Nadella's comments were unusual in the way he raised doubts about the data protection assurances of leading AI providers. Palantir CEO Alex Karp went further, telling CNBC earlier this month that something had gone "completely wrong." He said he was channeling the voice of American businesses privately "livid" about paying so much for tokens that create no value. "The basic view among enterprises in this country is, 'I'm going to chillax and waste my time with tokens. I'm going to get no value and they're going to get my IP,'" Karp said. Workplaces look more for better 'routing' of their AI work Bain & Company management consultant Jue Wang said many of the big businesses her firm advises have been taking a closer look at returns on their AI investments. "The token cost for them has been doubling, almost every other month," she said. "Let's say $200 per developer per month. Multiply that by 20,000 developers, which is often what we're dealing with at these companies, and that quickly gets you to a number that is not a line item that any general manager has planned for." Sometimes that just means not using the AI equivalent of a sledgehammer to crack a nut. "Not everything needs a Claude Opus 4.6," she said of one of Anthropic's more capable models suited to software engineering or deep research. "And yet you see so many companies, so many users, default to using Opus for everything, including generating emails." That's led to a search for tools that do AI "model routing" -- in which easier queries get automatically sent to cheaper and more efficient AI systems and more complex tasks go to more powerful models. Open-source AI models built in China offer less costly alternatives Software developer Hassan El Mghari said companies' sticker shock over the "ridiculous amount of money" spent on subscriptions to AI products from leading U.S. companies has led many away from rewarding high usage. "It's better to kind of just empower employees on how to use this stuff and let them use AI when and however much they need to," said El Mghari, who leads developer experience at the startup Together AI, which supplies developers with a variety of "open-source" AI models. At the same time, those who favor racking up as many tokens as possible are having a field day with new open-source models from Chinese startups like Moonshot's Kimi or Zhipu's GLM, which nearly match the capabilities of top U.S. models at a fraction of the price. "There is some validity to the theory that this could push tokenmaxxing a little bit further," said Raffi Krikorian, the chief technology officer at Mozilla. "But if we look at the industry overall, I think it's realizing that tokenmaxxing is a dumb thing." It's similar, Krikorian said, to how software companies once considered how many lines of code a programmer wrote to be a good metric of productivity. That later fell out of favor. "I think tokenmaxxing is moving through the exact same pattern," he said. "I think this is going to be an interesting blip that we're all going to look back to laugh at in a year."
[5]
A flex in corporate America, AI 'tokenmaxxing' fades as workplaces look to cut tech spending
A corporate fad of "tokenmaxxing" on artificial intelligence technology is hitting its limits as workplaces throwing AI at everything are seeing the costs rise without a similar spike in productivity. What started as tech industry-fueled springtime hype over squeezing as much AI-generated work as possible out of products like OpenAI's ChatGPT and Anthropic's Claude has shifted to a summertime backlash. "It's very easy to create something you don't need with AI," said Vincent Gusdorf, head of AI analytics at Moody's Ratings and author of a new report that recommends a more disciplined approach. "Tokenmaxxing" refers to maximizing usage of tokens -- the building blocks of generative AI that correspond to small pieces of text that an AI system reads or writes. Each token is about three quarters of a word. And there's typically a limit to how many you can use, with pricier versions of AI products offering higher caps. "As bills started to pile in, people realized that those new tools are quite expensive and you need to use them wisely," Gusdorf said. Tech executives cast high AI usage as a badge of honor Just a few months ago, Silicon Valley executives were promoting high token consumption as a signal of high-performing employees. The stereotypical tokenmaxxer was staying up late -- perhaps ignoring their significant other -- while orchestrating an army of 24-hour AI agents performing work on their behalf. OpenAI CEO Sam Altman said in May he was "excited to see what will happen with tokenmaxxing startups, both for how they work internally and the products they can build." Nvidia CEO Jensen Huang said "if your $500K engineer isn't burning $250K in tokens, something is wrong." Facebook parent Meta had an internal competition rewarding token usage. The trend boosted revenue for leading AI large language model developers like Anthropic and OpenAI, but it fizzled as it became apparent it wasn't necessarily the best strategy for everyone else. Microsoft CEO Satya Nadella has admitted that tokenmaxxing can be addictive but warned in a recent blog post that customers of those models are paying twice for AI, first in spending on tokens and second by feeding all their proprietary data to them. While promoting Microsoft's own approach, Nadella's comments were unusual in the way he raised doubts about the data protection assurances of leading AI providers. Palantir CEO Alex Karp went further, telling CNBC earlier this month that something had gone "completely wrong." He said he was channeling the voice of American businesses privately "livid" about paying so much for tokens that create no value. "The basic view among enterprises in this country is, 'I'm going to chillax and waste my time with tokens. I'm going to get no value and they're going to get my IP," Karp said. Workplaces look more for better 'routing' of their AI work Bain & Company management consultant Jue Wang said many of the big businesses her firm advises have been taking a closer look at returns on their AI investments. "The token cost for them has been doubling, almost every other month," she said. "Let's say $200 per developer per month. Multiply that by 20,000 developers, which is often what we're dealing with at these companies, and that quickly gets you to a number that is not a line item that any general manager has planned for." Sometimes that just means not using the AI equivalent of a sledgehammer to crack a nut. "Not everything needs a Claude Opus 4.6," she said of one of Anthropic's more capable models suited to software engineering or deep research. "And yet you see so many companies, so many users, default to using Opus for everything, including generating emails." That's led to a search for tools that do AI "model routing" -- in which easier queries get automatically sent to cheaper and more efficient AI systems and more complex tasks go to more powerful models. Open-source AI models built in China offer less costly alternatives Software developer Hassan El Mghari said companies' sticker shock over the "ridiculous amount of money" spent on subscriptions to AI products from leading U.S. companies has led many away from rewarding high usage. "It's better to kind of just empower employees on how to use this stuff and let them use AI when and however much they need to," said El Mghari, who leads developer experience at the startup Together AI, which supplies developers with a variety of "open-source" AI models. At the same time, those who favor racking up as many tokens as possible are having a field day with new open-source models from Chinese startups like Moonshot's Kimi or Zhipu's GLM, which nearly match the capabilities of top U.S. models at a fraction of the price. "There is some validity to the theory that this could push tokenmaxxing a little bit further," said Raffi Krikorian, the chief technology officer at Mozilla. "But if we look at the industry overall, I think it's realizing that tokenmaxxing is a dumb thing." It's similar, Krikorian said, to how software companies once considered how many lines of code a programmer wrote to be a good metric of productivity. That later fell out of favor. "I think tokenmaxxing is moving through the exact same pattern," he said. "I think this is going to be an interesting blip that we're all going to look back to laugh at in a year."
[6]
Workplaces Look for Cheaper AI as 'Tokenmaxxing' Fades as a Corporate Fad
A corporate fad of "tokenmaxxing" on artificial intelligence technology is hitting its limits as workplaces throwing AI at everything are seeing the costs rise without a similar spike in productivity. What started as tech industry-fueled springtime hype over squeezing as much AI-generated work as possible out of products like OpenAI's ChatGPT and Anthropic's Claude has shifted to a summertime backlash. "It's very easy to create something you don't need with AI," said Vincent Gusdorf, head of AI analytics at Moody's Ratings and author of a new report that recommends a more disciplined approach. "Tokenmaxxing" refers to maximizing usage of tokens -- the building blocks of generative AI that correspond to small pieces of text that an AI system reads or writes. Each token is about three quarters of a word. And there's typically a limit to how many you can use, with pricier versions of AI products offering higher caps. "As bills started to pile in, people realized that those new tools are quite expensive and you need to use them wisely," Gusdorf said. Tech executives cast high AI usage as a badge of honor Just a few months ago, Silicon Valley executives were promoting high token consumption as a signal of high-performing employees. The stereotypical tokenmaxxer was staying up late -- perhaps ignoring their significant other -- while orchestrating an army of 24-hour AI agents performing work on their behalf. OpenAI CEO Sam Altman said in May he was "excited to see what will happen with tokenmaxxing startups, both for how they work internally and the products they can build." Nvidia CEO Jensen Huang said "if your $500K engineer isn't burning $250K in tokens, something is wrong." Facebook parent Meta had an internal competition rewarding token usage. The trend boosted revenue for leading AI large language model developers like Anthropic and OpenAI, but it fizzled as it became apparent it wasn't necessarily the best strategy for everyone else. Microsoft CEO Satya Nadella has admitted that tokenmaxxing can be addictive but warned in a recent blog post that customers of those models are paying twice for AI, first in spending on tokens and second by feeding all their proprietary data to them. While promoting Microsoft's own approach, Nadella's comments were unusual in the way he raised doubts about the data protection assurances of leading AI providers. Palantir CEO Alex Karp went further, telling CNBC earlier this month that something had gone "completely wrong." He said he was channeling the voice of American businesses privately "livid" about paying so much for tokens that create no value. "The basic view among enterprises in this country is, 'I'm going to chillax and waste my time with tokens. I'm going to get no value and they're going to get my IP," Karp said. Workplaces look more for better 'routing' of their AI work Bain & Company management consultant Jue Wang said many of the big businesses her firm advises have been taking a closer look at returns on their AI investments. "The token cost for them has been doubling, almost every other month," she said. "Let's say $200 per developer per month. Multiply that by 20,000 developers, which is often what we're dealing with at these companies, and that quickly gets you to a number that is not a line item that any general manager has planned for." Sometimes that just means not using the AI equivalent of a sledgehammer to crack a nut. "Not everything needs a Claude Opus 4.6," she said of one of Anthropic's more capable models suited to software engineering or deep research. "And yet you see so many companies, so many users, default to using Opus for everything, including generating emails." That's led to a search for tools that do AI "model routing" -- in which easier queries get automatically sent to cheaper and more efficient AI systems and more complex tasks go to more powerful models. Open-source AI models built in China offer less costly alternatives Software developer Hassan El Mghari said companies' sticker shock over the "ridiculous amount of money" spent on subscriptions to AI products from leading U.S. companies has led many away from rewarding high usage. "It's better to kind of just empower employees on how to use this stuff and let them use AI when and however much they need to," said El Mghari, who leads developer experience at the startup Together AI, which supplies developers with a variety of "open-source" AI models. At the same time, those who favor racking up as many tokens as possible are having a field day with new open-source models from Chinese startups like Moonshot's Kimi or Zhipu's GLM, which nearly match the capabilities of top U.S. models at a fraction of the price. "There is some validity to the theory that this could push tokenmaxxing a little bit further," said Raffi Krikorian, the chief technology officer at Mozilla. "But if we look at the industry overall, I think it's realizing that tokenmaxxing is a dumb thing." It's similar, Krikorian said, to how software companies once considered how many lines of code a programmer wrote to be a good metric of productivity. That later fell out of favor. "I think tokenmaxxing is moving through the exact same pattern," he said. "I think this is going to be an interesting blip that we're all going to look back to laugh at in a year."
Share
Copy Link
The corporate trend of tokenmaxxing—maximizing AI token usage—is collapsing under mounting costs. What began as a Silicon Valley badge of honor has become an AI cost sink, with companies now prioritizing model routing and efficient AI strategies over raw consumption. Industry leaders warn that unchecked AI token usage threatens ROI while data privacy concerns loom large.
A corporate phenomenon that swept through Silicon Valley just months ago is rapidly losing momentum. Tokenmaxxing, the practice of maximizing AI tokens consumed through platforms like OpenAI's ChatGPT and Anthropic's Claude, emerged as a status symbol among tech workers and executives who viewed high token consumption as proof of productivity
4
. OpenAI CEO Sam Altman declared in May that he was "excited to see what will happen with tokenmaxxing startups," while Nvidia's Jensen Huang proclaimed that "if your $500K engineer isn't burning $250K in tokens, something is wrong"4
. Meta even launched internal competitions rewarding AI token usage4
. But as summer arrived, the honeymoon ended. Companies discovered that their AI cost was spiraling without corresponding productivity gains, transforming what seemed like innovation into an unsustainable AI cost sink.AI tokens represent the fundamental unit of information processed by AI models—small chunks of text and data that AI systems read and generate with each query
2
. Each token corresponds to roughly three-quarters of a word5
. As AI as a utility becomes reality—with Sam Altman envisioning "a future where intelligence is a utility, like electricity or water, and people buy it from us on a meter"—AI tokens are emerging as the kilowatt-hour of AI, the standard measure for consumption and billing2
. While AI companies still offer flat-rate subscriptions to consumers, they increasingly charge businesses through token-based pricing2
. This shift has exposed the true economic impact of corporate AI adoption, with agentic AI systems consuming orders of magnitude more tokens than traditional generative AI applications1
.
Source: ZDNet
Steve Lucas, CEO of integration specialist Boomi, experienced the financial shock firsthand. "Last year, I personally spent at Boomi 10 times the amount on Claude that I did the previous year—10 times; that's not sustainable. I can't do that every year," he told ZDNET
1
. Bain & Company management consultant Jue Wang reported that token costs for large enterprises have been doubling almost every other month. At $200 per developer per month multiplied by 20,000 developers, companies face bills reaching millions—expenses no general manager had budgeted for4
. Vincent Gusdorf, head of AI analytics at Moody's Ratings, observed that "as bills started to pile in, people realized that those new tools are quite expensive and you need to use them wisely"5
. The reality check has been stark: managing AI expenses is now a C-suite priority, with the practice of "tokenomics"—measuring, pricing, and managing token consumption—becoming essential business activity1
.Beyond cost, data privacy concerns have fueled skepticism about unchecked token consumption. Microsoft CEO Satya Nadella warned that customers are "paying twice for AI"—first in token spending and second by feeding proprietary data to AI providers—while raising unusual doubts about data protection assurances from leading AI companies
4
. Palantir CEO Alex Karp told CNBC that American businesses are "livid" about paying for tokens that create no value while risking their intellectual property. "The basic view among enterprises in this country is, 'I'm going to chillax and waste my time with tokens. I'm going to get no value and they're going to get my IP,'" Karp said4
.Companies are pivoting toward efficient AI strategies that prioritize outcomes over volume. The key technique gaining traction is model routing, which automatically directs simple queries to cheaper, efficient AI systems while reserving powerful models for complex tasks
4
. "Not everything needs a Claude Opus 4.6," explained Jue Wang, referring to Anthropic's advanced model suited for software engineering. "And yet you see so many companies, so many users, default to using Opus for everything, including generating emails"5
. Open-source AI models from Chinese startups like Moonshot's Kimi and Zhipu's GLM offer capabilities nearly matching top U.S. models at a fraction of the price, providing alternatives for cost-conscious organizations4
.
Source: Fast Company
While some companies chased tokenmaxxing glory, BNY Mellon took a fundamentally different approach. CFO Dermot McDonogh told Fortune that token costs remain "modest within modest" relative to the bank's engineering budget, and the tokenmaxxing conversation never took hold internally
3
. "I couldn't tell you how many prompts we did last week," McDonogh said. "I'm focused more on outcomes" . Those outcomes are measurable: in the first quarter of 2026, more than 40% of BNY's code was authored by AI, rising to roughly 50% more recently. Revenue per employee climbed from $338,000 in 2022 to $401,000 in 2025, while pre-tax income per employee jumped from $99,000 to $143,000 . BNY's approach demonstrates that AI productivity metrics focused on capacity creation and business outcomes deliver more value than raw token consumption.
Source: Fortune
Related Stories
Snowflake CEO Sridhar Ramaswamy acknowledged concerns about AI token usage but emphasized the need to give staff room to explore agentic AI systems. "Are we worried about how much we are spending on AI inference across our different internal teams? Absolutely. But do I see that spend as a reason not to use AI? Absolutely not," he said at the company's Summit 2026 event
1
. Matt Luizzi, VP of analytics at Whoop, echoed this balanced approach: "We have guardrails in place, and monitoring and observability to let people know what they're spending"1
. The consensus emerging from business leaders centers on enabling innovation within managed boundaries rather than constraining exploration entirely.Researchers are discovering that AI tokens offer unprecedented visibility into AI's economic impact. Economists Nicola Borri, Aleh Tsyvinski, and Yukun Liu analyzed data from 380 trillion AI tokens to understand how AI consumption reshapes financial markets, identifying an "AI premium" in stock prices of companies benefiting from AI adoption
2
. Their research suggests that "the story of AI is no longer just a Silicon Valley story," with financial markets expecting Main Street businesses across industries to feel the impact2
. Platforms like OpenRouter, which aggregate access to hundreds of AI models, are creating rich datasets that allow researchers to track AI adoption with precision impossible in previous technological revolutions2
.Mozilla CTO Raffi Krikorian predicts tokenmaxxing will become "an interesting blip that we're all going to look back to laugh at in a year," comparing it to outdated metrics like lines of code written
4
. Hassan El Mghari from Together AI suggests the path forward lies in empowering "employees on how to use this stuff and let them use AI when and however much they need to"4
. Steve Lucas frames the fundamental question: "What matters now is, 'Can I operate AI at a return?'"1
. Organizations must develop strategies that balance exploration with fiscal discipline, implement model routing to optimize costs, and focus on productivity gains rather than consumption metrics. The shift from tokenmaxxing to strategic token management marks a maturation point in corporate AI adoption, where the focus moves from hype to sustainable value creation.Summarized by
Navi
21 Jul 2026•Business and Economy

20 Mar 2026•Business and Economy

17 Jun 2026•Business and Economy

1
Technology

2
Policy and Regulation

3
Technology
