4 Sources
[1]
Uber's tech chief says the AI 'tokenmaxxing' era is ending
The maximalist phase of AI spending is drawing to a close, Uber's CTO argues, as companies start asking what all those tokens actually bought them. The mood around corporate AI spending is turning, and Uber is saying so out loud. The company's chief technology officer, Praveen Neppalli Naga, says the era of tokenmaxxing is coming to an end. The term needs a translation. Tokenmaxxing is the practice of spending ever more on AI, measured in the tokens that models consume, on the assumption that more usage automatically means more value. Uber knows the habit from the inside. The company burned through its entire Claude Code budget for 2026 by April, an eye-catching sign of how quickly AI bills can balloon once teams are set loose. That experience has bred caution and Uber's leadership has started questioning whether all that spending is actually producing the productivity gains and new products it was supposed to unlock. The company's president put it starkly earlier this year. Andrew Macdonald said the link between higher AI spending and shipping successful features simply is not there yet, even as usage statistics soared. The headline numbers on AI adoption, he said, make your head explode, and yet nothing had meaningfully gained traction in the products customers actually use. So Uber is pumping the brakes. Rather than chase maximal usage, it is adopting a more disciplined approach, continuing to work with the big model providers but demanding clearer evidence of value. Uber is far from alone in this rethink. Atlassian has begun putting its engineers on AI budgets as the cost of tokenmaxxing bites, a sign that the free-for-all is giving way to spreadsheets. The wider evidence backs the skeptics. Study after study has found that most enterprise AI spending never leaves the pilot stage, producing demos and proofs of concept rather than shipping products. The uncomfortable question is one vendors would rather avoid. There is a reason it is the question AI providers hope engineering leaders never ask, namely whether the output justifies the invoice. The economics are catching up with the hype. GitHub recently froze new Copilot sign-ups because agentic usage blew past what its pricing could bear, a concrete example of the strain tokenmaxxing puts on providers too. None of this means companies are abandoning AI. The shift is from unlimited experimentation toward efficiency, smaller and cheaper models, and a harder look at which use cases genuinely pay for themselves. The timing matters for the whole industry. Model makers have justified enormous valuations on the assumption that enterprise AI spending only rises, so a heavy customer signalling restraint is a data point the market will notice. It also reframes what progress looks like. If the next phase rewards efficiency over raw consumption, the advantage may shift from whoever has the biggest model to whoever delivers the most useful work per dollar. That would suit buyers and unsettle sellers. Cheaper, smaller models and tighter budgets are good news for companies footing the bill, and a harder sell for providers whose revenue grows with every token burned. Coming from Uber, the message carries weight. This is a company that spends heavily on technology and works closely with the leading labs, so its caution is not a laggard's excuse but a heavy user's verdict. Skeptics will note the caveat, of course. Pumping the brakes is not the same as stopping, and Uber is still spending heavily with the big labs even as it preaches discipline about where the money goes. The tokenmaxxing era was always going to meet a budget. Neppalli Naga is simply naming the moment when the industry stops asking how much AI it can buy and starts asking what it is getting in return.
[2]
After blowing its entire 2026 AI budget in months, Uber CTO says 'We're coming to the end of the so-called 'tokenmaxxing' era' | Fortune
Uber believes it's found a solution to its AI spending problem after it blew through its budget for the technology in just the first few months of the year. In an interview with The Information earlier this year, Uber Chief Technology Officer Praveen Neppalli Naga admitted he went "back to the drawing board" on allotted spending after the rideshare giant encouraged employees to use its tools, particularly Anthropic's Claude Code, as much as possible, even devising "leader boards" to rank software engineers on their usage. The blitz was part of a trend of "tokenmaxxing," or companies incentivizing workplace AI use, only for many to back off from the practice as they found it wasn't offering the returns on investment to justify the rapid spending. While Uber was no exception, Naga said the company has now figured out a better way to deploy AI without breaking the bank. "We're seeing some very interesting trends on AI costs," he wrote in an X post on Wednesday. "I think it's another signal that we're coming to the end of the so-called 'tokenmaxxing' era." Uber quadrupled the number of employees who use frontier AI tools, Naga explained, which brought down the cost per token. It was able to do this by improving prompt caching, as well as adjusting its default model setting, evaluating new models for efficiency, and allowing engineers to see their AI usage and costs per hour. "You might expect costs to rise as adoption accelerates," Naga continued. "We've seen the opposite. Not because we've restricted access, but because we've treated efficiency as an engineering problem rather than a budget problem." AI's rising ROI stakes The stakes are increasing for companies to deliver on their massive AI investments. Last month, Jim Reid, Deutsche Bank Research Institute's global head of macro and thematic research, warned AI productivity gains were still years away. Profit margins for the Magnificent Seven swelled from 15% to 25% between the first quarters of 2023 to 2026, while the rest of the S&P 500 index saw only 10% margin growth over the same period, indicating little widespread returns on investment in AI outside of the immediate tech sector. As of May, Uber was still trying to unlock the innovation AI promised. "That link is not there yet," Uber President and Chief Operating Officer Andrew Macdonald said in an interview on the Rapid Response podcast at the time. "Maybe implicitly there's more that is getting shipped, but it's very hard to draw a line between one of those stats and 'Okay now we're actually producing like 25% more useful consumer features.'" The threat of Jevons paradox Even as Uber unlocks strategies to lower the cost per token to make its AI use more sustainable, it risks falling into a trap economists have warned about: Jevons paradox, in which spending on a resource, in this case tokens, actually increases even as its cost decreases. Named for 19th century economist William Stanley Jevons, the phenomenon originally referred to his observation of coal consumption skyrocketing in 1865, despite the Watt steam engine making coal use more efficient. The same dynamic is playing out today with AI: According to the Silicon Data Token Expenditure Index, the price of a single token dropped more than 90% since 2023, but large language model spending has doubled since late last year. "As tokens get cheaper, companies don't spend less but instead run more AI agents, automate more workflows and generate more code, pushing aggregate expenditure higher even as the unit cost of intelligence collapses," Apollo Chief Economist Torsten Slok wrote in a recent blog post. A Bain and Co. brief published in June punctuated Slok's claim. It found that token costs halved from December 2024 to 2025, but tokens consumed grew by 450% over the same period as companies upgraded AI tools. Naga, for his part, noted a shift in company philosophy to put quality over quantity when it comes to token spending, but did not say if Uber was using more or less computing than earlier this year. "This is the future of applied AI at enterprise scale," he concluded. "The next phase, whatever we call it, will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible."
[3]
Uber CTO Says Thousands of Engineers Now Use Advanced AI Tools as Company Cuts Cost Per Token: 'The Next
Uber Technologies Inc. (NYSE:UBER) says the next phase of AI will focus less on consuming more tokens and more on improving efficiency, reducing costs and maximizing value. Uber Lowers AI Cost Per Token On Wednesday, Uber CTO Praveen Neppalli said the company has seen a major increase in employees using advanced AI tools while simultaneously lowering its cost per token. "As our CFO @_balaji_km mentioned at earnings today, we're seeing some very interesting trends on AI costs," he wrote He added, "I think it's another signal that we're coming to the end of the so-called 'tokenmaxxing' era." Neppalli said Uber has more than quadrupled the number of people using frontier AI tools since the beginning of the year, with thousands of engineers using AI daily. Despite the increased adoption, he said the company's cost per token has declined. He attributed the improvement to engineering-focused optimizations rather than restricting access to AI tools. Uber has improved its prompt caching and reuse systems to reduce input token spending, adjusted default model settings and context sizes, and provided engineers with real-time visibility into AI usage and costs. The company is also testing open-weight AI models and selecting different models based on specific use cases. "The next phase, whatever we call it, will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible," Neppalli said. AI Race Accelerates as Engineers Adapt Earlier, venture capitalist Chamath Palihapitiya said AI may have entered a recursive self-improvement cycle, allowing systems to help build more advanced models, accelerate breakthroughs and reduce costs. He said the trend could represent a step toward the AI singularity. Anthropic CEO Dario Amodei predicted AI could soon handle most software engineering tasks, noting that some engineers are already using AI to generate and edit code. AI pioneer Andrew Ng warned that engineers who fail to adopt AI tools could fall behind, saying developers who combine technical experience with AI skills will be better positioned as the industry evolves. Disclaimer: This content was partially produced with the help of AI tools and was reviewed and published by Benzinga editors. Photo courtesy: Shutterstock Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
[4]
Uber says AI spending now has to prove it pays off: "tokenmaxxing" is ending
Uber Technologies says the industry's "tokenmaxxing" phase is winding down. The mood is getting a lot less loose: if AI costs keep climbing, companies want proof that the spending turns into real business results. CTO Praveen Neppalli Naga says teams are moving away from treating more prompts, more tokens, and heavier use of coding assistants or chatbots as automatic signs of progress. That's the backdrop even after, according to reports, Uber had already burned through its 2026 budget for Anthropic's Claude Code by April. At Uber, use of advanced generative AI tools has quadrupled since the start of the year. Even so, the company says it's bringing down cost per interaction with prompt caching, more efficient default models, and better visibility into consumption. The timing matters. Global generative AI spending is heading toward about 2.5 trillion in 2026, and according to Uber, only around 25% of 2025 projects actually hit payback targets, even though 84% of executives say returns are improving. If you're paying for these tools on a regular basis, this shift deserves attention, especially after Uber Technologies COO Andrew Macdonald said in May 2026 that spending trade-offs were getting harder to justify without matching productivity gains. Uber points to its internal "Agentic Pods" in finance, legal, and marketing: one planning process dropped from 15 hours to 30 minutes, and report creation fell from two days to 10 minutes. That's where this is showing up, in model switching, semantic caching, token reduction, and smaller, cheaper models.
Share
Copy Link
Uber burned through its entire 2026 AI budget by April, prompting CTO Praveen Neppalli Naga to declare the end of the tokenmaxxing era. The rideshare giant is now demanding measurable returns on AI investments, shifting from unlimited experimentation to disciplined spending that prioritizes efficiency over raw consumption.
Uber is declaring the end of the tokenmaxxing era, marking a significant turning point in how companies approach AI spending
1
. Chief Technology Officer Praveen Neppalli Naga announced that the rideshare giant has moved away from the maximalist phase of AI adoption, where companies spent freely on tokens without demanding clear evidence of value. This shift comes after Uber burned through its entire Claude Code budget for 2026 by April, an eye-opening example of how quickly AI spending can spiral when teams operate without constraints2
.
Source: The Next Web
The term tokenmaxxing refers to the practice of spending ever more on AI, measured in the tokens that models consume, on the assumption that more usage automatically translates to more value
1
. Uber actively encouraged this behavior, creating leaderboards to rank software engineers on their usage of Anthropic's Claude Code and other frontier AI tools2
. The company quadrupled the number of employees using advanced AI tools since the beginning of the year, with thousands of engineers now using AI daily3
.The uncomfortable reality driving this shift is that productivity gains have not materialized as expected. Uber President Andrew Macdonald stated bluntly in May 2026 that the link between higher AI spending and shipping successful features simply is not there yet, even as usage statistics soared
1
. "The headline numbers on AI adoption make your head explode, and yet nothing had meaningfully gained traction in the products customers actually use," Macdonald explained1
.This assessment aligns with broader industry evidence. Study after study has found that most enterprise AI spending never leaves the pilot stage, producing demos and proofs of concept rather than shipping products that deliver measurable business outcomes
1
. According to Uber, only around 25% of 2025 projects actually hit payback targets4
. The economics are catching up with the hype, forcing companies to ask whether the output justifies the invoice.Rather than restrict access to AI tools, Uber is treating efficiency as an engineering problem rather than a budget problem
2
. The company has successfully reduced its cost per token while expanding adoption through several technical optimizations. These include improved prompt caching and reuse systems to reduce input token spending, adjusted default model settings and context sizes, and real-time visibility for engineers into their AI usage and costs per hour2
3
.Uber is also evaluating new models for efficiency, testing open-weight AI models, and selecting different models based on specific use cases through strategic model selection
2
3
. "You might expect costs to rise as adoption accelerates. We've seen the opposite," Naga wrote2
. Where Uber has achieved operational efficiency, the results are striking. Internal "Agentic Pods" in finance, legal, and marketing have reduced one planning process from 15 hours to 30 minutes, and report creation from two days to 10 minutes4
.
Source: Benzinga
Uber is far from alone in this rethink. Atlassian has begun putting its engineers on AI budgets as the cost of tokenmaxxing bites, signaling that the free-for-all is giving way to spreadsheets and accountability
1
. GitHub recently froze new Copilot sign-ups because agentic usage blew past what its pricing could bear, a concrete example of the strain tokenmaxxing puts on providers too1
.The shift toward a more disciplined approach to AI spending matters for the entire industry. Model makers have justified enormous valuations on the assumption that enterprise AI spending only rises, so a heavy customer signaling restraint is a data point the market will notice
1
. Global generative AI spending is heading toward approximately $2.5 trillion in 20264
, but the pressure for return on investment is intensifying.Related Stories
Even as companies pursue cost-efficiency, they face a counterintuitive risk: Jevons paradox, where spending on a resource actually increases even as its cost decreases
2
. Named for 19th century economist William Stanley Jevons, who observed coal consumption skyrocketing in 1865 despite efficiency improvements, the phenomenon is playing out in AI today. According to the Silicon Data Token Expenditure Index, the price of a single token dropped more than 90% since 2023, but large language model spending has doubled since late last year2
.
Source: Fortune
A Bain and Co. brief published in June found that token costs halved from December 2024 to 2025, but tokens consumed grew by 450% over the same period as companies upgraded AI tools
2
. "As tokens get cheaper, companies don't spend less but instead run more AI agents, automate more workflows and generate more code, pushing aggregate expenditure higher even as the unit cost of intelligence collapses," Apollo Chief Economist Torsten Slok explained2
.The end of the tokenmaxxing era reframes what progress looks like in AI adoption. If the next phase rewards efficient AI usage over raw consumption, the advantage may shift from whoever has the biggest model to whoever delivers the most useful work per dollar
1
. "The next phase, whatever we call it, will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible," Naga concluded2
3
.Coming from Uber, this message carries weight. This is a company that spends heavily on technology and works closely with leading labs, so its caution is not a laggard's excuse but a heavy user's verdict
1
. The shift toward cost-conscious AI adoption and demanding proof that AI spending now has to prove it pays off represents a maturation of the market. Companies are continuing to work with big model providers but demanding clearer evidence of value before writing blank checks. The tokenmaxxing era was always going to meet a budget constraint. Uber is simply naming the moment when the industry stops asking how much AI it can buy and starts asking what it is getting in return1
.Summarized by
Navi
[1]
[3]
17 Jun 2026•Business and Economy

21 Jul 2026•Business and Economy

26 May 2026•Business and Economy

1
Technology

2
Policy and Regulation

3
Technology
