3 Sources
[1]
Uber's tech chief says the AI 'tokenmaxxing' era is ending
The maximalist phase of AI spending is drawing to a close, Uber's CTO argues, as companies start asking what all those tokens actually bought them. The mood around corporate AI spending is turning, and Uber is saying so out loud. The company's chief technology officer, Praveen Neppalli Naga, says the era of tokenmaxxing is coming to an end. The term needs a translation. Tokenmaxxing is the practice of spending ever more on AI, measured in the tokens that models consume, on the assumption that more usage automatically means more value. Uber knows the habit from the inside. The company burned through its entire Claude Code budget for 2026 by April, an eye-catching sign of how quickly AI bills can balloon once teams are set loose. That experience has bred caution and Uber's leadership has started questioning whether all that spending is actually producing the productivity gains and new products it was supposed to unlock. The company's president put it starkly earlier this year. Andrew Macdonald said the link between higher AI spending and shipping successful features simply is not there yet, even as usage statistics soared. The headline numbers on AI adoption, he said, make your head explode, and yet nothing had meaningfully gained traction in the products customers actually use. So Uber is pumping the brakes. Rather than chase maximal usage, it is adopting a more disciplined approach, continuing to work with the big model providers but demanding clearer evidence of value. Uber is far from alone in this rethink. Atlassian has begun putting its engineers on AI budgets as the cost of tokenmaxxing bites, a sign that the free-for-all is giving way to spreadsheets. The wider evidence backs the skeptics. Study after study has found that most enterprise AI spending never leaves the pilot stage, producing demos and proofs of concept rather than shipping products. The uncomfortable question is one vendors would rather avoid. There is a reason it is the question AI providers hope engineering leaders never ask, namely whether the output justifies the invoice. The economics are catching up with the hype. GitHub recently froze new Copilot sign-ups because agentic usage blew past what its pricing could bear, a concrete example of the strain tokenmaxxing puts on providers too. None of this means companies are abandoning AI. The shift is from unlimited experimentation toward efficiency, smaller and cheaper models, and a harder look at which use cases genuinely pay for themselves. The timing matters for the whole industry. Model makers have justified enormous valuations on the assumption that enterprise AI spending only rises, so a heavy customer signalling restraint is a data point the market will notice. It also reframes what progress looks like. If the next phase rewards efficiency over raw consumption, the advantage may shift from whoever has the biggest model to whoever delivers the most useful work per dollar. That would suit buyers and unsettle sellers. Cheaper, smaller models and tighter budgets are good news for companies footing the bill, and a harder sell for providers whose revenue grows with every token burned. Coming from Uber, the message carries weight. This is a company that spends heavily on technology and works closely with the leading labs, so its caution is not a laggard's excuse but a heavy user's verdict. Skeptics will note the caveat, of course. Pumping the brakes is not the same as stopping, and Uber is still spending heavily with the big labs even as it preaches discipline about where the money goes. The tokenmaxxing era was always going to meet a budget. Neppalli Naga is simply naming the moment when the industry stops asking how much AI it can buy and starts asking what it is getting in return.
[2]
Uber CTO Says Thousands of Engineers Now Use Advanced AI Tools as Company Cuts Cost Per Token: 'The Next
Uber Technologies Inc. (NYSE:UBER) says the next phase of AI will focus less on consuming more tokens and more on improving efficiency, reducing costs and maximizing value. Uber Lowers AI Cost Per Token On Wednesday, Uber CTO Praveen Neppalli said the company has seen a major increase in employees using advanced AI tools while simultaneously lowering its cost per token. "As our CFO @_balaji_km mentioned at earnings today, we're seeing some very interesting trends on AI costs," he wrote He added, "I think it's another signal that we're coming to the end of the so-called 'tokenmaxxing' era." Neppalli said Uber has more than quadrupled the number of people using frontier AI tools since the beginning of the year, with thousands of engineers using AI daily. Despite the increased adoption, he said the company's cost per token has declined. He attributed the improvement to engineering-focused optimizations rather than restricting access to AI tools. Uber has improved its prompt caching and reuse systems to reduce input token spending, adjusted default model settings and context sizes, and provided engineers with real-time visibility into AI usage and costs. The company is also testing open-weight AI models and selecting different models based on specific use cases. "The next phase, whatever we call it, will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible," Neppalli said. AI Race Accelerates as Engineers Adapt Earlier, venture capitalist Chamath Palihapitiya said AI may have entered a recursive self-improvement cycle, allowing systems to help build more advanced models, accelerate breakthroughs and reduce costs. He said the trend could represent a step toward the AI singularity. Anthropic CEO Dario Amodei predicted AI could soon handle most software engineering tasks, noting that some engineers are already using AI to generate and edit code. AI pioneer Andrew Ng warned that engineers who fail to adopt AI tools could fall behind, saying developers who combine technical experience with AI skills will be better positioned as the industry evolves. Disclaimer: This content was partially produced with the help of AI tools and was reviewed and published by Benzinga editors. Photo courtesy: Shutterstock Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
[3]
Uber says AI spending now has to prove it pays off: "tokenmaxxing" is ending
Uber Technologies says the industry's "tokenmaxxing" phase is winding down. The mood is getting a lot less loose: if AI costs keep climbing, companies want proof that the spending turns into real business results. CTO Praveen Neppalli Naga says teams are moving away from treating more prompts, more tokens, and heavier use of coding assistants or chatbots as automatic signs of progress. That's the backdrop even after, according to reports, Uber had already burned through its 2026 budget for Anthropic's Claude Code by April. At Uber, use of advanced generative AI tools has quadrupled since the start of the year. Even so, the company says it's bringing down cost per interaction with prompt caching, more efficient default models, and better visibility into consumption. The timing matters. Global generative AI spending is heading toward about 2.5 trillion in 2026, and according to Uber, only around 25% of 2025 projects actually hit payback targets, even though 84% of executives say returns are improving. If you're paying for these tools on a regular basis, this shift deserves attention, especially after Uber Technologies COO Andrew Macdonald said in May 2026 that spending trade-offs were getting harder to justify without matching productivity gains. Uber points to its internal "Agentic Pods" in finance, legal, and marketing: one planning process dropped from 15 hours to 30 minutes, and report creation fell from two days to 10 minutes. That's where this is showing up, in model switching, semantic caching, token reduction, and smaller, cheaper models.
Share
Copy Link
Uber burned through its entire 2026 Claude Code budget by April, prompting a major shift in strategy. CTO Praveen Neppalli Naga now says the era of unlimited AI spending is over as companies demand clear evidence that tokenmaxxing delivers real business value, not just impressive usage statistics.
Uber is pumping the brakes on unlimited AI spending
1
, with Chief Technology Officer Praveen Neppalli Naga declaring the end of the 'tokenmaxxing' era. The ride-hailing giant burned through its entire 2026 Claude Code budget by April1
, a striking example of how quickly enterprise AI spending can spiral when teams operate without constraints. This experience has fundamentally changed how Uber approaches advanced AI tools, shifting from maximizing token consumption to demanding measurable returns on every dollar spent.
Source: The Next Web
Tokenmaxxing refers to the practice of spending ever more on AI, measured in the tokens that models consume, on the assumption that more usage automatically translates to more value
1
. Uber's leadership now questions whether this approach actually produces the productivity gains and new products it was supposed to unlock. Company president Andrew Macdonald stated earlier this year that the link between higher AI spending and shipping successful features simply isn't there yet, even as usage statistics soared1
. The headline numbers on AI adoption make your head explode, he said, yet nothing had meaningfully gained traction in the products customers actually use.Despite the cautionary tone, Uber has more than quadrupled the number of people using frontier AI tools since the beginning of the year, with thousands of engineers now using AI daily
2
. The company's cost per token has simultaneously declined through engineering-focused optimizations rather than restricting access2
. Uber improved its prompt caching and reuse systems to reduce input token spending, adjusted default model settings and context sizes, and provided engineers with real-time visibility into AI usage and costs2
.The company is also testing open-weight AI models and selecting different models based on specific use cases
2
. This disciplined approach to AI spending represents a fundamental shift in strategy. Neppalli emphasized that the next phase will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible2
. Uber continues working with big model providers but now demands clearer evidence of value before committing resources.Uber is far from alone in this rethink. Atlassian has begun putting its engineers on AI budgets as the cost of tokenmaxxing bites, signaling that the free-for-all is giving way to spreadsheets
1
. Study after study has found that most enterprise AI spending never leaves the pilot stage, producing demos and proofs of concept rather than shipping products1
. Global generative AI spending is heading toward approximately $2.5 trillion in 2026, yet according to Uber, only around 25% of 2025 projects actually hit payback targets, even though 84% of executives say returns are improving3
.
Source: Benzinga
The economics are catching up with the hype. GitHub recently froze new Copilot sign-ups because agentic usage blew past what its pricing could bear, a concrete example of the strain tokenmaxxing puts on providers too
1
. This uncomfortable reality forces the question AI providers hope engineering leaders never ask: whether the output justifies the invoice. The shift is from unlimited experimentation toward efficiency, smaller and cheaper models, and a harder look at which use cases genuinely pay for themselves1
.Related Stories
Uber points to its internal Agentic Pods in finance, legal, and marketing as proof that cost-conscious AI adoption can deliver results. One planning process dropped from 15 hours to 30 minutes, and report creation fell from two days to 10 minutes
3
. These improvements showcase where efficient AI usage is showing up: in model selection, semantic caching, token reduction, and smaller, cheaper models3
.The timing matters for the whole industry. Model makers have justified enormous valuations on the assumption that enterprise AI spending only rises, so a heavy customer signaling restraint is a data point the market will notice
1
. If the next phase rewards operational efficiency over raw consumption, the advantage may shift from whoever has the biggest model to whoever delivers the most useful work per dollar. That would suit buyers and unsettle sellers, as cheaper, smaller models and tighter cost management are good news for companies footing the bill, and a harder sell for providers whose revenue grows with every token burned1
. Coming from Uber, the message carries weight as a heavy user's verdict rather than a laggard's excuse.Summarized by
Navi
[1]
[2]
17 Jun 2026•Business and Economy

21 Jul 2026•Business and Economy

26 May 2026•Business and Economy

1
Technology

2
Technology

3
Science and Research
