Uber CTO declares end of AI 'tokenmaxxing' era as companies demand proof spending pays off

3 Sources

Share

Uber burned through its entire 2026 Claude Code budget by April, prompting a major shift in strategy. CTO Praveen Neppalli Naga now says the era of unlimited AI spending is over as companies demand clear evidence that tokenmaxxing delivers real business value, not just impressive usage statistics.

Uber Signals Major Shift in AI Spending Strategy

Uber is pumping the brakes on unlimited AI spending

1

, with Chief Technology Officer Praveen Neppalli Naga declaring the end of the 'tokenmaxxing' era. The ride-hailing giant burned through its entire 2026 Claude Code budget by April

1

, a striking example of how quickly enterprise AI spending can spiral when teams operate without constraints. This experience has fundamentally changed how Uber approaches advanced AI tools, shifting from maximizing token consumption to demanding measurable returns on every dollar spent.

Source: The Next Web

Source: The Next Web

Tokenmaxxing refers to the practice of spending ever more on AI, measured in the tokens that models consume, on the assumption that more usage automatically translates to more value

1

. Uber's leadership now questions whether this approach actually produces the productivity gains and new products it was supposed to unlock. Company president Andrew Macdonald stated earlier this year that the link between higher AI spending and shipping successful features simply isn't there yet, even as usage statistics soared

1

. The headline numbers on AI adoption make your head explode, he said, yet nothing had meaningfully gained traction in the products customers actually use.

Cost Per Token Drops Despite Quadrupled Usage

Despite the cautionary tone, Uber has more than quadrupled the number of people using frontier AI tools since the beginning of the year, with thousands of engineers now using AI daily

2

. The company's cost per token has simultaneously declined through engineering-focused optimizations rather than restricting access

2

. Uber improved its prompt caching and reuse systems to reduce input token spending, adjusted default model settings and context sizes, and provided engineers with real-time visibility into AI usage and costs

2

.

The company is also testing open-weight AI models and selecting different models based on specific use cases

2

. This disciplined approach to AI spending represents a fundamental shift in strategy. Neppalli emphasized that the next phase will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible

2

. Uber continues working with big model providers but now demands clearer evidence of value before committing resources.

Industry-Wide Rethink of AI Economics

Uber is far from alone in this rethink. Atlassian has begun putting its engineers on AI budgets as the cost of tokenmaxxing bites, signaling that the free-for-all is giving way to spreadsheets

1

. Study after study has found that most enterprise AI spending never leaves the pilot stage, producing demos and proofs of concept rather than shipping products

1

. Global generative AI spending is heading toward approximately $2.5 trillion in 2026, yet according to Uber, only around 25% of 2025 projects actually hit payback targets, even though 84% of executives say returns are improving

3

.

Source: Benzinga

Source: Benzinga

The economics are catching up with the hype. GitHub recently froze new Copilot sign-ups because agentic usage blew past what its pricing could bear, a concrete example of the strain tokenmaxxing puts on providers too

1

. This uncomfortable reality forces the question AI providers hope engineering leaders never ask: whether the output justifies the invoice. The shift is from unlimited experimentation toward efficiency, smaller and cheaper models, and a harder look at which use cases genuinely pay for themselves

1

.

Operational Efficiency Through Agentic Pods

Uber points to its internal Agentic Pods in finance, legal, and marketing as proof that cost-conscious AI adoption can deliver results. One planning process dropped from 15 hours to 30 minutes, and report creation fell from two days to 10 minutes

3

. These improvements showcase where efficient AI usage is showing up: in model selection, semantic caching, token reduction, and smaller, cheaper models

3

.

The timing matters for the whole industry. Model makers have justified enormous valuations on the assumption that enterprise AI spending only rises, so a heavy customer signaling restraint is a data point the market will notice

1

. If the next phase rewards operational efficiency over raw consumption, the advantage may shift from whoever has the biggest model to whoever delivers the most useful work per dollar. That would suit buyers and unsettle sellers, as cheaper, smaller models and tighter cost management are good news for companies footing the bill, and a harder sell for providers whose revenue grows with every token burned

1

. Coming from Uber, the message carries weight as a heavy user's verdict rather than a laggard's excuse.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved