10 Sources
[1]
AI is becoming a bargain hunter's market, with a few luxury models on top
The price of AI tokens is fluctuating widely, with some becoming cheaper and others more expensive, leaving users of AI services struggling to assess if the price is right. Aman Panjwani, an AI engineer based in India, says that GPT-4-class model output cost about $20 per million tokens in late
[2]
Employers pushed staff to use AI more. That has backfired
Some Amazon developers must be suffering whiplash about how they use AI. The technology group sets targets for most of its developers to use AI tools. This year it started monitoring consumption and posting usage figures on internal leader boards. But after the FT wrote in May that some workers
[3]
Token maxxing is your AI program's quiet failure mode
Most organizations have settled on those numbers because they're easy to track and show up well in reports AI investment is accelerating across enterprises. Budgets are increasing, boardrooms are asking hard questions, and the answer they're getting back is token costs, prompt counts, and copilot
[4]
DeepSeek cut prices 75%. The 100x problem remains
DeepSeek's recent decision to drastically cut pricing on its V4-Pro model by 75% should have been unequivocally good news for enterprise AI vendors and developers. Instead, many are discovering that cheaper models don't automatically translate into healthier margins. The reason is simple: While
[5]
Cognition CEO says tech companies got 'carried away' with token leaderboards and should measure employees on output instead | Fortune
The trend of tokenmaxxing has gone too far. That's at least according to Cognition CEO Scott Wu, who argues that as companies scramble to rein in AI spending, they should focus on employee productivity instead of AI use. In an episode of the "Founders" podcast with David Senra, Wu said that as
[6]
How to embrace the spirit of 'Tokenmaxxing' without breaking the bank
"Tokenmaxxing" - the idea that AI coding success comes down to using as many tokens as possible - is an appealing metric. Tokens are the fundamental unit that AI coding tools use to read, write, and reason. So on the surface, more tokens should mean more output, more productivity, and more
[7]
Enterprise hits and misses - harness engineering takes over enterprise AI, and tokenomics is in full swing
Last week, I addressed reader pushback on my real-time organizational truth missive. I also unfurled something of a call to action: * Studies have shown that smaller models can reduce LLM operating costs by as much as 90 percent. * Even Google just admitted, in a major new playbook, that proper
[8]
The token economy: The state of AI mid-2026
Gigawatt factories in the Texas scrub. Thirty trillion tokens a day. A $30 billion storage company, a search engine built for machines, and a continent trying to buy its independence one graphics processing unit at a time. Three years after ChatGPT, the artificial intelligence business has stopped
[9]
The (not so) Hidden Cost of Using GenAI: By Steve Morgan
It's no secret that during the first wave of GenAI adoption, most of the fees were hidden behind subsidised pricing, big allowances and a relentless focus on driving usage. The message to financial services and all other industries was clear: use more GenAI. No doubt excitement and experimentation
[10]
Token Unmaxxing: AI's Price War Is Starting to Reshape the Trade
The issue is not that demand for AI is evaporating. It is that the price of intelligence is beginning to fall faster than many investors expected. Takeaways * Token demand is not the same thing as token pricing power. AI usage can keep compounding while realised spend per token falls under free
Share
Copy Link
Companies are discovering that cheaper AI models don't automatically mean lower bills. While token prices have plummeted 55x since 2022, AI costs have surged 10x since January as employees consume tokens faster than prices decline. Tech giants like Amazon and Meta scrapped internal leaderboards after tokenmaxxing led to wasteful spending, prompting a industry-wide rethink of how to measure AI value.
The AI market is experiencing a paradox that's reshaping how companies approach AI adoption: while token prices have collapsed, actual spending is skyrocketing. GPT-4-class model output that cost $20 per million tokens in late 2022 now costs just $0.40, representing a 55x decline in less than four years, according to AI engineer Aman Panjwani
1
. When DeepSeek released its R1 reasoning model in January 2025 at $0.55 per million input tokens against OpenAI's o1-preview at $15, the entire market repriced overnight with a 97 percent discount1
. Yet despite DeepSeek's recent 75% price cut on its V4-Pro model, companies are discovering that cheaper models don't automatically translate into healthier margins .
Source: SiliconANGLE
The reality is that AI costs have increased approximately 10x between January and now, especially in engineering operations, according to Ameya Kanitkar, CTO of Larridin, an AI measurement platform
1
. Companies are now spending between 10 and 20 percent of their labor costs on token usage, translating to $2,000 to $4,000 per month for a software engineer paid $200,000 annually1
. Uber burned through its entire AI budget for 2026 in just four months and capped token spending for employees at $1,500 per month5
.The surge in AI costs stems partly from a phenomenon called token maxxing, where employees compete to use as many tokens as possible rather than focus on actual business outcomes
2
. Amazon set targets for developers to use AI tools and started monitoring consumption through internal leaderboards. But after reports emerged that workers were automating non-essential tasks to improve their rankings, Amazon took the leaderboards offline2
. "Please don't use AI just for the sake of using AI," Amazon senior vice-president Dave Treadwell told staff2
.
Source: FT
Cognition CEO Scott Wu argues that companies got "carried away" with token leaderboards and should measure employees on output instead. "People are like, 'We rank our engineers by how many tokens they're spending.' Well, let's try and rank people by how much output they're actually producing," Wu said
5
. This shift reflects a broader industry reckoning with how to properly incentivize AI adoption without creating wasteful spending patterns.The transition from subscription to token-based billing has fundamentally changed the economics of AI adoption. As model providers switched pricing models, the cost of profligate usage started hitting corporate balance sheets hard
2
. The problem is amplified by agentic workflows, which consume tokens voraciously. While a chatbot turns one user question into one model call, an agent turns it into a chain of planning, retrieval, tool use, verification, summarization, and follow-up decisions4
.
Source: VentureBeat
This creates what industry observers call the "100x problem": the same user-visible request can cost exponentially more to serve as an agentic workflow than as a simple chatbot response
4
. In multi-step agent workflows, the input-to-billed ratio routinely lands at 1:700 or higher, compared to about 1:5 for single-turn chatbots4
. Token amplification breaks traditional SaaS pricing models, as power users running 50 agent invocations daily on a $40/seat plan can cost more in inference costs than the plan charges4
.The token market is splitting into two distinct segments: commodity inference heading toward zero while frontier inference costs rise. OpenAI doubled the price of GPT-5.5 to $5 input and $30 output per million tokens, while Google's Gemini Flash 3.5 arrived three to six times more expensive than the model it replaced
1
. Anthropic's Claude Sonnet 5, despite having a lower per-token price than Claude Opus 4.8, uses more tokens to produce the same results1
.Open-weight models are emerging as a cost lever for companies seeking alternatives. Kimi 2.6/2.7 and GLM 5.2 are almost at parity with Opus 4.7 or 4.8 and are 10x cheaper in theory, or about 5x cheaper in practice
1
. Nearly 75 percent of companies now use multi-model strategies, switching between providers based on task requirements1
.Related Stories
Higher token usage doesn't necessarily correlate with higher productivity. Larridin data shows that between 15 and 30 percent of AI users account for more than 50 percent of total AI spend, and often that spending does not correlate with gains in output
1
. When Larridin plotted token spend against developer productivity, it found an inflection point at about 35 to 40 percent of client spending where burning more tokens failed to boost productivity1
.The challenge extends beyond token consumption to how companies measure success. Most organizations track token costs, prompt counts, and deployment numbers, but these are activity logs, not business metrics
3
. They show how much AI is being used, not whether the business is performing better3
. Boston Consulting Group's 2026 Global AI at Work report found that 42% of workers reported regular AI use saving them eight hours per week, but 66% said they received little to no guidance on how to invest that saved time5
.As AI costs hit balance sheets, companies are rethinking their approach to strategic AI implementation. The era of experimentation has given way to demands for clear AI return on investment. "This year for the first time I started to see a maturing of the AI conversation. Now we need to see a return on it, rather than spraying and praying," says Ravin Jesuthasan, senior partner at Mercer
2
.The scale of investment required is evident in OpenAI's proposed program to give every Y Combinator startup $2 million in API credits, an amount that would have funded an entire seed round in any prior tech cycle
4
. For established enterprises, the absolute numbers are larger still. The dominant AI business model assumed by most company plans doesn't survive contact with agentic workloads, forcing a fundamental rethink of how AI creates value at the organizational-level change required goes beyond individual productivity gains to transforming how work moves across teams and systems3
.Summarized by
Navi
[4]
21 Jul 2026•Business and Economy

17 Jun 2026•Business and Economy

14 Aug 2026•Business and Economy
1
Technology

2
Policy and Regulation

3
Technology
