6 Sources
[1]
The AI race is shifting from bigger models to cheaper, smarter systems
Benchmark's Peter Fenton says open-weight models could soon handle most AI usage, putting pressure on the economics of the biggest model providers. For the past two years, the artificial intelligence race has been easy to score: bigger models, better benchmarks and whichever company could claim the lead, at least until the next launch. That scorecard is starting to look incomplete. As companies move from testing AI to using it in real products and workflows, it's not longer about tapping the best model, but accessing the one that's the best fit for a specific job, at the right cost, with the necessary data and in a chosen environment. That shift is opening the door for a new kind of AI competition, one focused less on model size and more on routing, cost, control and compute. "The model alone is no longer the product," Perplexity CEO Aravind Srinivas told CNBC. "It is the harness, the orchestration system that puts the model inside a very capable harness and pairs the model with a lot of tools." That means AI products are becoming systems that can decide which model to use, when to use it and what outside tools or company data sources are necessary. A customer service task might not need the most expensive model. A complex coding problem might. A routine internal workflow could run on a cheaper open model. A harder step could be escalated to a more powerful one. "The answer is always use whatever is the best for the task," Srinivas said. The emergence of alternative models comes as corporate America tightens its belt on AI spending, and presents another challenge for OpenAI and Anthropic, which have flourished over the past few years by selling the most cutting-edge technology. Perplexity this week previewed a new system for its computer-use product built around GLM 5.2, an open model from China's Z.ai. The system is designed to let a cheaper model handle more of the work while calling in a stronger model only when needed. That approach reflects a broader change in the market. Open-weight models, which can be downloaded, tuned and run by companies themselves, are becoming more capable. They are also cheaper to run than premium proprietary models from the biggest AI labs. Benchmark general partner Peter Fenton said the shift could be dramatic. "A maybe contrarian view that is becoming consensus is our belief that 90-plus percent of the tokens created will come out of open-weight models over the next 18 to 24 months, possibly even by the end of the year," Fenton told CNBC. Tokens are the units of data AI models process and generate. "The inference margins generated by the frontier model companies, I think, are going to come under pressure when you can run those without the markup that they're providing, when you have good enough models from open weights," Fenton said. Fenton said the move to open models is not only about saving money. In some cases, smaller models that are tuned for a specific task can be faster and perform better than larger general-purpose models. That is one reason Benchmark invested in Ollama, a company that makes it easier for developers and enterprises to download, run and manage open models. "One thing is where the model's from and where it was created and trained," Ollama CEO Jeff Morgan said. "But the more important thing to these businesses we speak to is where it runs and how it runs." Morgan said Ollama has been adopted by more than 85% of the Fortune 500, including companies in regulated industries such as aviation, insurance and health care. He said many companies start with smaller models running close to their own data, then expand to larger open models as they get more comfortable. The rise of open models also creates a strategic challenge for the U.S. Many of the most competitive open-weight models are coming from Chinese labs, including Z.ai and DeepSeek. That has made open-source AI a business issue, a policy issue and a national competitiveness issue. Srinivas said the U.S. should support open models because they make AI more affordable and accessible. "If you want the benefits of AI to be widely distributed to small businesses in America and American allied countries, then you really need AI to be a lot more affordable," Srinivas said. "And open source is the only way to do that." The shift could also affect the massive data center buildout underway across the tech industry. The current AI boom assumes demand will keep flowing to large cloud data centers filled with high-end chips. Srinivas says some AI work may eventually run locally instead, on devices owned by consumers or businesses. That wouldn't eliminate the need for data centers, but it could create a more hybrid AI system, with routine tasks run locally and the most difficult work getting sent to a more powerful model in the cloud. For investors, the question is whether the biggest AI labs can maintain their pricing power as open models get better and companies become more selective about what they use. Choose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.
[2]
The AI race is no longer about the biggest model
The assumption that the biggest AI model wins is breaking down, with enterprises now choosing models by task, cost, and control rather than leaderboard rank. Driving it are model bills running to millions a month, the rise of model routing, and specialised task-specific agents, which Gartner expects in 40% of enterprise applications by end-2026, up from under 5%. If capability is commoditising, the margin moves to whoever runs inference cheapest. For years the industry ran on one assumption, that the biggest model wins. That belief is now breaking down, CNBC reports. Companies are choosing models by task, cost, and control instead of benchmark position. The frontier still matters, but it is no longer the only thing being bought. The reason is unromantic. At enterprise scale, model bills run into millions of dollars a month. The rise of good enough The operating principle is now the cheapest model that clears the quality bar. Buyers have worked out that most tasks do not need a frontier system. Model routing has emerged to automate that judgment, sending each request to whichever model suits it. A summarisation job and a multi-step reasoning job no longer go to the same place. Specialised, industry-specific models are filling the rest of the gap. Gartner expects 40% of enterprise applications to embed task-specific AI agents by the end of 2026, up from under 5% a year earlier. Why the bills forced this The economics stopped adding up. Per-token prices have collapsed, yet enterprise AI bills have tripled anyway, because agentic tools consume vastly more tokens per task. Buyers noticed. Palo Alto Networks chief executive Nikesh Arora has said token prices need to fall by as much as 90% for adoption to scale. Some firms gave up waiting and started rationing. A wave of "tokenminimizing" has companies capping employee AI spending outright. Where the value moves next If capability is commoditising, the margin migrates to whoever runs it cheapest. Inference optimisation has quietly become one of AI infrastructure's most valuable layers. Open and cheap models sharpen the point. Chinese models are closing in on the US frontier labs at a fraction of the price, which caps what anyone can charge for merely competent output. This is uncomfortable for the scaling thesis. Hundreds of billions in capex were justified by the premise that bigger models would stay decisively better, and buyers are now voting otherwise. None of it means frontier models are finished. It means the industry is discovering that most work is boring, and boring work does not need the most expensive tool in the shop.
[3]
Companies are shifting toward cheaper open‑source AI models to rein in costs, Amazon CTO says | Fortune
Companies worried about mounting AI bills are increasingly shifting to cheaper, open-source models, according to Amazon's chief technology officer, Werner Vogels. "We see a shift happening between the cheaper open source models and the bigger expensive models," Vogels said in an interview on the sidelines of the UN's AI for Good summit. Stories of runaway AI bills have been making some executives skittish about building systems on the most advanced models from companies such as OpenAI, Anthropic, and Google DeepMind, that bill by the token. (A token is the basic unit of data an AI model processes, equivalent to about a word and a half of English language text.) Uber said it burned through its entire 2026 AI budget in four months, while company reportedly burned through half a billion dollars in a single month after failing to cap AI usage for employees have caused concern across industries. Fears of runaway spending are forcing companies to rethink how -- and where -- they deploy the most powerful frontier models. While large models from companies like OpenAI, Anthropic, and Google often deliver top-tier performance, they also come with significantly higher operating costs, particularly when deployed at scale. Open source models (also sometimes referred to as "open weight" models) can usually be downloaded for free but then users have to pay for their own cloud computing infrastructure on which to run the models. Still, this often works out to be cheaper than using the most advanced proprietary models. "Cost is a very important part of your architecture, you need to take that into account," Vogels said. "Do you really need to have the biggest, highest‑end model to solve this? The answer is no, you don't." The shift also reflects a broader maturation in how companies are thinking about AI adoption. After an initial wave of experimentation fueled by hype and rapid advances in large language models, many organizations are now entering a more pragmatic phase focused on return on investment. That means scrutinizing not just what AI can do, but what it costs to deploy and maintain over time. While customers may be shifting toward open‑source models as a cheaper option, Vogels also said companies were also putting a premium on transparency and trust in how models are trained and deployed "Transparency becomes extremely important," he said. "People want to know what is the data that goes into it." That demand is particularly acute in sectors like healthcare, government, and humanitarian work, where understanding how an AI system was trained -- and how it makes decisions -- can be as important as its performance. "If these people serve vulnerable communities. If they don't trust the system, they won't use it," Vogels said. Open-source models, which allow developers to inspect and modify code and more easily fine tune the model on their own data, are often seen as better aligned with those needs. But even most open weight model providers do not fully reveal all of the data on which the model was initially trained. At the Summit, Vogels also launched a new Amazon open-source AI tool designed to help researchers find relevant scientific datasets quickly. The system connects the AWS Registry of Open Data -- home to more than 1,100 datasets from organizations including NASA, NOAA, and the NIH -- to AI assistants, allowing users to search using natural language rather than navigating complex data catalogs. By enabling queries such as requests for satellite imagery or genomics datasets with specific licensing, the tool aims to replace processes that would have previously taken hours, lowering technical barriers for scientists, particularly those at under-resourced institutions, and accelerating research across fields such as climate science and public health.
[4]
AI models could soon get cheaper as OpenAI, Meta, and xAI enter a new price war
Businesses are shifting their focus from raw AI power to cost efficiency, and OpenAI, Meta, and SpaceXAI are capitalizing on the trend. All three companies have recently released new AI models that emphasize lower operational costs, a move that could put pressure on Anthropic in the AI enterprise space. According to Bloomberg, OpenAI's latest model, GPT-5.6, is designed to complete more work while using fewer tokens, and this shift represents a fundamental shift in the AI race as the costs for more sophisticated or higher-power models begin to be felt by AI companies. Meta and SpaceXAI have also launched updated models, with SpaceXAI debuting Grok 4.5, claiming improved token efficiency, which directly targets the exponential cost for enterprise clients. With business customers increasingly scrutinizing AI spending, efficiency is now a major selling point, and, according to Bloomberg, is the new direction AI companies are heading. This new emphasis on cost comes as companies seek to justify AI budgets amid shifting economic expectations. Token efficiency, which measures the amount of data processed and billed, is now a primary metric for enterprise users and is the main figure that companies using AI models are looking to increase as much as possible, even at the cost of using a less powerful model. A model that uses fewer tokens for the same task can mean significant savings at scale, which is what Meta, Anthropic, SpaceXAI, and any other tech company that has its hat in the mainstream AI ring is focusing on, whether that be through hardware advancement or software refinements. As the race to deliver cost-effective AI heats up, Anthropic may find itself challenged unless it can match or exceed the latest efficiency benchmarks set by its competitors, but in the meantime it's attempted to maintain users by extending access to its renowned Fable 5 model. The next few months will reveal whether these new models can reshape the enterprise AI landscape, or if they are simply another rung in the ostensibly endless ladder of AI advancement.
[5]
OpenAI, Meta, SpaceXAI compete for more cost-efficient AI models
Three prominent artificial intelligence developers released new models over the past week. They all promise to be more advanced, but their biggest immediate selling point may not be what they can do, but how little they charge to do it. OpenAI said its most advanced offering, GPT-5.6, is designed to complete more work while using significantly fewer tokens, a unit of data processed by AI models. This will make the software far more cost efficient for customers. Grok 4.5, from Elon Musk's SpaceXAI, is billed as having twice the token efficiency as comparable models from other firms. And Meta Platforms is making the pricing for its Muse Spark 1.1 very "attractive," Chief Executive Officer Mark Zuckerberg told Bloomberg.
[6]
Enterprise AI model race shifts: cost, speed and fit now win
Smaller models handle routine tasks as inference costs dominate budgets Between 2022 and 2024, enterprise AI buying got a lot less fixated on benchmark wins and a lot more focused on task fit and cost. You can see that in enterprise use cases like support-ticket classification and contract extraction, where smaller or more specialized models are often fast enough and accurate enough at a much lower price. Inference bills can still reach the millions, and multi-step agent workflows keep pushing total spend higher even as per-token prices keep falling. That's leading to a pretty straightforward strategy: use the cheapest model that still clears the quality bar. Mistral AI's Mistral 7B can be more than 200 times cheaper per request than OpenAI's GPT-4, so companies send summarization and tagging to lightweight models and keep the expensive systems for legal work, coding, and more complex reasoning. Gartner says 40% of apps will use task-specific agents by the end of 2026. A year earlier, that number was under 5%. The research points in the same direction. Benchmark-size needs fell 142x from 2022 to 2024, and with a fixed compute budget, smaller models trained on more data can outperform larger ones. If you buy or deploy AI, this is worth downloading. Palo Alto Networks CEO Nikesh Arora says token prices may need to fall 90%. Some companies are capping usage. Meta's Llama and Alibaba's Qwen have passed 1 billion downloads. Cheaper Chinese rivals are closing the gap. And the value is shifting toward inference, orchestration, governance, and infrastructure. You can see it across enterprise tools.
Share
Copy Link
The AI industry is experiencing a fundamental shift as companies move away from simply chasing the most powerful models toward prioritizing cost efficiency and task-specific performance. Major players like OpenAI, Meta, and SpaceXAI are now competing on token efficiency and pricing, while open-source models gain traction with enterprises looking to control AI spending that has spiraled into millions monthly.
The AI race has entered a new phase where cost efficiency matters more than model size. For two years, artificial intelligence development followed a straightforward pattern: bigger models produced better benchmarks, and companies competed to launch the most powerful system
1
. That scorecard now looks incomplete as enterprises move from testing to deploying AI models in real products and workflows. Companies are choosing AI models based on task-specific performance, computational costs, and control rather than leaderboard rankings2
. The shift reflects an uncomfortable reality: at enterprise scale, model bills run into millions of dollars monthly, forcing organizations to rethink their approach2
.
Source: Japan Times
The economics of AI deployment have forced a fundamental recalculation across corporate America. Stories of runaway spending have made executives skittish about building systems on the most advanced proprietary models from companies like OpenAI, Anthropic, and Google DeepMind
3
. Uber burned through its entire 2026 AI budget in four months, while one company reportedly consumed half a billion dollars in a single month after failing to cap AI usage for employees3
. These cautionary tales have accelerated the shift to open-source AI as organizations seek alternatives to expensive frontier models. Amazon CTO Werner Vogels confirmed this trend, stating that companies are increasingly moving toward cheaper open-source models to rein in costs3
. "Cost is a very important part of your architecture, you need to take that into account," Vogels said. "Do you really need to have the biggest, highest-end model to solve this? The answer is no, you don't"3
.
Source: Fortune
The operating principle for enterprise applications has become selecting the cheapest model that clears the quality bar. Model routing systems have emerged to automate this judgment, directing each request to whichever model suits it best
2
. A customer service task might not need the most expensive model, while a complex coding problem might require more power1
. Perplexity CEO Aravind Srinivas explained that "the model alone is no longer the product. It is the harness, the orchestration system that puts the model inside a very capable harness and pairs the model with a lot of tools"1
. Token efficiency has become a primary metric for enterprise users, measuring the amount of data processed and billed. Companies are now willing to use less powerful models if they deliver significant savings at scale through improved token efficiency4
. Gartner expects 40% of enterprise applications to embed task-specific AI agents by end-2026, up from under 5% a year earlier2
.
Source: Softonic
Related Stories
OpenAI, Meta, and SpaceXAI have all released new models emphasizing lower operational costs, entering what amounts to an AI price war
4
. OpenAI's GPT-5.6 is designed to complete more work while using significantly fewer tokens, making the software far more cost-efficient for customers5
. Grok 4.5 from Elon Musk's SpaceXAI claims twice the token efficiency as comparable models from other firms5
. Meta is making the pricing for its Muse Spark 1.1 very attractive, according to CEO Mark Zuckerberg5
. This competition could put pressure on Anthropic in the AI enterprise space unless it can match or exceed the latest efficiency benchmarks4
.Benchmark general partner Peter Fenton predicts that 90-plus percent of tokens created will come from open-weight models over the next 18 to 24 months, possibly even by year's end
1
. Open-source models can be downloaded, tuned, and run by companies themselves, typically at lower costs than premium proprietary models from the biggest AI labs1
. "The inference margins generated by the frontier model companies are going to come under pressure when you can run those without the markup that they're providing," Fenton said1
. Inference optimization has quietly become one of AI infrastructure's most valuable layers as capability commoditizes2
. Benchmark invested in Ollama, a company making it easier for developers and enterprises to download, run, and manage open models. Ollama CEO Jeff Morgan reports adoption by more than 85% of the Fortune 500, including companies in regulated industries such as aviation, insurance, and healthcare1
. Chinese models from labs including Z.ai and DeepSeek are closing in on US frontier capabilities at a fraction of the price, creating both business and national competitiveness concerns1
.Summarized by
Navi
[2]
26 Jul 2026•Policy and Regulation

17 Jun 2026•Business and Economy

16 Aug 2025•Technology
