5 Sources
[1]
OpenAI and Anthropic in price war as Chinese AI rivals gain ground
Leading US AI labs such as OpenAI and Anthropic are releasing cheaper models as they fight to retain cost-conscious customers who are switching to cut-price alternatives from Chinese rivals. The price war comes as rising AI bills push companies to curb usage and seek cheaper models, helping Chinese developers including Moonshot and DeepSeek make inroads with users from Silicon Valley to Europe. OpenAI recently said that it was slashing prices for GPT-5.6 Luna, its "fastest and most affordable model", by 80 percent. Anthropic has launched Claude Opus 5, touting the system's "frontier intelligence... at half the price" of Fable 5, the company's most capable model. The moves have helped decrease prices that customers are paying for models from leading US labs by almost a quarter since mid-July, according to Silicon Data's token price index. Tokens are the units of data processed by language models and are used to calculate many customers' bills. The cuts mark a shift for US AI groups that make proprietary "closed" models that have, until now, competed heavily on performance. Increasingly capable "open" Chinese models -- which can be freely downloaded and tweaked by developers -- have contributed to pressure on prices. The moves also come as OpenAI and Anthropic plot initial public offerings at trillion-dollar valuations while investors seek evidence that the industry's vast spending on AI can generate returns. Corporate AI users face cost pressures as Anthropic and OpenAI shift some enterprise customers away from flat subscriptions and toward usage-based billing, under which companies pay according to the computational resources they consume. Some businesses have responded to a rise in bills by imposing caps on AI usage or testing cheaper alternatives. Companies such as DoorDash and Airbnb have said they have started to use Chinese-made models in an effort to rein in bills. That shift has coincided with a flurry of releases from Chinese labs that have narrowed the performance gap with leading US models, raising concerns in the US tech industry that American developers could lose customers even as they spend heavily to maintain their technological edge. AI labs offer a range of models with different capabilities and prices, with costs varying further according to the version of a model and the "effort" settings used. Customers are typically charged for input tokens, used to measure data fed into a model, and output tokens, which measure what it generates in response. The latest price cuts from US labs apply to mid-tier products and make them more competitive with Chinese offerings. OpenAI, for example, cut the price of GPT-5.6 Luna from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output tokens. Anthropic launched Opus 5 at $5 per million input tokens and $25 per million output tokens -- half the price of its Fable 5 model. This week, the company called off a planned rise in prices for its Sonnet 5 model, which had been due to take effect from September. Headline token prices do not provide a straightforward comparison between AI models, however. More capable models can sometimes complete a task using fewer tokens or with fewer attempts, meaning a model that appears more expensive based on the headline price of tokens can ultimately cost less. Additionally, most can operate at different "effort" settings, which vary the computing power used to answer a question and can affect both performance and the ultimate cost of completing a task. Artificial Analysis, which benchmarks models across areas including math, science, coding, and reasoning, found Anthropic's Opus 5 at "medium" effort delivered similar performance and cost per task to Moonshot's Kimi K3 at "max" effort. OpenAI's GPT-5.6 Luna at "max" effort performed similarly to DeepSeek's V4 Flash at "max," but cost just under twice as much per task. Anthropic and OpenAI declined to comment. A person close to Anthropic said Opus 5's pricing below its flagship Fable 5 was how the startup's "family of models is built, so there's no connection to competitors." Mantas Lukauskas, AI tech lead at Hostinger, a website hosting provider that has used large language models since 2020, noted that prices for the very best models were "flat to rising." He added that the recent pricing changes are the "first real test" of whether groups such as Anthropic and OpenAI can protect the cost of their most advanced offerings: "The US labs have cut the middle and are defending the top."
[2]
OpenAI and Anthropic in price war as Chinese AI rivals gain ground
Leading US AI labs such as OpenAI and Anthropic are releasing cheaper models as they fight to retain cost-conscious customers who are switching to cut-price alternatives from Chinese rivals. The price war comes as rising AI bills push companies to curb usage and seek cheaper models, helping Chinese developers including Moonshot and DeepSeek make inroads with users from Silicon Valley to Europe. OpenAI recently said that it was slashing prices for GPT-5.6 Luna, its "fastest and most affordable model", by 80 per cent. Anthropic has launched Claude Opus 5, touting the system's "frontier intelligence . . . at half the price" of Fable 5, the company's most capable model. The moves have helped decrease prices that customers are paying for models from leading US labs by almost a quarter since mid-July, according to Silicon Data's token price index. Tokens are the units of data processed by language models and are used to calculate many customers' bills. The cuts mark a shift for US AI groups that make proprietary "closed" models which have, until now, competed heavily on performance. Increasingly capable "open" Chinese models -- which can be freely downloaded and tweaked by developers -- have contributed to pressure on prices. The moves also come as OpenAI and Anthropic plot initial public offerings at trillion-dollar valuations while investors seek evidence that the industry's vast spending on AI can generate returns. Corporate AI users face cost pressures as Anthropic and OpenAI shift some enterprise customers away from flat subscriptions and towards usage-based billing, under which companies pay according to the computational resources they consume. Some businesses have responded to a rise in bills by imposing caps on AI usage or testing cheaper alternatives. Companies such as DoorDash and Airbnb have said they have started to use Chinese-made models in an effort to rein in bills. That shift has coincided with a flurry of releases from Chinese labs that have narrowed the performance gap with leading US models, raising concerns in the US tech industry that American developers could lose customers even as they spend heavily to maintain their technological edge. AI labs offer a range of models with different capabilities and prices, with costs varying further according to the version of a model and the "effort" settings used. Customers are typically charged for input tokens, used to measure data fed into a model, and output tokens, which measure what it generates in response. The latest price cuts from US labs apply to mid-tier products and make them more competitive with Chinese offerings. OpenAI, for example, cut the price of GPT-5.6 Luna from $1 to $0.20 per mn input tokens and from $6 to $1.20 per mn output tokens. Anthropic launched Opus 5 at $5 per mn input tokens and $25 per mn output tokens -- half the price of its Fable 5 model. This week, the company called off a planned rise in prices for its Sonnet 5 model, which had been due to take effect from September. Headline token prices do not provide a straightforward comparison between AI models, however. More capable models can sometimes complete a task using fewer tokens or with fewer attempts, meaning a model that appears more expensive based on the headline price of tokens can ultimately cost less. Additionally, most can operate at different "effort" settings, which vary the computing power used to answer a question and can affect both performance and the ultimate cost of completing a task. Artificial Analysis, which benchmarks models across areas including maths, science, coding and reasoning, found Anthropic's Opus 5 at "medium" effort delivered similar performance and cost per task to Moonshot's Kimi K3 at "max" effort. OpenAI's GPT-5.6 Luna at "max" effort performed similarly to DeepSeek's V4 Flash at "max", but cost just under twice as much per task. Anthropic and OpenAI declined to comment. A person close to Anthropic said Opus 5's pricing below its flagship Fable 5 was how the start-up's "family of models is built, so there's no connection to competitors". Mantas Lukauskas, AI tech lead at Hostinger, a website hosting provider that has used large language models since 2020, noted that prices for the very best models were "flat to rising". He added that the recent pricing changes are the "first real test" of whether groups such as Anthropic and OpenAI can protect the cost of their most advanced offerings: "The US labs have cut the middle and are defending the top."
[3]
OpenAI and Anthropic May Be Cheaper Than Chinese AI After All
For months, cheaper token prices have helped Chinese AI models build a reputation as the budget-friendly alternative to OpenAI and Anthropic. But a new study suggests enterprises chasing the lowest sticker price could end up paying more. Research from the AI-powered market intelligence platform AlphaSense found that OpenAI's GPT-5.6 Sol and Anthropic's Opus 4.8 outperformed Chinese rivals Kimi K3 and GLM-5.2 on both cost and quality in complex financial analysis tasks. The findings challenge the widely held assumption that lower token prices automatically translate into lower AI costs. The Cheapest Model Isn't Always the Cheapest to Use At first glance, Chinese models appear significantly cheaper. Moonshot charges $15 per one million output tokens for Kimi K3, compared with $25 for Anthropic's Opus 4.8 and $30 for OpenAI's GPT-5.6 Sol. However, AlphaSense found that token pricing tells only part of the story. After testing 246 financial analysis tasks -- including analyzing earnings transcripts, SEC filings, analyst estimates and acquisition activity -- the company concluded that OpenAI's GPT-5.6 Sol delivered answers with roughly 20% higher quality while costing about 13% less than Kimi K3 on a median basis. Anthropic's Opus 4.8 performed even better, generating responses that scored around 13% higher in quality at roughly half the overall cost of Kimi K3. The reason, according to the study, is that more capable models often require fewer tokens and fewer processing steps to complete the same task, reducing the total cost despite charging higher prices per token. "Some of the more expensive models, the ones that look more expensive based on just their price per token, actually ended up being less costly because they were more efficient in using tokens," AlphaSense CEO Jack Kokko said. AI Buyers May Need a New Way to Measure Cost The findings arrive as enterprises increasingly weigh whether to adopt frontier AI models from companies like OpenAI and Anthropic or to opt for lower-priced, open-weight alternatives from Chinese developers. Rather than focusing solely on token prices, AlphaSense argues businesses should evaluate the total cost of completing a task alongside the quality of the output. In knowledge-intensive workloads such as financial research, more intelligent models may justify their premium by delivering accurate answers more efficiently. At the same time, the report doesn't dismiss open models altogether. Companies running AI workloads on their own infrastructure can avoid per-token charges entirely, while less demanding tasks -- such as email summarization -- may not require frontier models. Instead, AlphaSense says the most cost-effective strategy may be using multiple models together. Its AI search platform employs a routing system that automatically assigns different parts of a query to different models -- for example, using a more capable model to plan a response before handing execution to a smaller, cheaper model. The broader takeaway is that as enterprises move beyond benchmark scores and token pricing, the AI industry's next pricing battle may be decided less by who charges the least -- and more by who delivers the lowest cost per completed task. Image via Shutterstock Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
[4]
AI Labs Stop Competing on Smarts and Start Competing on Price | PYMNTS.com
Venture capital firm Andreessen Horowitz has tracked the same phenomenon since 2024, coining the term "LLMflation" to describe a roughly 10-fold price decline every year for a model of fixed capability, according to its analysis. That trajectory looks familiar from broadband internet: as infrastructure improves and competition intensifies, the price of access keeps falling while what a user can do with it keeps expanding. What changed this summer is the pace. The three largest AI labs all moved in the same direction within weeks of each other this summer. OpenAI cut the price of GPT-5.6 Luna, its fastest and cheapest model, by 80%, from $1 to $0.20 per million input tokens. It cut GPT-5.6 Terra, its mid-tier model, by 20%, from $2.50 to $2, CNBC reported. Input tokens are what a company sends into a model. Output tokens, what the model sends back, run five to six times higher across all three labs. OpenAI left its flagship model, GPT-5.6 Sol, unchanged at $5 per million input tokens. Google Cut Its Gemini Flash Price Weeks After Launch Google arrived at a similar place from a different starting point. Gemini 3.6 Flash, a mid-tier model built to balance cost and capability, launched July 21 at $1.50 per million input tokens, and Gemini 3.5 Flash-Lite, Google's cheapest and fastest model, launched the same day at $0.30 per million input tokens, according to Google's pricing page. Google then cut Gemini 3.6 Flash's price to an introductory $0.75 within weeks. It priced its newest model, Gemini 3.7 Flash, which launched Thursday (Aug. 13), at that same introductory $0.75 per million input tokens through the end of the year, VentureBeat reported. Both rates double to $1.50 on Jan. 1, according to Google's pricing page. That means the newest model does not undercut its immediate predecessor. It undercuts only Flash's original launch price. Anthropic Canceled a Price Increase It Had Already Announced Anthropic went further. It reversed a price increase the company had already scheduled and publicly announced. Claude Sonnet 5 launched June 30 at $2 per million input tokens, explicitly framed as a temporary introductory rate, with a posted increase to $3 scheduled for Sept. 1. On Aug. 10, roughly three weeks before that increase was set to take effect, Anthropic added an editor's note to that same launch announcement canceling the increase and making the $2 rate permanent. The reversal took the only scheduled price increase among the three labs off the table. The Same Pattern That Drove Broadband Adoption Is Now Driving AI The pattern across all three companies is consistent: prices keep falling at the low and middle tiers even as flagship, most-capable models hold steady or fall more slowly. That mirrors broadband internet, where the cost of a baseline connection fell for years while the fastest, most premium tiers commanded a persistent premium. Falling prices at the bottom of the market did more to expand who could use the technology at all than any improvement to the top tier did. Enterprise AI workloads now range from simple jobs that barely tax a model, like sorting emails into categories, to multistep reasoning. The price gap between the cheapest and most expensive tiers has widened even as absolute prices fall across the board. At OpenAI, the spread between its cheapest and most expensive models widened to 25-to-1 from 5-to-1 in a single announcement. That gap is what is letting AI spread into uses that were not affordable a year ago. A task that costs a fraction of a cent to run through a discounted, high-volume model can now be applied at a scale that would have been too expensive at 2023 pricing. It is the same dynamic that allowed cheaper broadband to unlock video streaming and cloud computing once connectivity costs fell far enough. For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.
[5]
US AI Labs Cut Prices 25% in 1 Month to Fend Off Chinese Rivals | PYMNTS.com
OpenAI cut the price of its mid-tier model, GPT-5.6 Luna, by 80%, while Anthropic launched its Claude Opus 5 at half the price of its Fable 5, according to the report. These price cuts came at a time when companies are becoming more cost-conscious after seeing their AI bills rise, and when Chinese developers such as Moonshot and DeepSeek are gaining ground among U.S. and European customers by offering models that are cheaper and increasingly closer to the capabilities of leading U.S. models, the report said. The report said that while OpenAI and Anthropic have lowered the prices of their mid-tier models, they're maintaining or increasing the prices of their top-tier models. It added that a more capable, more expensive model may be able to complete a task with fewer tokens and therefore at a lower cost. PYMNTS has reported over the last two months that the AI price war is real and has reached American consumers, that soaring AI costs have pushed enterprise buyers to cheaper Chinese models and that the era of "tokenmaxxing," or pushing employees toward the biggest AI models and the heaviest usage, is ending after two years of unchecked growth. When OpenAI announced July 30 that it was cutting prices on select models, the company said in a blog post that it made the changes to improve the models' performance per dollar across enterprise workloads. On that day, OpenAI cut the price of GPT-5.6 Luna by 80%, cut the price of GPT-5.6 Terra by 20% and provided faster performance of GPT-5.6 Sol in the API while leaving its price unchanged. GPT-5.6 Sol is the company's frontier model. "We are building a resilient infrastructure portfolio and matching each workload to the systems best suited to run it," OpenAI said in the post. "That approach supports both ends of the price-performance curve." It was reported Thursday (Aug. 13) that DeepSeek is adding peak-hour pricing for its flagship V4 models that will quadruple the current levels, though the new prices remain lower than those of its main competitors.
Share
Copy Link
OpenAI cut GPT-5.6 Luna prices by 80% while Anthropic launched Claude Opus 5 at half its flagship cost as the AI price war intensifies. US labs reduced token prices by nearly 25% since mid-July to compete with Chinese rivals Moonshot and DeepSeek, who are capturing cost-conscious enterprise customers from DoorDash to Airbnb.
The AI price war has escalated dramatically as OpenAI and Anthropic implement aggressive price cuts to defend their market share against surging Chinese AI rivals
1
. OpenAI slashed prices for GPT-5.6 Luna, its fastest and most affordable model, by 80%—dropping from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output tokens2
. Anthropic launched Claude Opus 5 at $5 per million input tokens and $25 per million output tokens, positioning it at half the price of its flagship Fable 5 model1
. These moves helped decrease prices that customers are paying for models from leading US labs by almost 25% since mid-July, according to Silicon Data's token price index2
.
Source: PYMNTS
The price reductions mark a strategic shift for US AI groups that make proprietary closed models and have historically competed heavily on performance rather than cost
1
. Anthropic went further by reversing a scheduled price increase for its Claude Sonnet 5 model, canceling a planned rise from $2 to $3 per million input tokens that was set to take effect in September4
. OpenAI also cut GPT-5.6 Terra, its mid-tier model, by 20% from $2.50 to $2 per million input tokens while leaving its flagship GPT-5.6 Sol unchanged at $54
.Rising AI usage costs have pushed companies to impose caps on AI usage or test cheaper alternatives, with major enterprises like DoorDash and Airbnb publicly stating they have started using Chinese-made models to rein in bills
2
. Chinese developers including Moonshot and DeepSeek are making significant inroads with users from Silicon Valley to Europe by offering increasingly capable open models that can be freely downloaded and tweaked by developers1
. This shift has coincided with a flurry of releases from Chinese labs that have narrowed the performance gap with leading US models, raising concerns in the US tech industry that American developers could lose customers even as they spend heavily to maintain their technological edge2
.
Source: Benzinga
The competitive pressure intensified as Anthropic and OpenAI shift some enterprise customers away from flat subscriptions toward usage-based billing, under which companies pay according to the computational resources they consume
1
. Venture capital firm Andreessen Horowitz has tracked this phenomenon since 2024, coining the term "LLMflation" to describe a roughly 10-fold price decline every year for a model of fixed capability4
.Headline token prices don't provide a straightforward comparison between AI models, as more capable models can sometimes complete a task using fewer tokens or with fewer attempts
2
. Research from AlphaSense found that OpenAI's GPT-5.6 Sol delivered answers with roughly 20% higher quality while costing about 13% less than Moonshot's Kimi K3 on a median basis when analyzing 246 financial analysis tasks3
. Anthropic's Opus 4.8 performed even better, generating responses that scored around 13% higher in quality at roughly half the overall cost of Kimi K33
.Artificial Analysis, which benchmarks models across areas including math, science, coding, and reasoning, found Anthropic's Opus 5 at medium effort delivered similar performance and cost per task to Moonshot's Kimi K3 at max effort
1
. OpenAI's GPT-5.6 Luna at max effort performed similarly to DeepSeek's V4 Flash at max, but cost just under twice as much per task1
. Most models can operate at different effort settings, which vary the computing power used to answer a question and can affect both performance and the ultimate cost of completing a task2
.Related Stories
Mantas Lukauskas, AI tech lead at Hostinger, a website hosting provider that has used large language models since 2020, noted that prices for the very best models were "flat to rising"
1
. He characterized the recent pricing changes as the "first real test" of whether groups such as Anthropic and OpenAI can protect the cost of their most advanced offerings, stating: "The US labs have cut the middle and are defending the top"2
. At OpenAI, the spread between its cheapest and most expensive models widened to 25-to-1 from 5-to-1 in a single announcement4
.
Source: Ars Technica
The price cuts arrive as OpenAI and Anthropic plot initial public offerings at trillion-dollar valuations while investors seek evidence that the industry's vast spending on AI can generate returns
1
. AlphaSense argues businesses should evaluate the total cost of completing a task alongside the quality of the output, suggesting the most cost-effective strategy may be using a multi-model strategy that employs a routing system to automatically assign different parts of a query to different models3
. As enterprise AI adoption accelerates, the pattern mirrors broadband internet, where falling prices at the bottom of the market expanded access more than any improvement to premium tiers4
.Summarized by
Navi
10 Jul 2026•Business and Economy

26 Jul 2026•Policy and Regulation

31 Jul 2026•Business and Economy
