19 Sources
[1]
AI companies are now racing to the bottom -- crashing token prices and competitive models push companies to cut costs
For the rest of us, there's now significant AI intelligence to be found in more affordable models. We've entered a new phase of the AI industry's development, with all the major players heavily cutting costs and boosting the capabilities of their entry-level models in order to compete with new models from China, like Moonshot's Kimi K3 and DeepSeek's V4 Flash. OpenAI did so most recently, cutting the price of its base frontier model, ChatGPT 5.6 Luna, by 80% per million tokens, and its mid-range 5.6 Terra by 20%. This comes just over a week after Google introduced its more-affordable Gemini 3.6 Flash and 3.5 Flash-Lite models. Anthropic hasn't cut prices, but replaced its most-affordable Opus 4.8 model with a more capable Claude 5.0 at the same price point. Intelligence is getting more affordable thanks to increased global competition, but this can come at the cost of margin for these major companies. This follows months of major AI businesses announcing cuts and limits on their use of the technology, even by major AI boosters like Elon Musk's xAI. Despite more workers using AI than ever before, productivity gains are reported to have been less than ideal. Intensifying competition The story of Chinese and American AI development efforts has been somewhat emblematic of the countries' historic strengths. While American firms burn through enormous amounts of money to push frontier technologies, Chinese developers have leveraged their industrial base to develop models that are cheaper, leaner, and almost as good at the top end. DeepSeek gave Western AI developers a shock in 2025, and Kimi K3 did much the same in 2026. Alone, these events would cause concern for companies like OpenAI, Google, and Anthropic. Still, after months of companies that use AI heavily complaining about skyrocketing token costs, the news of an almost-as-good model at a much lower price really made a splash. Now, the big AI developers can't just compete by throwing more parameters and training data at the problem. Now they're having to really compete on price, and to do it, OpenAI has massively reduced the price of its models. Not its most powerful and capable -- the faster version of that is actually becoming more expensive -- but models in its frontier range are now the cheapest they've ever been, and the timeline for this transition of intelligence and pricing is wild. OpenAI launched ChatGPT 5.4 in March with powerful new agentic capabilities for $2.50 per million input tokens and $15 per million output tokens. GPT 5.6 Luna is now just $0.20 and $1.20, respectively. That's a less-than-four-month window for a frontier model to remain cutting-edge and priced accordingly. These latest cuts bring Luna into the realm of DeepSeek V4, with its pro model costing $0.435 per million input tokens and $0.87 per million output tokens. GPT 5.6 Terra is a more capable model, but after its 20% price cut, it's now $2.0 per million input tokens and $12.00 per million output tokens. That undercuts the headline-grabbing K3, which is $3.00 and $15.00, respectively. Meanwhile, GPT 5.6 Sol remains $5 and $30 per million input/output tokens, and OpenAI has actually raised the price of its top model, with 5.6 Sol in Fast mode charging $10 and $60, respectively, to deliver the same kind of intelligence but at a lower latency -- competing directly with other flagship frontier models like Claude Fable 5 and Mythos 5. But is any of this actually going to make OpenAI any money? Bills are coming due After OpenAI announced that it was effectively abandoning its idea of owning first-party data centers earlier this year, the lease contracts it held with Neoclouds became more important than ever. Deals like the enormous $300 billion compute commitment with Oracle became paramount for the very existence of OpenAI's service as a company. But if there were questions about how OpenAI would afford such a venture at the time, they're even more pronounced now. OpenAI is already losing money on its subscription-based accounts and missed key revenue targets earlier this year. And that's after losing 10s of billions in 2025, despite revenue rising consistently throughout the year. OpenAI has committed to some $600 billion in compute spend by 2030. Even if revenue is rising, it might not be rising anywhere near quickly enough to cover these kinds of bills, and cutting the price of the most popular, affordable models suggests margins will either shrink dramatically or disappear altogether. This may be why there's also a lot of talk of Nvidia backstopping OpenAI with a $250 billion investment. OpenAI is far from alone here, either. Google spent around nine times its cloud revenue on AI infrastructure over the past year, while Anthropic has only been able to post profits on annualized revenue recently because of a limited cut-price deal with xAI to rent its Colossus data center. AI is not suddenly cheaper to run or cheaper to build for, and yet companies are slashing prices and making faster, more capable models available for less. On the surface, the numbers just don't add up. Betting on Jevons Paradox The AI industry often cites the Jevons Paradox when it comes to accelerating AI adoption and mass-market use. Where in Jevons' time making more efficient coal-powered engines resulted in more coal use, rather than less of it, AI developers claim that as AI use becomes more efficient, greater uses for it will be found, leading to greater overall use. That may be the future that the token cost-cutting may be hoping to rush us towards. If tokens are cheap, people will use more of them overall, leading to higher earnings. Throw in next-generation AI accelerators becoming more prevalent within AI data centers towards the end of the year, and we could have 10x more tokens per watt, making slimmer margins more profitable by volume. Then there's Vera Rubin to look forward to, which Nvidia claims will deliver another 10x increase in token performance efficiency. It is certainly possible that the advantages of Blackwell and Vera Rubin platforms will make AI a more potentially profitable industry for inference servers. But even then, it's hard to imagine the big companies covering anything close to their enormous investments with direct AI earnings. Especially as increasing competition drives down token pricing.
[2]
China turns up the heat with open model blitz as US model makers panic
The AI arms race reached a fever pitch on Monday after Chinese e-commerce and cloud provider Alibaba called into question America's technological lead with the launch of Qwen 3.8-Max, a 2.4 trillion-parameter model that goes toe-to-toe with the best models from Anthropic and OpenAI. The new model comes just days after the launch of DeepSeek V4 Flash 0731, which, according to independent benchmarks by Artificial Analysis, performs within a single point of OpenAI's budget-friendly GPT-5.6 Luna while costing 40 percent less per task. What's more, at just 284 billion parameters, it's small enough to run on relatively modest enterprise servers and workstations. Chinese model devs like Moonshot, Alibaba, and DeepSeek are now attacking their American counterparts on both price and performance. The pincer movement comes as US model devs like Anthropic and OpenAI stoke fears over the origins and safety of China-made AI models. In a recent blog post, Anthropic CEO Dario Amodei insisted he's not opposed to open models, just ones made in China, ones distilled from proprietary models, and ones that do not meet rigorous safety metrics. In other words, anything that actually competes with Anthropic's own models. The safety bit is particularly disingenuous, as the company's fearmonger-in-chief has gone out of his way to stoke fears among US government officials. Proprietary models can be controlled, but open weights, once released in the wild, are impossible to claw back. But neither Amodei's comments nor commitments from major American and European tech giants change the fact that China is providing the only meaningful competition in the open weights arena. China dominates here. Speaking on CNBC Monday, Clément Delangue, CEO of Hugging Face, the biggest and most influential model repo in the world, said as much. "They're clearly dominating on open models right now, and I wouldn't be surprised if they start dominating at the frontier either by the end of this year or next year at the rate of progress," he told the financial TV news network. America's most capable open weights model, Inkling, is just shy of a billion parameters, and still can't keep up with DeepSeek's finest. Models like Moonshot's Kimi K3 and Z.ai's GLM 5.2 are in an entirely different orbit. In fact, an enterprise's only credible alternatives to proprietary models and their dubious security policies are Chinese models. DeepSeek, Alibaba, Moonshot, MiniMax, and Z.ai aren't going to pass that opportunity up. Alibaba lets its most powerful model loose on the world Of the Chinese model devs, Alibaba is arguably the most like its American competition. Its open-weight models are well regarded, and their range of sizes and permissive licensing have made them attractive for fine-tuning application-specific systems. However, like Google and OpenAI, its top models have been locked behind an API - until now. Seizing the moment, Alibaba is making its most capable model weights available for download for the first time with the launch of Qwen 3.8-Max. The blog post contains the usual array of vaguely intelligible bar charts showcasing how the model compares to OpenAI and Anthropic's own models, as well as a slew of demos that would have made Billy Mays proud. If you want specifics, we recommend checking out the launch blog here. You don't have to look that closely to get what Alibaba is selling: anything OpenAI and Anthropic can do, we can do as well, if not better, cheaper, and on the hardware you own. With that said, 2.4 trillion parameters is a rather big lift for most enterprises, likely requiring 48-64 Nvidia B200-class GPUs if deploying in a customer-facing capacity. For internal workloads, 8-16 B300 or AMD MI355X GPUs would be adequate. If that's a little rich for your blood, Alibaba hasn't forgotten its roots and will be releasing a 27-billion parameter version of the model alongside the Max variant. With benchmarks, model devs can usually find a collection of tests that paint their model in a positive light. And unsurprisingly, Artificial Analysis' own Intelligence leaderboard tells a slightly different story than Alibaba's, with Qwen 3.8-Max matching Anthropic's less-capable, but still formidable, Claude Sonnet 5. Beyond the marketing and pitch demos, Qwen's blog post is surprisingly short on detail. What we do know is it's a multi-modal mixture of experts (MoE) model. That means of the 2.4 trillion total parameters, only 95 billion are actually used to generate tokens for any one request. We can also assume that the model uses the same hybrid Transformer+Mamba architecture as previous Qwen models in order to maintain performance across large contexts. Speaking of which, the model will support context windows up to 1 million tokens, though it's still not clear whether that relies on techniques like rope scaling to extend it or not. The model is currently available via Alibaba's API service QwenCloud for $2 per million input tokens, and $6 per million output tokens. Cached tokens are charged on a sliding scale with $0.25 charged per million implicit cached tokens, $0.17 per million explicit cache token reads, and $2.5 per million explicit cache tokens created. For comparison, Anthropic's Claude Sonnet 5 will set you back $2/M input tokens, $0.20/M cache hits, and $10/M output tokens. And Sonnet 5 pricing is set to increase 50 percent starting September 1. Meanwhile OpenAI's GPT 5.6 Luna, which falls just behind Qwen 3.8-Max and Sonnet 5 on the Artificial Analysis leaderboard, will run you $0.20/M input, $0.02/M cached input, $0.25/M cached write, and $1.20/M output tokens for short context lengths under 272,000 tokens and double that for jobs exceeding that context. The model weights are slated for release on popular model repos, including Hugging Face, starting next week. DeepSeek undercuts OpenAI with flashy new V4 refresh While Alibaba joins Kimi K3-maker Moonshot.AI's assault on frontier models, DeepSeek has taken a very different tack with its latest open weights model: squeeze every ounce of performance from as few weights as possible. The result is DeepSeek V4-Flash-0731, a 284-billion parameter model that can fit into around 142 GB of GPU memory (at FP4). That means enterprises can easily run this model at scale on a single system. And despite its smaller stature, the Flash model actually outperforms the 1.6 trillion-parameter DeepSeek V4 Pro by nearly 14 percent on Artificial Analysis' Intelligence leaderboard. With that said, the refinements made to DeepSeek V4 Flash will no doubt find their way into the Pro model before long. Like Qwen 3.8-Max, DeepSeek V4 Flash undercuts OpenAI and Anthropic on pricing - this time by a considerable margin. For API access, DeepSeek is currently asking $0.14/M input tokens, $0.0028/M cached tokens, and $0.28/M output tokens. However, in an agentic world filled with reasoning models, API pricing doesn't paint a complete picture. A model may appear cheaper, but if it consumes twice as many tokens as another higher priced model, it may not actually be less expensive. Because of this, it's important to look at how efficiently the model solves real world problems. Alarmingly for the US LLM makers, according to Artificial Analysis, DeepSeek V4-Flash isn't just cheaper per token; it is also incredibly efficient at its job. Compared to OpenAI's GPT 5.6 Luna, which is among the most efficient and least expensive models in Sam Altman's current lineup, DeepSeek's latest model is a full 40 percent less expensive, with a cost to solve of just three cents versus five cents. That difference may feel small, but it's worth remembering that for vibe coders consuming tens or hundreds of millions of tokens a day, that difference adds up quite quickly. One of the secrets to DeepSeek's efficiency is the integration of DSpark speculative decoding directly into the model weights. We've explored speculative decoding in the past, but in a nutshell, it involves using a smaller draft model to predict the outputs of a larger, more capable one. When it works, inference performance increases. When it guesses wrong, it falls back to the base model. More importantly, because the smaller model is only predicting the output of the larger one, it's entirely lossless and therefore requires no compromise in terms of performance. Alibaba and others have implemented similar speculative decoding mechanisms, like multi-token-prediction (MTP), for this reason. With DSpark, DeepSeek claims it can extract 57-85% more per-user speed on the exact same hardware, which is an impressive claim, and one that at least in our testing rings true. Your own personal frontier model? DeepSeek V4 Flash is just small enough that running it at home is entirely possible if you've got some deep pockets, or a heck of a lot of memory lying around. Testing on a 128 GB DGX Spark, we were able to get the model running at a respectable 128,000 token context window using Unsloth's IQ3-XXS quant in Llama.cpp. Three bits per weight gives us just enough room to pack the DSPARK draft model into memory, but it's also a bit more compression than we typically recommend for homelab use. Because of this, Unsloth warns that the model's outputs could show some signs of quality loss, but it does run. Unsloth has a full guide on how to get the model up and running, assuming you've got beefy enough hardware. We admit, the DGX Spark is not a cheap box. At $4,699, it's squarely in workstation territory, as is the $3,999 Ryzen AI Halo we looked at early last month. But the fact that you can even contemplate running a model like DeepSeek V4-Flash at home is impressive in its own right. The original IBM PC kitted out with a monitor and diskette drive cost about $3,735 in 1981. Adjusting for inflation, that works out to about $13,700 in today's money. If history repeats itself, the hardware necessary to run models like DeepSeek V4-Flash should become much more accessible within the next decade.®
[3]
OpenAI drops GPT-5.6 Luna and Terra API prices by up to 80%
The lower prices are more likely to accelerate enterprise AI deployments than reduce CIO budgets, analysts say. OpenAI has cut API prices for its GPT-5.6 Terra and Luna models by 20% and 80%, respectively, while also reducing the number of usage credits the models consume in ChatGPT Work and Codex, in an effort to effectively increase the amount of AI work enterprise subscribers can perform without paying more. "Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra, and $0.20 per million input tokens and $1.20 per million output tokens for Luna," the company wrote in a blog post. Prior to this update, enterprise customers paid $2.50 per million input tokens and $15 per million output tokens for GPT-5.6 Terra. For Luna, pricing stood at $1 per million input tokens and $6 per million output tokens.
[4]
OpenAI cuts prices on smaller models as businesses scrutinize AI spend
July 30 (Reuters) - OpenAI slashed prices of its low- and mid-tier AI models on Thursday, a move that may intensify competition in the industry as U.S. companies battle cheaper Chinese rivals for customers increasingly wary of the technology's ballooning costs. The ChatGPT maker lowered the cost of its smaller GPT-5.6 Luna model by 80% and its mid-tier Terra by 20%, while leaving the price of its biggest and flagship Sol model unchanged. The cuts show that rising cost scrutiny by businesses facing hefty AI bills is forcing American labs to rethink pricing. Many tech CEOs have also said in recent months that cheaper AI options are key to the technology's widespread adoption. OpenAI's new pricing also turns up the heat on Anthropic, whose Claude models dominate enterprise and developer use but sit at the costlier end of the market. Both companies have been under pressure from open-source Chinese rivals such as Z.ai's GLM-5.2 that nearly match their performance at a lower cost. Analysts have said that cutting prices could boost usage of OpenAI's and Anthropic's technology, but strain their finances ahead of highly anticipated initial public offerings. While Thursday's cuts affect only OpenAI's smaller and mid-tier models, the company said it would still benefit businesses broadly as those models can now do work that recently required a top-tier system at far lower cost. The new pricing means businesses using OpenAI's technology will have to pay less for every million "tokens", or the units used to measure AI usage, they run through the models. Sending text to Luna drops to 20 cents per million tokens from $1 and Terra's to $2 from $2.50, while generating responses falls to $1.20 and $12 from $6 and $15, respectively. Anthropic's mid-tier Claude Sonnet 4.6 model, meanwhile, costs $3 per million input tokens and $15 per million output tokens, above the rates for Terra. OpenAI said the lower prices were partly enabled by efficiency gains from GPT-5.6, including the model's ability to improve code and optimize performance during internal development. Overall, prices of tokens have been falling in the past year, but the cost of completing a task is rising as AI firms shift from flat subscriptions to usage-based pricing. That is leaving companies with unpredictable and often higher bills as usage per task becomes harder to estimate. Reporting by Aditya Soni in Bengaluru and Deepa Seetharaman in San Francisco; Editing by Devika Syamnath Our Standards: The Thomson Reuters Trust Principles., opens new tab
[5]
OpenAI cuts prices for two of its GPT-5.6 AI models as companies grow sensitive to costs
OpenAI on Thursday announced it is slashing the price of two of its latest artificial intelligence models, GPT-5.6 Terra and GPT-5.6 Luna, roughly three weeks after their public release. The company is facing pressure to cater to a more cost-sensitive customer base, where enterprises have been less inclined to deploy expensive models without a clear picture of the return on their investments. It's also working to fend off competition from Chinese startups and tech giants Google and Microsoft, which have been touting cost-effective models. OpenAI launched three models as part of its GPT-5.6 series, including Sol, the most powerful offering, Terra, the mid-tier model, and Luna, its fastest offering. The company said Thursday that it's reducing the price of Terra by 20% to $2 per million input tokens and $12 per million output tokens. It's cutting the cost of Luna by 80% to 20 cents per million input tokens and $1.20 per million output tokens. Sol's pricing remains the same. "Our strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost," OpenAI said in a release. OpenAI kickstarted the AI boom with the launch of ChatGPT in 2022, prompting companies across the U.S. to rush to deploy the technology and incentivize adoption within their workforces. The era of so-called tokenmaxxing was born, where employers encouraged staffers to use as much AI as possible without worrying about costs.
[6]
Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use
Chinese e-commerce and cloud giant Alibaba's famed Qwen team of AI researchers last night unveiled Qwen3.8-Max, a new flagship 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM) that targets one of the most competitive corners of the frontier AI market: autonomous software engineering and long-horizon enterprise work. If the company's published benchmarks hold up under broader independent testing, Qwen3.8-Max doesn't merely compete with today's leading proprietary models -- it surpasses several of them on some key benchmarks in agentic computing. Most notably, Qwen reports that Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark measuring how well ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0), while also posting the highest reported score on PaperBench and leading or remaining highly competitive across software engineering, research reproduction, multimodal reasoning, and visual web development benchmarks. The release also signals a potentially significant strategic shift for Alibaba: the company says open weights for Qwen3.8-Max will be released next week, alongside Qwen3.8-27B. If that happens under a permissive license, it would represent the first time a Max-class Qwen model becomes available for self-hosted deployment -- a move that could substantially reshape enterprise adoption. One important caveat remains, however: Alibaba has not yet disclosed the licensing terms, leaving open the possibility that the release could use a more restrictive custom license, as we saw recently with Chinese rival Moonshot's open Kimi K3 frontier model, rather than a broadly permissive one such as Apache 2.0. A different definition of 'frontier' Over the past year, the competitive landscape for foundation models has become increasingly specialized. OpenAI has largely focused its GPT series on general reasoning, multimodal interaction and enterprise productivity. Anthropic's Claude series has emphasized coding and dependable long-context reasoning. Google continues to push Gemini toward multimodal productivity and web-native workflows. Moonshot AI's Kimi K3 recently entered the conversation by pairing frontier-class performance with an open-weight release. Qwen3.8-Max attempts to combine many of these strengths into a single model aimed squarely at enterprise automation. Rather than emphasizing conversational intelligence, Alibaba is positioning the model as an autonomous coworker capable of executing projects that span days rather than minutes. According to the company, Qwen3.8-Max can autonomously complete software projects lasting more than 10 days, reproduce research papers involving thousands of lines of code, perform iterative chip-design optimization, and continuously revise plans using multimodal feedback loops. Those demonstrations remain company-produced and have not yet been broadly replicated by independent evaluators. Nevertheless, they illustrate a growing industry trend: frontier models are increasingly competing on their ability to finish entire workflows rather than answer individual prompts. Benchmarks increasingly reward autonomous execution The benchmark suite released alongside Qwen3.8-Max reflects this shift. Instead of focusing solely on traditional reasoning exams or coding puzzles, many of the highlighted evaluations measure long-horizon execution. On OSWorld-Verified, which evaluates computer-use agents interacting with desktop environments, Qwen3.8-Max posts 86.1, ahead of GPT-5.6 Sol Max's 83.2, Fable 5's 85.0, and Gemini 3.1 Pro's 76.2. The model also leads: * PaperBench: 93.0 * TerminalBench 2.1: 86.6 * Vision2Web: 69.0 * LVBench: 81.8 * ERQA: 77.8 Elsewhere, it remains competitive with proprietary leaders while trailing in several categories. On the professional software engineering benchmark SWE-Pro, for example, OpenAI's model posts the highest reported score, while Opus 4.8 continues to lead on certain software engineering evaluations and Agents' Last Exam. Rather than dominating every benchmark, Qwen appears to offer one of the broadest balanced performance profiles currently available. That balance may ultimately matter more for enterprise buyers than isolated benchmark wins. Many organizations increasingly evaluate models based on how reliably they complete heterogeneous workflows -- writing code, reading documents, navigating interfaces, generating reports, inspecting images and coordinating multiple subtasks -- rather than optimizing for one narrow capability. Where Qwen3.8-Max appears strongest Assuming Alibaba's published results translate into production deployments, several enterprise workloads stand out as particularly well suited for Qwen3.8-Max. 1. Long-running software engineering Alibaba's primary demonstration involves autonomous software development extending beyond ten days. While enterprises should treat these demonstrations as vendor claims until independently reproduced, they align with a growing interest in persistent coding agents that operate continuously rather than interactively. Organizations experimenting with autonomous engineering teams, CI/CD automation, repository maintenance, regression testing or feature implementation may find Qwen particularly attractive if its agentic performance proves consistent outside laboratory settings. 2. Computer-use agents The strongest differentiator may be computer use. OSWorld has rapidly become one of the industry's most closely watched benchmarks because it measures a model's ability to interact with operating systems instead of simply generating text. Models capable of reliably navigating desktop software can automate countless repetitive business processes, including document processing, enterprise software integration, internal operations and legacy workflows where APIs may not exist. Leading OSWorld could therefore translate into real operational advantages if benchmark performance generalizes to production environments. 3. Research automation Qwen's PaperBench leadership suggests strong potential for organizations performing scientific computing, literature review, experiment reproduction and technical analysis. Research institutions, pharmaceutical companies and industrial R&D teams increasingly use LLMs not only for summarization but also for executing reproducible computational workflows. Models capable of maintaining context across extended sessions become increasingly valuable in these environments. 4. Multimodal industrial workflows Unlike earlier multimodal systems that primarily analyze uploaded images, Qwen describes vision as an ongoing feedback mechanism integrated into planning and execution. That architecture could prove particularly useful in manufacturing, logistics, engineering inspection and design review, where visual inputs continuously inform operational decisions rather than serving as isolated prompts. The economics may prove just as important Perhaps the biggest competitive pressure comes not from benchmark scores but from pricing through Qwen's application programming interface (API) on QwenCloud (based in China): Qwen3.8-Max launches at $2/$6 per million input/output tokens, a mid-priced model but undercutting the top U.S. proprietary offerings to which it is benchmarked against by meaningful percentages, less than 1/3 the combined in/out price of Claude Opus 5 and less than 1/4 the price of GPT-5.6 Sol Max. Lower inference costs increasingly matter because agentic systems consume dramatically more tokens than conventional chatbots -- a reality that likely factored into OpenAI's decision late last week to cut the API prices of its mid- and lower-end GPT-5.6 lineup of models (Terra and Luna) by 20% and 80%, respectively. Indeed, as those running these systems can attest, multi-hour autonomous workflows, iterative planning and continuous self-correction can generate millions of tokens during a single task. For enterprises deploying hundreds or thousands of agents simultaneously, inference costs often become one of the largest operational expenses. Small reductions in per-token pricing therefore compound rapidly. How it compares with American frontier models Despite headline benchmark comparisons, Qwen3.8-Max should not necessarily be viewed as a wholesale replacement for leading American models. Instead, its strengths suggest different deployment strategies. OpenAI's GPT family continues to excel as a broadly capable enterprise reasoning platform with mature tooling, ecosystem integration and extensive commercial deployment. Organizations already invested in Microsoft ecosystems or OpenAI's enterprise offerings may continue to value those operational advantages even if Qwen leads on selected agent benchmarks. Anthropic's Claude Opus remains widely regarded as one of the strongest coding assistants, particularly for careful software engineering and long-context reasoning. Some enterprises may still prefer Claude for human-in-the-loop development where reliability and predictable behavior outweigh raw autonomy. Google Gemini continues to differentiate itself through deep Workspace integration, multimodal capabilities and Google Cloud services, making it attractive for organizations already standardized on Google's enterprise stack. Where Qwen appears most compelling is for enterprises prioritizing autonomous execution, extended planning horizons and favorable inference economics without sacrificing frontier-level performance. The open-weight question remains unanswered The largest unknown surrounding Qwen3.8-Max has little to do with benchmarks. Alibaba says open weights are coming next week. However, neither the announcement nor the provided documentation specifies the license that will govern those weights. That distinction could prove critical. A permissive license such as Apache 2.0 would significantly broaden enterprise adoption by allowing organizations to self-host, fine-tune and integrate the model into proprietary products with relatively few restrictions. A custom license -- similar to approaches used by several recent frontier releases -- could impose limitations on commercial deployment, redistribution, field of use or model modification. Such restrictions would narrow the appeal for enterprises seeking long-term infrastructure investments, regardless of the model's technical performance. Moonshot AI's recent Kimi K3 release illustrates why this distinction matters. While Kimi K3 made its weights openly available to all, its licensing terms included specific terms including a disclosure and a commercial license requirement for those offering it as a "Model as a Service." Until Alibaba publishes Qwen3.8-Max's license, organizations considering self-hosting should treat the open-weight announcement as promising but incomplete. An increasingly crowded frontier Qwen3.8-Max arrives during one of the fastest-moving periods in the history of foundation models. Within weeks, developers have seen major releases from Moonshot AI, OpenAI, Anthropic and others, each emphasizing different strengths: reasoning, coding, multimodality, autonomous agents or economics. Alibaba's contribution is notable because it combines competitive benchmark performance, aggressive pricing, a million-token context window and a stated commitment to releasing weights for its flagship model. Whether it becomes the preferred platform for enterprise autonomous agents will ultimately depend less on leaderboard positions than on broader independent validation, production reliability and the licensing terms accompanying the forthcoming weight release. Those factors -- not benchmark charts alone -- will determine whether Qwen3.8-Max becomes a genuine alternative to the leading American proprietary models or simply another impressive entrant in an increasingly crowded frontier AI race.
[7]
DeepSeek just slashed AI prices again, and China's AI race is getting even messier
China's AI race has never really been short on competition, but it's increasingly looking like a battle over who can charge the least rather than who can build the best model. DeepSeek has unveiled its latest open-weight AI model, called V4-Flash, alongside a dramatic price reduction that makes it significantly cheaper for developers to use. The company has cut token pricing by 50% and has also decided against introducing a previously announced dynamic pricing system that would have increased costs during periods of heavy demand. The new changes paint a clear picture: DeepSeek appears more interested in winning over developers and expanding its user base than maximizing revenue in the short term. DeepSeek is making AI cheaper, but not necessarily better The new V4-Flash model isn't just about lower prices. DeepSeek says it also delivers stronger agent capabilities, allowing AI systems to better handle multi-step tasks and more complex workflows with less user intervention. Even so, the model isn't currently viewed as the strongest option coming out of China. That title is still widely associated with Moonshot AI's Kimi K3, which continues to lead the domestic field in overall performance. Instead of trying to leapfrog competitors on raw capability alone, DeepSeek appears to be taking a different approach. By making its latest model substantially more affordable while keeping it openly available, the company is lowering the barrier for developers, startups, and businesses looking to build AI-powered products. For developers, that's welcome news. Lower operating costs mean cheaper experimentation, larger deployments, and fewer worries about usage bills climbing during busy periods. China's AI price war shows no signs of slowing down DeepSeek's latest move also highlights a broader trend that's reshaping China's AI industry. Competition has become so aggressive that companies are increasingly undercutting one another on pricing in an effort to capture market share. The situation has become serious enough that Chinese officials have publicly warned technology companies about "involution" -- a term used to describe destructive competition where businesses continue cutting prices without creating proportional value. The irony, however, is difficult to ignore. While Beijing has cautioned firms against this race to the bottom, it has also invested heavily in China's AI ecosystem and continues supporting compute infrastructure and energy costs through subsidies. According to industry experts, those incentives can make it easier for companies to keep operating even when profits are thin or nonexistent. For now, developers are the biggest winners, with increasingly powerful AI models becoming far more affordable. Whether that pricing strategy leads to a healthier AI industry in the long run, however, remains a much bigger question.
[8]
AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost
To quote an ancient Jedi Master "Begun, the AI price wars have!" OpenAI is sharply reducing the prices of two models in its GPT-5.6 frontier series, cutting GPT-5.6 Luna, the smallest and fastest model in the series, by 80% and GPT-5.6 Terra, the mid-tier model, by 20%, while adding a premium Fast mode for its flagship GPT-5.6 Sol model. The cuts place Luna much closer to the lowest-cost commercial models in the market and arrive just a few days after Anthropic released its highly performant Claude Opus 5 at the same price as Opus 4.8, and Google introduced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, two rival models built around lower inference costs, faster execution and more efficient agent workloads. OpenAI is successfully undercutting Google's price per intelligence and attempting to sway Anthropic users, who may not mind paying more, with a speed boost. OpenAI says Luna will now cost $0.20 per million input tokens and $1.20 per million output tokens, for a combined input-plus-output price of $1.40 per million tokens. Terra will cost $2 per million input tokens and $12 per million output tokens, for a combined price of $14. Pricing for Sol Standard remains unchanged at $5 per million input tokens and $30 per million output tokens. OpenAI is also adding Sol Fast mode at twice the Standard price: $10 per million input tokens and $60 per million output tokens. The company says Fast mode delivers up to 2.5 times the throughput without changing the model's underlying intelligence. OpenAI co-founder and CEO Sam Altman took to X to announce the changes as "major price cuts today." VentureBeat Frontier AI model API pricing comparison Pricing is shown per one million tokens. Total cost is calculated as input price plus output price. Cached-input pricing is excluded to keep the comparison consistent across providers. OpenAI moves Luna into the low-cost tier The most consequential change is the Luna price cut. When OpenAI introduced the GPT-5.6 series, Luna was priced at $1 per million input tokens and $6 per million output tokens, for a combined total of $7. The new pricing reduces that combined figure to $1.40. That places Luna below Google's Gemini 3.5 Flash-Lite, which costs a combined $2.80 per million input and output tokens, and far below Gemini 3.6 Flash at $9. Luna also now costs less than OpenAI's own GPT-5.4 and Terra models by a wide margin. It is not the cheapest model in the broader market. Xiaomi's MiMo-V2.5 Flash, DeepSeek's flash model and several other APIs remain less expensive on a pure token basis. But the reduction brings an OpenAI frontier-series model into direct competition with the market's low-cost inference tier. OpenAI says the GPT-5.6 series represents its frontier model family, with Sol positioned at the top of the lineup, Terra as the middle tier and Luna as the smallest and fastest option. The lineup was initially released in late June 2026 through a limited rollout by U.S. government request, before broader access, with each model intended to offer a different tradeoff among intelligence, latency and cost. Sol is aimed at the most complex reasoning-heavy and agentic workloads, including advanced coding, multi-step planning and tool-using systems, while Terra is designed for general production use where a balance of capability and efficiency is required. Luna is positioned for high-throughput, low-latency tasks such as summarization, classification, routing, and lightweight real-time assistants where cost per request is the primary constraint. Terra drops to match Google's Gemini 3.1 Pro pricing Terra's 20% reduction moves its combined price from $17.50 to $14 per million tokens. At that level, Terra now matches Google's Gemini 3.1 Pro Preview pricing for context windows of 200,000 tokens or less. It also undercuts OpenAI's GPT-5.4, which remains priced at $2.50 per million input tokens and $15 per million output tokens, offering the same intelligence for about 1/13th the cost, as Krea AI's Nic Dunz noted on X: The adjustment creates a wider separation between OpenAI's three GPT-5.6 tiers. Luna costs one-tenth as much as Terra on a simple combined input-plus-output basis, while Terra costs 60% less than Sol Standard. Sol Fast moves in the opposite direction. At a combined $70 per million tokens, it is the most expensive model configuration in the comparison below, reflecting OpenAI's decision to charge a premium for latency-sensitive workloads rather than lower Sol's base price. Cuts follow Google's low-cost Gemini releases and Anthropic's Claude Opus 5 OpenAI's pricing changes come only about a week and a half after Google introduced its own low-cost Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens. Google framed both models around the economics of agent deployment, arguing that lower token usage, fewer reasoning steps and reduced tool calls could lower the total cost of long-running software engineering and knowledge-work tasks. Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching as high as 65% on some long-horizon engineering workloads. Gemini 3.5 Flash-Lite is positioned as the fastest model in Google's 3.5 series. However, OpenAI's models are more performant than Google's, according to third party analysis outfits like Artificial Analysis, with even the Luna model outperforming Gemini 3.6 Flash and the older Gemini 3.1 Pro model, making the cost-per intelligence much more favorable to OpenAI. As AI coding startup Cognition noted on X, GPT-5.6 now "sits on the pareto curve of price/performance efficiency," posting an animation of the GPT-5.6 series moving left on a chart representing intelligence on the y axis and cost on the x, showing that the models now offer among the most superior intelligence for lowest cost on the market. And yet, rival Anthropic's Claude Opus 5 remains about as performant as GPT-5.6 Sol, yet is 6% cheaper. The model costs $5 per million input tokens and $25 per million output tokens -- the same rates as Opus 4.8 -- but Anthropic says it delivers nearly all the intelligence of its more expensive Fable 5 model at roughly half the cost. Unlike OpenAI's Luna and Terra changes, Anthropic did not reduce the Opus API sticker price. Instead, it effectively lowered the price per unit of capability by replacing Opus 4.8 with a more capable model at the same $30 combined input-and-output rate. Anthropic also added an adjustable effort setting that allows developers to trade reasoning depth for speed and token savings. That distinction matters for enterprise buyers. OpenAI is directly cutting per-token rates, Google is pairing lower prices with reductions in token use and tool calls, and Anthropic is emphasizing stronger task performance at an unchanged price. All three approaches target the same operational metric: the total cost of completing production work, rather than the advertised cost of an individual token alone. The timing highlights how quickly pricing has become a competitive lever among frontier model providers. OpenAI's response does not introduce a new model generation. Instead, it changes the economics of deploying models that were released only recently. The market shifts from model access to model economics The cuts indicate that access to frontier-level capability is no longer the only point of competition. The next question for enterprises is how cheaply and predictably those models can run in production. OpenAI is still not the lowest-priced provider on a pure token basis. But Luna's 80% reduction materially changes its position, moving it from the middle of the market into a pricing tier populated by smaller models from Google, Xiaomi, DeepSeek, MiniMax and other vendors. That matters most for high-volume applications, where relatively small differences in token pricing can compound across coding agents, document systems, internal search tools and automated workflows. OpenAI's latest move therefore looks less like a routine adjustment and more like a repositioning of the GPT-5.6 series. Sol remains the premium option, Terra moves closer to competing pro-tier systems, and Luna becomes the company's direct answer to the industry's growing low-cost model segment.
[9]
OpenAI cuts prices on smaller models as businesses scrutinise AI spend
The ChatGPT maker lowered the cost of its smaller GPT-5.6 Luna model by 80% and its mid-tier Terra by 20%, while leaving the price of its biggest and flagship Sol model unchanged. OpenAI slashed prices of its low- and mid-tier AI models on Thursday, a move that may intensify competition in the industry as US companies battle cheaper Chinese rivals for customers increasingly wary of the technology's ballooning costs. The ChatGPT maker lowered the cost of its smaller GPT-5.6 Luna model by 80% and its mid-tier Terra by 20%, while leaving the price of its biggest and flagship Sol model unchanged. The cuts show that rising cost scrutiny by businesses facing hefty AI bills is forcing American labs to rethink pricing. Many tech CEOs have also said in recent months that cheaper AI options are key to the technology's widespread adoption. OpenAI's new pricing also turns up the heat on Anthropic, whose Claude models dominate enterprise and developer use but sit at the costlier end of the market. Both companies have been under pressure from open-source Chinese rivals such as Z.ai's GLM-5.2 that nearly match their performance at a lower cost. Analysts have said that cutting prices could boost usage of OpenAI's and Anthropic's technology, but strain their finances ahead of highly anticipated initial public offerings. While Thursday's cuts affect only OpenAI's smaller and mid-tier models, the company said it would still benefit businesses broadly as those models can now do work that recently required a top-tier system at far lower cost. The new pricing means businesses using OpenAI's technology will have to pay less for every million "tokens", or the units used to measure AI usage, they run through the models. Sending text to Luna drops to 20 cents per million tokens from $1 and Terra's to $2 from $2.50, while generating responses falls to $1.20 and $12 from $6 and $15, respectively. Anthropic's mid-tier Claude Sonnet 4.6 model, meanwhile, costs $3 per million input tokens and $15 per million output tokens, above the rates for Terra. OpenAI said the lower prices were partly enabled by efficiency gains from GPT-5.6, including the model's ability to improve code and optimize performance during internal development. Overall, prices of tokens have been falling in the past year, but the cost of completing a task is rising as AI firms shift from flat subscriptions to usage-based pricing. That is leaving companies with unpredictable and often higher bills as usage per task becomes harder to estimate.
[10]
OpenAI Slashes GPT 5.6 Luna API Pricing by a Massive 80%
OpenAI's recent moves are reshaping the AI landscape, with significant price cuts and performance upgrades to its GPT 5.6 models. Notably, the GPT 5.6 Luna model has seen an 80% price reduction, made possible by advancements in the GPT 5.6 Sol infrastructure, which enhances efficiency without sacrificing quality. Meanwhile, the introduction of GPT 5.6 Sol Fast Mode offers results 2.5 times faster than standard processing, catering to industries like finance and healthcare that rely on real-time decision-making. Universe of AI explores how these updates reflect OpenAI's strategy to balance affordability, speed and precision in an increasingly competitive market. Dive into this explainer to understand how these developments intersect with the upcoming release of GLM 5.5 from Z.ai, a large language model boasting 1.6 trillion parameters and optimized for coding long-horizon agents. You'll also gain insight into GLM 5.5's compatibility with Huawei's domestic chips and the challenges its scale presents for smaller users. Additionally, learn how OpenAI's experimental Zinc and Magnesium models are pushing boundaries in AI-driven game development, signaling new possibilities for interactive and complex environments. GLM 5.5: Redefining Large Language Models The upcoming release of GLM 5.5, scheduled for early August, represents a major milestone in the development of large language models (LLMs). With an impressive 1.6 trillion parameters, more than double the size of its predecessor, GLM 5.2, this model is designed to push the boundaries of AI capabilities. It is positioned as a direct competitor to China's Kimmy K3 and is specifically optimized for coding long-horizon agents, making it a powerful tool for solving complex AI challenges. One of the most notable features of Z.ai's GLM 5.5 is its compatibility with Huawei's domestic chips, reflecting a strategic emphasis on hardware optimization. This alignment with local hardware solutions highlights its potential to strengthen China's AI ecosystem. However, the model's massive size introduces challenges for smaller-scale users, as its deployment requires substantial computational resources. Despite these limitations, GLM 5.5 is expected to dominate the open-weight model space, solidifying its role as a key player in the global AI landscape. OpenAI's Aggressive Price Reductions OpenAI has taken decisive steps to make its GPT 5.6 models more accessible by implementing significant price cuts. The GPT 5.6 Luna model now features an 80% price reduction, while the GPT 5.6 Terra model has seen a 20% decrease in cost. These reductions are made possible by advancements in the GPT 5.6 Sol infrastructure, which enhances processing efficiency without compromising performance. The Luna model, in particular, delivers high-quality performance at a fraction of the cost of competing models, making it an attractive option for developers and businesses seeking cost-effective AI solutions. This pricing strategy not only reflects OpenAI's commitment to affordability but also serves as a calculated response to the growing competition from open source AI initiatives. By lowering barriers to entry, OpenAI is positioning itself as a leader in the widespread access of advanced AI technologies. Here are additional guides from our expansive article library that you may find useful on GPT 5.6. GPT 5.6 Sol Fast Mode: Speed Without Compromise The introduction of GPT 5.6 Sol Fast Mode marks a significant leap forward in processing speed. This new API mode is capable of delivering results 2.5 times faster than standard processing, making it particularly valuable for applications that require rapid decision-making. While Fast Mode comes at twice the cost of the standard mode, it maintains the same level of intelligence and accuracy, making sure that users do not have to compromise on quality. Fast Mode is especially beneficial for industries that rely on real-time data analysis and decision-making, such as financial markets, healthcare diagnostics and customer support. By combining speed with precision, this feature sets a new benchmark for high-performance AI applications, allowing businesses to operate more efficiently in time-sensitive environments. AI in Game Development: Zinc and Magnesium Models OpenAI is expanding its reach into the gaming industry with the experimental Zinc and Magnesium models. These models are being tested on platforms like Design Arena, with the goal of creating playable, AI-generated games, including Minecraft-like clones. Although early results have been mixed, the potential for AI-driven game design remains substantial. The application of AI in game development represents a promising frontier, offering the possibility of generating complex, interactive environments and streamlining the game creation process. OpenAI's efforts in this area aim to establish a foothold in a competitive market, where the ability to innovate could redefine the gaming experience for developers and players alike. Frontier Labs vs Open source Models The competition between proprietary frontier labs, such as OpenAI and open source AI initiatives is becoming increasingly intense. OpenAI's focus on efficiency and affordability is a direct response to the growing influence of open source models like China's Kimmy K3. By making advanced AI technologies more accessible, OpenAI seeks to maintain its leadership position in the rapidly evolving AI ecosystem. Conversely, open source models continue to drive innovation and collaboration, challenging proprietary systems and fostering advancements across the industry. This dynamic interplay between frontier labs and open source initiatives is accelerating the pace of AI development, making sure that the technology evolves in diverse and impactful ways. Additional Innovations from OpenAI OpenAI has introduced several updates to its existing tools, further enhancing their functionality and accessibility. The ChatGPT app and Codec CLI, now powered by GPT 5.6 Luna, have undergone significant improvements. These updates have reduced operational costs by a factor of 10, making them more accessible to developers and end-users alike. The ChatGPT app now features a new auto-review capability, which streamlines content evaluation and improves user experience. These enhancements reflect OpenAI's ongoing commitment to delivering practical, user-friendly solutions that cater to a wide range of needs. By continuously refining its offerings, OpenAI is making sure that its technologies remain relevant and valuable in an increasingly competitive market. The Road Ahead The rapid advancements in AI, exemplified by OpenAI's cost-efficient models and the new GLM 5.5, highlight the fantastic potential of this technology. As competition intensifies between proprietary frontier labs and open source initiatives, the future of AI promises to be both dynamic and accessible. These innovations are reshaping industries and redefining the role of AI in everyday life, offering new opportunities for developers, businesses and enthusiasts alike. Media Credit: Universe of AI Disclosure: Some of our articles include affiliate links. If you buy something through one of these links, Geeky Gadgets may earn an affiliate commission. Learn about our Disclosure Policy.
[11]
OpenAI Goes After The Jugular Of China's Open-Weight AI Models, Cuts Token Prices By Up To 80%
Let the AI price wars begin. OpenAI has just fired a veritable volley across the bow of China's growing number of AI labs, sacrificing its sky-high margins to try to starve the budding open-weight AI economy. OpenAI is now trying to starve out China's AI labs by cutting token prices by as much as 80% on GPT-5.6 Luna OpenAI has just slashed the token-based pricing for GPT-5.6 Luna by 80 percent, with input tokens now priced at just $0.2 per 1 million from their earlier perch at $1, and output tokens priced at just $1.20 per 1 million vs. the earlier price of $6. With these pricing cuts, OpenAI's GPT-5.6 Luna now gives the biggest bang for your buck, eclipsing even the relative attractiveness of DeepSeek's V4 Pro model. Of course, this comes as China's open-weight models appear to be going in the opposite direction, with Moonshot having priced the Kimi K3 at $3 per 1 million of input tokens and a whopping $15 for every 1 million of output tokens. With around 3GW of dedicated inference-related compute running on NVIDIA's GPUs, OpenAI had quite a lot of room heretofore to implement these cuts. It also has around 2GW of training compute. In contrast, China's AI labs are probably running with just around 500MW of inference-related compute, that too made up of older NVIDIA GPUs as well as local ones from Huawei. Even so, one should never bet against China in a pricing war. And so, we now wait for China's response to OpenAI's aggressive gambit. Follow Wccftech on Google to get more of our news coverage in your feeds.
[12]
QUICK SPARK: OpenAI Cuts AI Model Prices as Businesses Push Back on Rising Costs
OpenAI is cutting prices on some of its artificial intelligence models as businesses grow increasingly cautious about rising AI expenses and competition intensifies across the industry. The ChatGPT maker lowered the cost of its smaller GPT-5.6 Luna model by 80% and reduced pricing for its mid-tier Terra model by 20%, while keeping its flagship Sol model unchanged, Reuters reported. OpenAI said efficiency improvements in its latest models helped enable the price cuts, allowing smaller systems to handle tasks that previously required more expensive, higher-capability models. The price cuts come as enterprises face growing uncertainty around AI expenses, with many moving away from flat subscription plans toward usage-based models that can lead to unpredictable bills as adoption increases. For investors, the pricing reset highlights the intensifying battle for AI market share. OpenAI is competing against Anthropic's Claude models, which have gained traction among enterprise customers, as well as lower-cost open-source models from Chinese AI developers challenging U.S. AI leaders on price and performance. While cheaper models could drive broader adoption and increase usage volumes, analysts warn the strategy could pressure margins for AI companies investing billions into infrastructure and computing capacity. This content was partially produced with the help of AI tools and was reviewed and published by Benzinga editors. Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
[13]
OpenAI Cuts Prices on Select Models to Make High-Volume Work Economical | PYMNTS.com
The company said in a Thursday blog post that it made these changes to improve the models' performance per dollar across enterprise workloads. Users can select the right model for the outcome they seek, balancing the stakes, cost of error, urgency and scale of each workflow, and the changes to GPT-5.6 provides them with more flexibility to optimize that calculation, according to the post. "The GPT-5.6 family expands the range of those choices," OpenAI said in the post. "Businesses can apply the maximum useful intelligence at every stage while paying the right price for the value it creates." OpenAI described GPT-5.6 Luna as its fastest and most affordable model, GPT-5.6 Terra as its balanced model for everyday work, and Sol as its frontier model. "We are building a resilient infrastructure portfolio and matching each workload to the systems best suited to run it," OpenAI said. "That approach supports both ends of the price-performance curve. At the lower-cost end, the new Luna and Terra prices make high-volume work economical at much greater scale. At the frontier end, Fast mode gives API customers faster access to Sol when response time is important." OpenAI announced July 8 that it was publicly launching the GPT-5.6 Sol, Terra and Luna models the following day after initially limiting their release at the request of the U.S. government. The company had said on June 26 that it previewed the models' capabilities as part of its ongoing engagement with the government. When OpenAI released the models on July 9, OpenAI CEO Sam Altman told CNBC that GPT-5.6 Sol was 54% more token efficient on agentic coding jobs and "as good or better" than competing models on the market. "Every enterprise now is thinking about spend and the value they're getting in exchange for AI, and this is what we really want to do," Altman said. The PYMNTS Intelligence report "New Data Shows How Tech Sectors Are Turning AI Into Strategy" found that when choosing to fund AI projects over the next 12 months, a majority of firms filter the technology through their unique definition of value. The report found that 60% of cybersecurity firms, 65% of software-as-a-service firms and 65% payments firms seek near-term financial results. For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.
[14]
OpenAI cuts GPT-5.6 Luna pricing by 80%, introduces Fast mode for Sol API
OpenAI has announced pricing reductions for its GPT-5.6 model family, reducing the cost of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. The company says the changes are based on improvements across models, inference systems and agent workflows to reduce the cost and time required for AI tasks. The lower pricing for Luna and Terra is also reflected in how usage is counted against paid subscription limits when using Codex and ChatGPT Work. GPT-5.6 Luna supports tool usage and multi-step workflows for high-volume tasks, while GPT-5.6 Terra is designed for everyday workloads. OpenAI has also introduced Fast mode for GPT-5.6 Sol in the API, replacing Priority Processing. Fast mode provides up to 2.5× faster speeds compared with Standard processing at twice the price, without any change in intelligence. Existing API requests marked as priority will continue to work and will automatically use Fast mode. GPT-5.6 efficiency improvements OpenAI says GPT-5.6 improvements come from optimising models, inference systems and the agent systems that connect models with tools and context. The improvements include: * Better routing to improve hardware utilisation * Optimised production software for more efficient token generation * Improved context management to avoid repeating completed work * Reduced time, token usage and compute required for tasks OpenAI says GPT-5.6 Sol has helped with internal optimisation efforts by rewriting and optimising production kernels, designing and running hundreds of experiments to improve token generation, and monitoring training while intervening when problems occur. The company says these efforts reduced the end-to-end cost of serving GPT-5.6 Sol by 20%, while experiments improved token-generation efficiency by more than 15%. GPT-5.6 model selection and performance OpenAI says choosing an AI model depends on factors such as the expected outcome, cost of errors, urgency and workload scale. Different tasks may require different balances of intelligence, speed, reliability and cost. According to OpenAI, GPT-5.6 Luna delivers performance comparable to models that were frontier-class a year ago at around 6 cents per dollar per task and with nearly nine times faster execution. On the Agents' Last Exam evaluation for professional work, OpenAI says Luna outperforms Fable 5 with an estimated cost per task that is nearly 99% lower. The company says businesses can use evaluations to identify where additional intelligence improves results and where faster, lower-cost processing can provide the required outcome. For example, a coding workflow could use GPT-5.6 Sol to resolve uncertainty and create a plan, then use GPT-5.6 Luna to implement defined changes, write and run tests, and evaluate results. Pricing Starting July 30, GPT-5.6 API pricing is: * GPT-5.6 Terra: $2 per million input tokens and $12 per million output tokens * GPT-5.6 Luna: $0.20 per million input tokens and $1.20 per million output tokens * GPT-5.6 Sol: Pricing remains unchanged ChatGPT and Codex subscription prices and quota budgets remain unchanged, while GPT-5.6 Luna and Terra usage now consumes fewer credits under paid subscription plans. Availability GPT-5.6 Terra and Luna are available through ChatGPT Work, Codex and the OpenAI API. In ChatGPT Work and Codex, Free and Go users can access Terra, while Plus, Pro, Business and Enterprise users can choose Terra and Luna. The updated pricing will begin rolling out to AWS. Fast mode for GPT-5.6 Sol is available through the API and aligns with in Codex.
[15]
OpenAI sharply cuts prices for some AI models to accelerate adoption
OpenAI cut the price of its GPT-5.6 Luna model by 80% and its midtier Terra model by 20%, while keeping pricing for its top-end Sol model unchanged. Companies will now pay 20 cents per million input tokens for Luna, down from $1 previously, while the cost to generate responses falls from $6 to $1.20. For Terra, pricing drops to $2 per million input tokens and $12 for generated responses, versus $2.50 and $15 previously. The pricing revision reflects growing pressure from companies looking to keep AI spending under control, as well as competition from cheaper Chinese models. It also increases pressure on Anthropic, whose Claude models remain among the most expensive on the market. OpenAI says improvements in its models now allow tasks that were previously reserved for the most advanced systems to be performed at a significantly lower cost. The group said the price cuts were made possible by efficiency gains in GPT-5.6, notably in software development and the optimization of its own infrastructure. While the per-token cost continues to decline, overall corporate spending on AI is still trending higher, driven by a gradual shift to usage-based billing that makes costs more variable and harder to forecast.
[16]
OpenAI sharply cuts prices for some AI models to speed adoption
OpenAI has cut the price of its GPT-5.6 Luna model by 80% and its mid-tier Terra model by 20%, while keeping the price of its top-end Sol model unchanged. Companies will now pay 20 cents per million input tokens for Luna, down from $1 previously, while the cost to generate responses falls from $6 to $1.20. For Terra, prices are reduced to $2 per million input tokens and $12 for generated responses, versus $2.50 and $15 previously. This pricing overhaul reflects mounting pressure from companies looking to rein in their AI spending, as well as competition from cheaper Chinese models. It also increases pressure on Anthropic, whose Claude models remain among the most expensive in the market. OpenAI believes improvements in its models now make it possible to carry out tasks once reserved for the most advanced systems at a significantly lower cost. The group says these price cuts were made possible by efficiency gains in GPT-5.6, notably in software development and optimization of its own infrastructure. While the per-token cost continues to fall, overall corporate AI spending is still trending higher, driven by the gradual shift to usage-based billing, which makes costs more variable and harder to forecast.
[17]
OpenAI cuts prices on smaller models as businesses scrutinize AI spend
July 30 (Reuters) - OpenAI slashed prices of its low- and mid-tier AI models on Thursday, a move that may intensify competition in the industry as U.S. companies battle cheaper Chinese rivals for customers increasingly wary of the technology's ballooning costs. The ChatGPT maker lowered the cost of its smaller GPT-5.6 Luna model by 80% and its mid-tier Terra by 20%, while leaving the price of its biggest and flagship Sol model unchanged. The cuts show that rising cost scrutiny by businesses facing hefty AI bills is forcing American labs to rethink pricing. Many tech CEOs have also said in recent months that cheaper AI options are key to the technology's widespread adoption. OpenAI's new pricing also turns up the heat on Anthropic, whose Claude models dominate enterprise and developer use but sit at the costlier end of the market. Both companies have been under pressure from open-source Chinese rivals such as Z.ai's GLM-5.2 that nearly match their performance at a lower cost. Analysts have said that cutting prices could boost usage of OpenAI's and Anthropic's technology, but strain their finances ahead of highly anticipated initial public offerings. While Thursday's cuts affect only OpenAI's smaller and mid-tier models, the company said it would still benefit businesses broadly as those models can now do work that recently required a top-tier system at far lower cost. The new pricing means businesses using OpenAI's technology will have to pay less for every million "tokens", or the units used to measure AI usage, they run through the models. Sending text to Luna drops to 20 cents per million tokens from $1 and Terra's to $2 from $2.50, while generating responses falls to $1.20 and $12 from $6 and $15, respectively. Anthropic's mid-tier Claude Sonnet 4.6 model, meanwhile, costs $3 per million input tokens and $15 per million output tokens, above the rates for Terra. OpenAI said the lower prices were partly enabled by efficiency gains from GPT-5.6, including the model's ability to improve code and optimize performance during internal development. Overall, prices of tokens have been falling in the past year, but the cost of completing a task is rising as AI firms shift from flat subscriptions to usage-based pricing. That is leaving companies with unpredictable and often higher bills as usage per task becomes harder to estimate. (Reporting by Aditya Soni in Bengaluru and Deepa Seetharaman in San Francisco; Editing by Devika Syamnath)
[18]
The great AI price war: How Chinese open models are squeezing GPT-5.6 Sol and Claude Fable 5
The age when the frontiers of artificial intelligence research belonged to select American laboratories was one of great monopoly, where OpenAI and Anthropic set the trend and price standards, and others in the industry followed suit. All that has changed in July 2026, with two Chinese laboratories producing models that compete with the likes of Silicon Valley on benchmark performance but do so at a much lower cost, with perhaps most critically, access to the weights. This is not a tale of how China is catching up. This is a tale of how China has rewritten the rules of the game altogether. Also read: OpenAI's Astra Model solved 10 open math problems: This is how we know it's true The premium tier: Sol and Fable 5 play it closed GPT-5.6 Sol by OpenAI and Claude Fable 5 by Anthropic occupy the top rows of most benchmark tables. Sol has a slight edge in the Coding Agent Index of Artificial Analysis, and both are on top in GDPval-AA v2, which is an evaluation benchmark based on real-world applications across 44 occupations. However, both are expensive closed models. Sol costs $5 for every million input tokens and $30 for every million output tokens in the API pricing, but input cache costs only $0.50 per million tokens. Fable 5 is similarly priced at the frontier tier. None comes with model weights. You pay for access, but you do not have the model, and you never will. This used to be the price you had to pay for frontier-level AI. No longer. The challengers: cheaper, open, and closing the gap fast The Moonshot AI Kimi K3 came out on July 16 as, according to the organization itself, the biggest open-source model to be released: 2.8 trillion parameters, sparsity mixture of experts that activates only 16 out of 896 experts per token, and a massive 1-million-token context window. The cost of using it is $3 for every million input tokens and $15 for every million output, with a cached input at $0.30. According to GDPval-AA v2, its score is third overall, right below Fable 5 and Sol, and above Claude Opus 4.8. Also read: 65 pct of Samsung's flagship sales come from Tier 2 and Tier 3 markets, says JB Park on India's foldable future Along came DeepSeek, which approached the goal in a very different manner. The DeepSeek-V4-Flash-0731, launched on July 31, did not scale up the model. Instead, it maintained the same architecture of 284 billion total parameters and 13 billion active parameters of the previous version and only retrained the model more intensively on data related to agentic and reasoning abilities. The result is a performance score of 82.7 on Terminal-Bench 2.1 compared to 61.8 in the previous version and 7.3 to 54.4 on DeepSWE. All of this at a cost of $0.14 per million input tokens and $0.28 per million output tokens under an MIT license. Read those two pricing numbers again. DeepSeek's flagship agentic model costs roughly 1 percent of what Sol charges for output tokens. That is not a rounding difference. That is a different economic category altogether. Why this matters more than the benchmark charts It might be tempting to dismiss this as China's desperate attempt at keeping up with price, since it is not capable of matching in raw capability. However, this does not match the data from the benchmarking exercise. First, Kimi K3 is defeating Claude Opus 4.8, which until recently was believed to have frontier-level capabilities. Second, the Flash model created by DeepSeek is defeating its larger version, V4-Pro. Both of these organizations do not appear to be selling a second-rate product. They are selling a frontier-like capability at commodity prices since they do not have a subscription-based revenue model like OpenAI and Anthropic. There's the accessibility too. Kimi K3's weights are accessible. DeepSeek's model is MIT-licensed, which means that any firm from anywhere in the world, even India, can download it, fine-tune it using their proprietary data, and deploy it on their own infrastructure, completely free from reliance on a US API, a US data-retention regime, or a US export control regime. This isn't an abstract point. The release of GPT-5.6 required two weeks of vetting from the US government before OpenAI could release it. Anthropic's own Fable 5 model had to be taken down for three weeks in June under export controls from the Department of Commerce. When your American frontier model can be shut down by Washington overnight, the case for an alternative writes itself. The market share math is already shifting It is not an academic exercise in pricing. Moonshot's previous open release, Kimi K2.6, reached the second-most popular position among the models on OpenRouter, ahead of virtually all Western frontier labs in independent usage metrics, within months of its release. Moonshot is apparently raising up to $2 billion in funding, valued at $31.5 billion precisely because of its open-source momentum. Every developer who uses a $0.14 input model versus a $5 input model is a developer OpenAI and Anthropic do not bill. American labs currently hold an advantage at the upper end of the capability curve. However, "currently" is doing a lot of work in this sentence, and the gap between them and the rest of the world is narrowing with each release cycle. While Sol and Fable 5 may be the definitive versions when it comes to raw capabilities, Kimi K3 and DeepSeek V4 Flash are the definitive versions when it comes to intelligence per dollar, and the latter metric is ultimately what will drive investment decisions for most companies using AI at scale. There is no price war on the horizon. The price war is already upon us, and it is being fought by American companies.
[19]
OpenAI makes GPT-5.6 dramatically cheaper, putting pressure on Anthropic and rivals
The revised pricing undercuts Anthropic's Claude Sonnet 4.6 and comes amid growing pressure to lower AI deployment costs for businesses. OpenAI has announced a massive price cut for its GPT 5.6 family of AI models lowering API costs by as much as 80 pct for some offerings. This move is said to make the company's latest AI models more affordable for developers and businesses building AI-powered applications and also intensifies the competition with rivals such as Anthropic and Chinese AI firms. OpenAI CEO Sam Altman confirmed the changes in a post on X, saying the company aims to offer the best price/intelligence tradeoff across its AI portfolio. The decision comes as enterprises increasingly want lower AI operating costs amid rising demand for generative AI tools. The biggest price cut applies to GPT 5.6 Luna, the lightweight model. The input token pricing has dropped by 80 per cent to $0.20 per million tokens, while output tokens now cost $1.20 per million. GPT-5.6 Terra, the company's mid-tier model, has also become cheaper, with input token pricing reduced to $2 per million and output tokens priced at $12 per million. Also read: Apple Q3 earnings: Record iPhone sales, weaker forecast and Tim Cook's biggest takeaways But unlike the Luna and Terra, GPT 5.6 Sol has not received a price cut. Instead, OpenAI has introduced a new Fast API mode which promises up to 2.5 times quicker response times while maintaining the same reasoning capabilities as the standard version. The faster variant will cost twice as much as the regular Sol model, giving developers the option to prioritise lower latency for time-sensitive workloads. The revised pricing also narrows the gap with competitors. Anthropic's Claude Sonnet 4.6, for instance, currently charges $3 per million input tokens and $15 per million output tokens, making OpenAI's Terra model a more affordable option on paper. In related news, India's Sarvam AI has announced plans to build a one-trillion-parameter foundation model in India, launched its Epoch Builder Edition platform for India-focused AI development, announced lower AI model pricing than global rivals. The company also appointed former xAI researcher Devendra Singh Chaplot as an advisor to oversee AI research and multilingual model capabilities.
Share
Copy Link
OpenAI has dramatically reduced pricing for its GPT-5.6 Luna and Terra models by 80% and 20% respectively, responding to mounting pressure from Chinese competitors like DeepSeek and Alibaba. The move signals a fundamental shift in AI economics as companies prioritize affordability alongside capability.
OpenAI announced sweeping AI price cuts on July 30, slashing costs for its GPT-5.6 Luna model by 80% and GPT-5.6 Terra by 20%
3
4
. The aggressive pricing strategy represents a dramatic shift from March 2026, when GPT-5.4 launched at $2.50 per million input tokens and $15 per million output tokens1
. GPT-5.6 Luna now costs just $0.20 per million input tokens and $1.20 per million output tokens, while Terra drops to $2 and $12 respectively3
. This means frontier AI models remain cutting-edge for less than four months before pricing collapses1
. The flagship GPT-5.6 Sol maintains its $5 and $30 pricing, though a new Fast mode charges $10 and $60 per million input/output tokens for lower latency1
.
Source: VentureBeat
Chinese competitors have fundamentally altered the AI competition landscape. DeepSeek V4 Flash costs $0.435 per million input tokens and $0.87 per million output tokens, while Moonshot's Kimi K3 charges $3.00 and $15.00 respectively
1
. According to independent benchmarks by Artificial Analysis, DeepSeek V4 Flash 0731 performs within a single point of GPT-5.6 Luna while costing 40% less per task2
. Alibaba escalated competition further by releasing Qwen 3.8-Max, a 2.4 trillion-parameter model that matches capabilities from Anthropic and OpenAI2
. Hugging Face CEO Clément Delangue stated on CNBC that China is "clearly dominating on open models right now," predicting they could dominate at the frontier by year-end2
. These open-weight models from DeepSeek, Alibaba, Moonshot, MiniMax, and Z.ai provide enterprises their only credible alternatives to proprietary systems2
.
Source: The Register
Businesses have grown increasingly wary of AI spending as costs balloon without clear returns on investment
5
. Companies using AI heavily have complained about skyrocketing token costs for months, making cheaper alternatives particularly attractive1
. The shift from flat subscriptions to usage-based pricing leaves enterprises with unpredictable and often higher bills as usage per task becomes harder to estimate4
. OpenAI attributed the lower prices partly to efficiency gains from GPT-5.6, including improved code optimization during internal development4
. Analysts suggest these reductions will more likely accelerate enterprise AI deployments rather than reduce CIO budgets3
. Google introduced more-affordable Gemini 3.6 Flash and 3.5 Flash-Lite models, while Anthropic replaced its Opus 4.8 with a more capable Claude 5.0 at the same price point1
.Related Stories
OpenAI faces mounting financial pressures despite these competitive moves. The company is already losing money on subscription-based accounts and missed key revenue targets earlier this year after losing tens of billions in 2025
1
. OpenAI has committed to $600 billion in compute spend by 2030, including a $300 billion compute commitment with Oracle1
. Cutting prices on the most popular, affordable models suggests margins will either shrink dramatically or disappear altogether1
. Reports indicate Nvidia may backstop OpenAI with a $250 billion investment1
. Google spent approximately nine times its cloud revenue on AI infrastructure over the past year, while Anthropic only recently posted profits on annualized revenue through a limited cut-price deal with xAI to rent its Colossus data center1
. Analysts warn that cutting prices could boost usage but strain finances ahead of highly anticipated initial public offerings4
.The new pricing intensifies pressure on Anthropic, whose Claude Sonnet 4.6 model costs $3 per million input tokens and $15 per million output tokens, above Terra's rates
4
. Anthropic CEO Dario Amodei recently stated opposition to open models from China, ones distilled from proprietary models, and those not meeting rigorous safety metrics2
. However, neither safety concerns nor commitments from American and European tech giants change the reality that China provides the only meaningful competition in the open-weight models arena2
. Alibaba's Qwen 3.8-Max is now available via QwenCloud for $2 per million input tokens and $6 per million output tokens, with open weights released for the first time2
. Despite more workers using AI than ever before, productivity gains have been less than ideal, prompting major AI businesses to announce cuts and limits on technology use1
. Watch how OpenAI balances its massive compute costs commitments against shrinking margins, whether Anthropic adjusts pricing to remain competitive, and if Chinese developers continue expanding their open-weight model advantage.
Source: Tom's Hardware
Summarized by
Navi
[1]
26 Jul 2026•Policy and Regulation

24 May 2026•Technology

17 Jun 2026•Technology
