3 Sources
[1]
Z.ai pitches GLM-5.2 for long-running software engineering tasks
The open-source model combines a one million-token context window with architectural updates aimed at lowering the cost of repository-scale AI coding. Z.ai has released GLM-5.2, an MIT-licensed open-source AI model designed for long-running software engineering tasks, as the Chinese company seeks
[2]
Z.ai's open-weights GLM-5.2 beats GPT-5.5 on multiple long-horizon coding benchmarks for 1/6th the cost
Today, Chinese AI startup Z.ai (formerly Zhipu AI) announced the immediate release of GLM-5.2, a 753-billion parameter open-weights large language model (LLM) engineered specifically to dominate "long-horizon" autonomous coding and engineering tasks. Available immediately on Hugging Face, the Z.ai
[3]
China's Z.AI Releases GLM-5.2: A Model That Rivals Claude Opus -- Using Zero Nvidia Chips
Unsloth AI already released 2-bit GGUF quantizations that shrink the model from 1.51TB to 238GB. You'll still need 256GB of RAM or VRAM -- but at that point, you can run it. Z.ai dropped GLM-5.2 on June 16, promising top level performances, beating its already advanced GLM 5.1. The Beijing-based
Share
Copy Link
Chinese AI startup Z.ai launched GLM-5.2, a 753-billion parameter open-source AI model that outperforms GPT-5.5 on multiple coding benchmarks while offering significant cost advantages. Released under an MIT license, the model features a one million-token context window and was trained entirely on Huawei Ascend chips, positioning it as a competitive alternative to proprietary coding models from OpenAI and Anthropic.
Z.ai has released GLM-5.2, an open-source AI model engineered specifically for long-running software engineering tasks that challenges the dominance of proprietary coding models from American tech giants
1
. The 753-billion parameter open-weights large language model became immediately available on Hugging Face, the Z.ai API, and more than 20 third-party coding environments, with enterprise subscription tiers starting at just $12.60 per month2
. Released under an unrestricted MIT license, the model allows enterprises to download, customize, fine-tune, and potentially run it locally for only the cost of compute and electricity, offering an appealing path to bypass geographic fencing and commercial limitations2
.
Source: VentureBeat
GLM-5.2 delivers impressive results on industry-standard benchmarks, particularly excelling in agentic coding workflows and long-horizon autonomous coding tasks. On the FrontierSWE benchmark, which evaluates open-ended technical projects measured in hours, the model scored 74.4%, trailing Claude Opus 4.8 by just 1% at 75.1% while surpassing GPT-5.5's 72.6%
1
3
. The model demonstrated even stronger performance on SWE-bench Pro, scoring 62.1 and decisively beating GPT-5.5's 58.6, while also clearing its predecessor GLM-5.1's 58.4 by a significant margin2
3
. On MCP-Atlas tool-usage evaluation, GLM-5.2 achieved 77.0, outscoring GPT-5.5's 75.3 and performing just shy of Claude Opus 4.8's 77.82
. The quality jump makes it the best open-source model to date in the Artificial Analysis Intelligence Index, which aggregates results from nine different scores3
.
Source: Decrypt
The model supports a one million-token context window with up to 131,072 output tokens, positioning it for repository-scale AI coding that requires reasoning across large codebases
1
. This represents a fivefold increase over GLM-5.1's 200K limit3
. Under the hood, GLM-5.2 operates as a mixture-of-experts model and introduces a major architectural optimization called IndexShare, which reuses the identical indexer across every four sparse attention layers2
. At maximum one million-token context length, this innovation reduces per-token compute FLOPs by 2.9 times2
. The model also features an upgraded Multi-Token Prediction layer for speculative decoding, boosting accepted token length by up to 20% during inference, and flexible selectable Thinking Modes that allow users to toggle between "Max" for peak logical problem-solving and "High" for balanced performance with token efficiency2
.Related Stories
GLM-5.2 was trained entirely on Huawei Ascend chips with no Nvidia hardware in the pipeline
3
. Emad Mostaque, founder of Stability AI, estimated total training costs at around $25 million, with 80% spent on post-training, making it extremely cheap compared to its peers3
. API pricing runs $1.40 per million input tokens and $4.40 per million output, delivering substantial cost efficiency compared to Claude Opus 4.8's $5 input and $25 output pricing3
. For enterprises seeking to run the model locally, Unsloth AI released 2-bit GGUF quantizations that compress the model from 1.51TB down to 238GB while retaining approximately 82% accuracy, though this still requires 256GB of unified memory or matching RAM/VRAM combination3
.
Source: InfoWorld
The Beijing-based lab, which has been on the U.S. Entity List since January 2025, appears to be benefiting from growing concerns over America's approach to AI
3
. Following the Trump Administration's export control directive prohibiting foreign nationals from using Anthropic's Claude Fable 5 model, which led Anthropic to take the models entirely offline for all users, Z.ai's offering provides enterprises a highly capable path to host frontier-level AI locally2
. Over the past week, the ban on Anthropic Fable and the release of this new model helped drive Z.ai's stock up 90%, sending it to a new all-time high3
. For multi-shot generation workflows and agentic pipelines where output diversity matters more than polish, the economics at open-source pricing levels present a compelling value proposition, though gaps remain on the hardest sustained tasks like SWE-Marathon, where GLM-5.2 scores 13.0 against Claude Opus 4.8's 26.03
.Summarized by
Navi
[2]
29 Jul 2025•Technology

11 Feb 2026•Technology

25 Jun 2026•Technology

1
Technology

2
Policy and Regulation

3
Technology
