11 Sources
[1]
The Powerful Chinese Model Experts Warned About -- and Waited for -- Is Here
It's now even easier to find -- and exploit -- vulnerabilities in computer systems using AI. Last Friday, the Chinese AI company Z.ai announced a powerful open-weight model that it says is capable of automating cutting-edge coding and cybersecurity tasks almost as well as the best publicly available models from Anthropic and OpenAI. The new model, GLM 5.3, could be a gift for companies looking to secure their systems against attacks, providing a cheaper way to scan for hidden bugs and other weaknesses. Open-weight -- or free-to-download -- models can be run on one's own hardware and are often significantly less costly than closed models like Claude and GPT. Alongside the new model, Z.ai released OpenVuln, a service for scanning code repositories for vulnerabilities using GLM 5.3. For now, the new model is in a limited release with trusted partners, but it shows how quickly open-weight models are gaining superhuman hacking skills. And that might pose problems if the model is harnessed by criminals and other bad actors. That prospect is especially sobering following a string of startling incidents involving rogue AI agents with advanced cyber-skills. In recent weeks, OpenAI, Anthropic, and independent security researchers have revealed examples of agents escaping from testing environments and autonomously hacking into outside systems, including the research platform Hugging Face, to complete tasks. On Monday, OpenAI president Greg Brockman warned in a blog post that the Hugging Face incident would go down as "a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months." Brockman argued that AI models are becoming so good at scouring codebases for unknown flaws and analyzing systems for misconfigurations that it's crucial for organizations to use AI to scan their systems and identify issues before they can be exploited. OpenAI would, of course, like companies to use its AI to do that. So far, it's moving carefully in providing access to its most capable AI. Like Anthropic, OpenAI has made its most advanced models available to a limited number of partners prior to full release. The US government is also wrestling with the issue and now reviews frontier models as part of their releases. Some believe that open-source AI will be crucial to shoring systems up from attack; Nvidia recently announced an alliance to promote the use of open AI for cybersecurity. A previous version of Z.ai's GLM was used by Hugging Face to shore up its systems after an unreleased OpenAI model went rogue and broke them last month. In a post on X, Guillermo Rauch, CEO of Vercel, a web design and hosting company, said his engineers had tested GLM 5.3 as a tool for scanning sites for bugs. "Given its lower costs, I expect this to be a boon for defensive security work," Rauch wrote in his post. "It's the new open frontier." Z.ai said in a post announcing GLM 5.3 that it had improved the model by "post-training," which involves giving a model examples of solved problems and letting it learn through experimentation. The company cited coding and cybersecurity benchmark scores that show GLM 5.3 nearing or even exceeding the scores of Anthropic and OpenAI's models in some cases, like one popular cybersecurity benchmark called CyberGym. Z.ai also acknowledged the risk of releasing powerful open models in its post. "These capabilities can help defenders identify weaknesses earlier, validate risks, and accelerate remediation," the company wrote. "They also create clear dual-use risks. We are therefore taking a staged approach to release. Selected security partners will first evaluate GLM-5.3 in controlled settings." Z.ai says that full access to the model will be available in two weeks. "This model looks exceptional, with a somewhat astounding increase in scores," Nathan Lambert, a prominent AI expert, wrote in a post about GLM 5.3. "This is another step towards the inevitable proliferation of very strong cyber capabilities across the economy." Z.ai's latest release also highlights China's edge in open-weight models. Although the US has sought to restrict the country's access to the most advanced chips for training AI models, recent months have seen the release of several extremely powerful open-weight models, including Qwen 3.8 Max from Alibaba and Kimi 3 from Moonshot AI. Z.ai has previously said that it used Chinese-made chips from Huawei to train some of its models. Meta, which appeared to have abandoned open-source AI, now seems poised to lead the US challenge with a powerful model called Muse Spark. The US government is developing a framework designed to mitigate the impact of AI's advancing cyber capabilities. A big remaining question is what it should do with open models -- especially as they introduce more potential risk.
[2]
Zhipu says new coding AI developed advanced cyber skills faster than expected
China's AI developer claims GLM-5.3 rivals leading Western models in vulnerability discovery and has identified thousands of security flaws across real-world software. Chinese AI developer Zhipu has launched GLM-5.3, a new coding-focused AI model that the company says has developed unexpectedly strong cybersecurity capabilities, putting it close to global leading models in vulnerability discovery while remaining behind them on deeper exploitation tasks. Zhipu's own testing places GLM-5.3 slightly ahead of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol on CyberGym, a benchmark that tests vulnerability identification and validation. GLM-5.3 scored 84.5%, compared with 83.8% for Mythos 5 and 83.6% for GPT-5.6 Sol. But the model trails both competitors by a much wider margin on ExploitBench, where it scored 54.4%, compared with 78% for Mythos 5 and 76.5% for GPT-5.6 Sol. "GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench," Zhipu said in a statement. "As we scaled post-training, cyber capability developed faster than we expected." The company said GLM-5.3 moved beyond identifying isolated vulnerabilities to "forming coherent plans for complete exploitation chains."
[3]
China's Z.ai says new model nears Anthropic's Mythos 5 in cyber-defence tests
BEIJING, Aug 14 (Reuters) - Chinese AI startup Z.ai said on Friday its open-source GLM-5.3 model had neared Anthropic's restricted Mythos 5 in identifying software vulnerabilities, bolstering the credentials of a Chinese AI challenger gaining traction among Western developers. Z.ai said GLM-5.3 scored 84.5% on CyberGym, a test of whether a model can review code, identify security flaws and confirm that they are real. That was slightly higher than the 83.8% it reported for Mythos 5. The results have not been independently verified. GLM-5.3 lagged behind Mythos 5 in converting discovered flaws into working attacks -- a standard part of defensive security research. Z.ai said its model scored 54.4% on the ExploitBench test of this capability, versus 78.0% for Mythos 5. In a separate timed test, Z.ai said GLM-5.3 completed 105 attack-development tasks in two hours and 130 in six hours. Mythos 5 completed 181 and 247 tasks, respectively. Anthropic has made Mythos, a version of its Claude Fable 5 model with cybersecurity safeguards removed, available only to vetted organisations. Such controls reflect concern that AI systems capable of finding and exploiting software flaws can assist defenders but may also lower barriers for attackers. Z.ai said it would release GLM-5.3 publicly in about two weeks after completing security assessments and strengthening its safeguards. Its most sensitive cybersecurity functions would be available only to verified users through a "trusted access" programme, it said. The company said it had added several layers of protection to GLM-5.3, including systems to screen risky requests, monitor the model's work and train it to reject malicious tasks. It said these were designed to distinguish harmful activity from legitimate uses such as fixing bugs, teaching cybersecurity or authorised security testing. But critics say those safeguards become harder to enforce once a model is released for others to download, alter or combine with outside tools. CHALLENGING CLOSED-SOURCE MYTHOS Z.ai framed the launch as a challenge to the restricted-access model represented by Mythos. It argued that advanced cyber-defence tools should be available to developers of open-source software and smaller security teams, rather than being controlled by a limited number of closed-model providers. It said it would begin an "Open Source Shield" initiative to audit selected open-source projects, provide model access for defensive work and add code-auditing functions to its ZCode programming product. Z.ai is not the first Chinese company to position a product as an answer to Mythos. Cybersecurity firm 360 said in June that its Tulongfeng vulnerability-discovery system had achieved Mythos-equivalent capabilities by combining AI models with security data and automated tools, though those claims were not independently verified. GLM-5.3 differs in that it is a general-purpose coding model that Z.ai says acquired cybersecurity capabilities through expanded post-training and reinforcement learning, rather than a purpose-built security system. The company said it used the same base model as GLM-5.2 but trained it in longer and more varied task environments. The launch builds on global interest in GLM-5.2, which gained attention in recent months among overseas developers for coding and agent capabilities that users and analysts said approached leading U.S. models but at a much lower cost. Reporting by Eduardo Baptista. Editing by Mark Potter Our Standards: The Thomson Reuters Trust Principles., opens new tab * Suggested Topics: * Cybersecurity Eduardo Baptista Thomson Reuters Eduardo Baptista is Chief Technology Correspondent, Greater China, for Reuters, based in Beijing. He covers artificial intelligence, semiconductors and emerging technologies. He holds a BA in History from the University of Cambridge.
[4]
GLM-5.3 is here with advanced cyber capabilities -- and reportedly already found a 'serious vulnerability' in Cursor
Chinese AI startup Z.ai, known internationally for its growing lineup of powerful, largely open source GLM series of language models, today released GLM-5.3 with substantial gains in long-horizon coding and a more consequential -- and potentially sensitive -- jump in cybersecurity capabilities. Already, GLM-5.3's cyber capabilities have found a "potentially serious vulnerability in Cursor," the AI coding startup recently acquired by SpaceX, according to z.ai developer advocate Lou, posting on X. VentureBeat also tagged Cursor for confirmation on X and is awaiting response. GLM-5.3 is available initially only through the company's GLM Coding Plan and ZCode coding environment, while API access and open weights are coming later, "once safety evaluation and hardening are complete," according to the company. Z.ai says it plans to release weights approximately two weeks after launch. For enterprise developers, the notable part of the release is not simply another round of benchmark improvements. Z.ai says GLM-5.3 uses the same base model as GLM-5.2, with the improvements coming entirely from scaling post-training across more environments, more diverse tasks and additional reinforcement-learning compute. That makes GLM-5.3 something of a test of how far a frontier-scale base model can be pushed without another expensive pretraining cycle. "Scaling post-training is all we did for GLM-5.3," Z.ai wrote in its technical announcement. The results suggest considerable headroom. But they have also produced an unusual problem for an open-model developer: according to Z.ai, cybersecurity capabilities improved faster than anticipated as training scaled, particularly as tasks progressed from vulnerability identification toward constructing complete exploitation chains. Reuters reported Friday that Z.ai is also introducing controls around some of the model's more advanced capabilities, including a "trusted access" approach for sensitive functionality. A large jump in coding without another base model GLM-5.3 builds on the 743-billion-parameter-scale base model behind GLM-5.2 rather than replacing it. Z.ai instead expanded the post-training system it had already assembled around long-horizon reinforcement learning. Those environments increasingly resemble complete engineering jobs rather than isolated programming exercises. Z.ai describes scenarios in which an agent receives access to codebases, documentation, compute clusters, storage systems and experimental results, then has to diagnose problems, modify systems, run experiments and demonstrate a measurable improvement while preserving correctness. Some tasks are designed to approximate several days of work for an experienced engineer. The approach produced sizable generation-over-generation improvements on Z.ai's reported evaluations. GLM-5.3 jumps from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 26.2 to 48.2 on AutomationBench. On Agents' Last Exam CLI, it improves from 23.8 to 28.5. The model does not dominate every frontier competitor. Z.ai's own benchmark table shows GPT-5.6 Sol at 34.6 and Claude Fable 5 at 33.7 on Terminal-Bench 3.0, compared with GLM-5.3's 28.3. On DeepSWE v1.1, GLM-5.3 scores 66.9, compared with 72.7 for GPT-5.6 Sol and 69.7 for Fable 5. But Z.ai is also emphasizing efficiency rather than benchmark position alone. On its private Z.ai Code Bench, GLM-5.3 reaches a 34.5% result at its Max reasoning setting while consuming roughly 75,000 output tokens per task. GLM-5.2 reaches 23.4% while consuming approximately 96,000. At High effort, GLM-5.3 reaches 31.4% at roughly 50,000 output tokens, compared with Z.ai's reported 29.5% for Claude Opus 4.8 using 120,000. Because Code Bench is Z.ai's own private evaluation, those comparisons should be treated as company-reported results rather than independent measurements. Still, reducing token consumption while improving task completion is operationally important for enterprises deploying coding agents, where long-running loops can make inference cost and latency compound quickly. Cyber capabilities developed faster than Z.ai expected The more unusual development is cybersecurity. Z.ai introduced vulnerability-discovery environments into GLM-5.3's post-training mix expecting the model to improve at finding software flaws. Instead, the company says capability began progressing further along the exploitation chain. "As we scaled post-training, cyber capability developed faster than we expected," Z.ai wrote. On CyberGym, which tests vulnerability discovery and validation against source code, GLM-5.3 scores 84.5%, compared with 77.2% for GLM-5.2. That also edges Z.ai's reported scores for GPT-5.6 Sol at 83.6% and Mythos 5 at 83.8%. The advantage does not extend across the entire exploitation stack. GLM-5.3 scores 54.4% on ExploitBench, more than twice GLM-5.2's 24.4%, but remains well behind the 76.5% Z.ai reports for GPT-5.6 Sol and 78% for Mythos 5. Similarly, on ExploitGym, GLM-5.3 completes 105 tasks under a normalized two-hour budget and 130 under six hours, up from 29 and 39 for GLM-5.2. Fable 5 reaches 181 and 247, while GPT-5.6 Sol reaches 216 and 293. The direction of travel may matter more than the leaderboard position. Z.ai says work with security teams in China has resulted in 2,436 vulnerability findings across 269 projects after expert review, screening and deduplication. Its disclosure ledger lists 1,097 as critical or high severity, with 53 publicly disclosed and 2,383 still under embargo at the time of the release. That creates a tension increasingly facing frontier model providers: the same long-horizon agent capabilities that make models more useful for software engineering can also make them more capable security researchers -- and potentially more capable offensive operators. GLM-5.3 also requires developers to change how they call the model Developers migrating existing GLM applications should pay attention to a breaking API behavior. GLM-5.3 supports three reasoning-effort levels -- , and -- with the default and Z.ai's recommended setting for coding. But unlike previous releases, thinking cannot be disabled. Applications currently sending must change the value to and specify a reasoning effort before switching the model identifier to GLM-5.3. Otherwise, Z.ai says the request will fail. That makes GLM-5.3 an actual migration rather than simply a model-name substitution for some production applications. From GLM-4.5 to GLM-5.3: Z.ai's rapid push into agentic engineering GLM-5.3 is the latest step in a rapid shift by Z.ai -- formerly known as Zhipu AI -- toward coding agents and long-running autonomous engineering workloads. GLM-4.5, released in July 2025, established much of that direction. The 355-billion-parameter mixture-of-experts model was designed to combine reasoning, coding and agent capabilities, while the smaller GLM-4.5-Air offered 106 billion total parameters. Z.ai released the models with open weights and emphasized integration with agent frameworks. GLM-4.6 followed in September, expanding context from 128,000 to 200,000 tokens and targeting coding, tool use and agent workflows in environments including Claude Code, Cline, Roo Code and Kilo Code. Z.ai also began placing greater emphasis on token efficiency in real-world coding evaluations rather than benchmark performance alone. The larger architectural jump came with GLM-5 in February 2026. Z.ai scaled the model from GLM-4.5's 355 billion parameters to 744 billion, with 40 billion active parameters, and increased pretraining data to 28.5 trillion tokens. It also introduced its "slime" asynchronous reinforcement-learning infrastructure and explicitly repositioned the GLM family around "agentic engineering" and long-horizon tasks. By June, GLM-5.2 had turned that strategy into a more direct enterprise proposition. The 753-billion-parameter model arrived with a stable 1-million-token context window, open weights under an MIT license and support across more than 20 coding environments. It also introduced IndexShare, which reuses an indexer across sparse-attention layers to reduce the computational burden of very long contexts. GLM-5.2 was priced at $1.40 per million API input tokens and $4.40 per million output tokens, with cached input priced substantially lower, positioning Z.ai as both a technical and pricing competitor to proprietary frontier labs. Z.ai's ambitions have been expanding outside model development as well. Reuters reported last month that Zhipu AI raised roughly HK$31.4 billion, or about $4 billion, through a Hong Kong share sale, with proceeds intended for areas including research and development, computing infrastructure, talent and business expansion. Taken together, the releases show a consistent progression: GLM-4.5 unified reasoning, coding and agents; GLM-5 substantially scaled the foundation model; GLM-5.2 attacked long-context and long-horizon engineering; and GLM-5.3 is now attempting to extract substantially more capability from that same foundation through post-training. Pricing, ZCode and availability GLM-5.3 is available now through Z.ai's GLM Coding Plan and ZCode. ZCode is the company's own coding-agent environment and supports long-running "Goal" tasks that plan, implement, test and verify work. It also offers remote control of running tasks and is available on macOS, Windows and Linux. Individual GLM Coding Plans currently start at a listed promotional price of $12.60 per month for Lite with 10,000 credits per week. Pro is listed at $56 per month with six times Lite usage, while Max costs $117.60 per month with 14 times Lite usage. Team Standard and Premium seats are listed at $88 and $188 per user per month, respectively. Z.ai has also moved the Coding Plan to a points-based quota system that separately accounts for input, cached-input and output tokens. Calls outside the company's weekday peak period consume 50% of the normal points. The company has not yet provided general GLM-5.3 API pricing in the supplied launch materials, making total production API cost difficult to compare directly with GLM-5.2 or competing frontier models until staged API access arrives. That staged release may ultimately be the most important part of GLM-5.3. Z.ai spent the past year pushing an open-model strategy centered on permissive weights, low-cost inference and compatibility with existing coding-agent ecosystems. GLM-5.3 demonstrates what happens when that strategy succeeds perhaps too well in one sensitive domain: better autonomous engineering also means better autonomous security research. The result is a model that advances Z.ai's coding ambitions while forcing the company to confront the same capability-versus-access tradeoff facing the largest closed frontier labs. For enterprise developers, GLM-5.3 is therefore worth watching for two reasons. Its coding results provide another indication that increasingly capable agents can emerge from better post-training and environments without continuously rebuilding the underlying foundation model. Its cybersecurity results show why deciding how those agents are distributed may become just as important as deciding how they are trained.
[5]
China's latest open-weight model rivals U.S. models at hacking
Why it matters: Cyber-capable AI models can bring huge gains to both defenders and hackers eager to scale up their operations. Driving the news: China-based AI lab Z.ai warned Friday that its latest model, GLM-5.3, is so capable at finding and exploiting security flaws that the company will delay the public release of the model weights for two weeks as it tests and strengthens safety and security controls. * The company specifically trained GLM-5.3 to get better at finding vulnerabilities by letting it practice cyber tasks in controlled environments. * Z.ai is implementing a tiered access program that only gives selected security partners access to GLM-5.3 in controlled environments. By the numbers: GLM-5.3 scored 84.5% on CyberGym, a benchmark that tests how well models can find known security vulnerabilities -- beating out Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol. * On ExploitBench, which tests models' ability to reason through and develop exploits for real vulnerabilities, GLM-5.3 trailed only Fable 5 and GPT-5.6 Sol, among the models Z.ai tested. The big picture: As some U.S. labs slow down model releases over security risks, Chinese labs are forging ahead and fine-tuning their open-weight models to be better at cyber tasks. * Meanwhile, hackers are experimenting more with AI models. Earlier this week, researchers at Dream found that open-source AI agents were used in an automated cyberattack against Taiwan's government. Zoom in: Z.ai claims its GLM models have found more than 2,400 security flaws, including over 1,000 that are considered critical and high, according to a new disclosure site released Friday. * Some of those vulnerabilities were found in the Linux kernel and across widely used VMware and Apache projects. Between the lines: Z.ai is framing the release of GLM-5.3 as a boon for defenders. * "An open world cannot have only open attack surfaces," the company wrote on X. "It must also have an open shield." * Z.ai also announced a new program Friday where open-source maintainers can have a GLM model scan their open repositories for bugs. The intrigue: At least one of the company's models has already proven useful to defenders: Hugging Face said it used GLM-5.2 to investigate a recent breach involving OpenAI's models after guardrails on U.S. frontier models declined to help. Yes, but: Once GLM-5.3's weights are public, Z.ai acknowledged in its X post that it won't be able to control how people modify or use the model. What we're watching: The Trump administration appears to be considering ways it can regulate open-source models as their capabilities start to rival their closed counterparts.
[6]
China's Z.AI Ships GLM-5.3, Calling It the Top Open-Weight Coding Model
GLM-5.3 is a 743-billion-parameter model built by scaling post-training on the GLM-5.2 base. Chinese AI lab Z.ai released GLM-5.3 on Thursday, a sizable coding model it's pitching as the strongest open-weights coder on the market. The model is live now through the GLM Coding Plan subscription and ZCode, with API access and downloadable weights following after a safety review. "Scaling post-training is all we did for GLM-5.3," the company wrote in its launch post. "With GLM-5.2 we built the stack... Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them." The team focused more on token efficiency, not raw dominance. GLM-5.3 stands at 743 billion parameters and consumes a lot less tokens per task than its predecessor. Parameters are the amount of dials a model handle while processing information while tokens are the basic unit of information a model can consume or generate. Z.ai says GLM-5.3 clears 34.5% on its in-house Z.ai Code Bench at Max effort while burning roughly 75,000 output tokens per task, against GLM-5.2's 23.4% at 96,000. Against closed models, the blog notes it beats Claude Opus 4.8 on token economy but "remains behind Claude Fable 5, which reaches 39.5% at Max effort." In terms of coding, GLM5.3 is a very good performer, beating fellow Chinese model Kimi K3 on the most relevant benchmarks. On Terminal Bench 3.0 -- a test of autonomous shell/tool use in real Linux environments -- GLM-5.3 scores 28.3, slightly behind closed models Fable 5 (33.7) and GPT-5.6 Sol (34.6). On DeepSWE v1.1, a benchmark for fixing real GitHub issues end-to-end, open rival Kimi K3 (67.5) and Fable 5 (69.7) both beat GLM-5.3's 66.9. The pattern can be more or less summed up like this: GLM-5.3 clears its own predecessor and some open peers, but closed U.S. models still lead the headline coding boards. The cybersecurity results show another important leap. GLM-5.3 leads CyberGym at 84.5% and more than doubles GLM-5.2 on exploitation benchmarks. Z.ai says the model flagged 2,436 vulnerabilities across 269 open-source projects, 1,097 of them medium-to-high severity. "GLM-5.3 takes agentic coding to the next level, delivering a dramatic improvement over GLM-5.2 while achieving better results with fewer output tokens," Z.ai posted on X. "GLM-5.3 is available now through GLM Coding Plan and ZCode. API access and open weights will be released in stages following rigorous safety evaluations." On price, the gap with U.S. frontier models is the open-weights draw. Z.ai's GLM Coding Plan runs on a points quota (off-peak calls cost half), with Zhipu's API priced at roughly a tenth of U.S. frontier per-token rates -- GLM-5.2's official rate was $1.40 in / $4.40 out per million tokens. That stacks against GPT-5.3-Codex at $1.75 / $14 and Claude Opus 4.8 near the top of Anthropic's tiers. Z.ai is a Beijing lab included on the U.S. Entity List, which means American firms cannot export controlled tech to it. Despite this, GLM is an extremely popular model and Chinese open-weight models already beat American ones on OpenRouter token usage. GLM-5.3 weights are set for public release in about two weeks, per the launch post -- the open-weights label applies to what's coming, not what's downloadable today.
[7]
Z.ai debuts GLM-5.3 with long-horizon coding, cybersecurity upgrades
Chinese artificial intelligence developer Z.ai Co. today debuted GLM-5.3, an open-source large language model that set records across several popular benchmarks. The LLM is based on an algorithm called GLM-5.2 that the company released in mid-July. The latter model features a mixture of experts architecture with 753 billion parameters and a context window of 1 million tokens. GLM-5.3 has an identical design, but went through a more extensive post-training process. Z.ai says that its training optimizations delivered significant performance improvements. GLM-5.3 achieved the highest score of any open-source AI model on Terminal Bench 3.0, which measures LLMs' command line scripting capabilities. It performed 50% better than GLM-5.2 on an internal Z.ai benchmark for evaluating coding agents. Notably, the model is also highly adept at cybersecurity research. It outperformed Claude Mythos 5 on CyberGym, a benchmark that evaluates LLMs' ability to find code vulnerabilities. GLM-5.3 fell behind Anthropic's flagship LLM on two other cybersecurity benchmarks. Z.ai says that the model has so far found more than 2,400 vulnerabilities in 269 software projects. About half of the flaws have a severity rating of medium or higher. According to the company, one of the vulnerabilities that GLM-5.3 found is in a piece of code authored 40 years ago. The post-training process through which Z.ai refined the model's coding capabilities involved sandboxes designed to mimic developer workstations. The company installed GLM-5.3 in the sandboxes and instructed it to complete complex coding tasks. Some of the exercises took days to complete, which improved the model's ability to tackle long-horizon tasks. Z.ai generated the sandboxes using specialized AI agents. According to the company, the agents modeled the environments they generated on real-world software projects. They then created programming exercises customized to each sandbox. A separate "judge agent" verified that the challenges can be solved before they were given to GLM-5.3. According to Z.ai, its engineers also automated certain other aspects of the training workflow. The company built pipelines capable of generating a reward signal, a piece of data that guides the LLM learning process. It provides feedback that helps the model being trained identify ways of refining its output. Z.ai built its training stack on two open-source technologies called slime and SAO. The former tool makes it easier to move an LLM from its training environment to production inference infrastructure. SAO, in turn, is an implementation of an AI method called asynchronous reinforcement learning that speeds up training runs. GLM-5.3 is currently available through Z.ai's GLM Coding Plan subscription service. The company plans to release the model's weights on Hugging Face under an open-source license within two weeks.
[8]
China's Z.ai unveils GLM-5.3, claims chart-leading scores
GLM-5.3 scores 84.5pc on CyberGym, ahead of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. Chinese AI company Z.ai has unveiled its latest open-weights model - GLM-5.3 - whose performance, it says, rivals that of OpenAI and Anthropic newest frontier launches - marking the latest in the ever-escalating battle between US and Chinese companies for AI dominance. The new model makes a significant leap in capability compared to its predecessor GLM-5.2, which launched just months ago in June. GLM-5.3 is "much better" at complex coding and long-horizon tasks, Z.ai said, boasting a 50pc improvement from the older version in its in-house code benchmarks. The company plans to release GLM-5.3's weights in two weeks' time, after wrapping up safety evaluation and hardening. "[GLM-5.3] delivers markedly stronger agentic coding results than GLM-5.2 at every effort level while consuming fewer output tokens," Z.ai said. GLM-5.3 also proves significantly more capable at exploiting vulnerabilities, scoring 84.5pc on CyberGym, above Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. "We expected this to make the model better at finding and reasoning about vulnerabilities. What surprised us was how quickly the capability continued to develop as training scaled," Z.ai said. Despite its lofty claims about the new model, the Hong Kong-listed company's shares slid nearly 4pc at market close today (14 August). "This firm remains on a completely unsustainable commercial footing," intelligence analyst Robert Lea told Bloomberg. "Rising agentic AI will drive Z.ai's inference costs and losses higher." Chinese-made models are increasingly offering models at par with its American counterparts at cheaper prices, in spite of the heavy constraints around chip exports into the country. Last month, Z.ai completed the construction of a data centre where has installed at least 10,000 Chinese-made chips to help develop its GLM line of models. Other frontier models from China, including Moonshot's latest Kimi model, DeepSeek's V4 models and Alibaba's Qwen series, have all ranked high in benchmarks. The Beijing-based start-up went public in January this year, raising funds at a market capitalisation of nearly $7bn. Shares at the company surged 2,000pc in June following the launch of GLM-5.2, taking Z.ai's valuation to $128bn. Though that has since fallen to $75bn. Don't miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic's digest of need-to-know sci-tech news.
[9]
Zhipu's new AI model takes aim at OpenAI, Anthropic in China's race to close the gap
Chinese AI firm Z.ai has launched its new GLM-5.3 model. This release intensifies competition with OpenAI and Anthropic. The model offers improved coding capabilities and open access for developers. Chinese AI companies are narrowing the performance gap with US rivals. Z.ai aims to attract users with its permissive licensing approach. Chinese artificial intelligence company Z.ai, also known as Zhipu, has released the latest version of its flagship model, stepping up competition with US leaders OpenAI and Anthropic with a focus on stronger coding capabilities and open access. The Beijing-based company's new GLM-5.3 builds on the same roughly 700-billion-parameter base model as GLM-5.2, which was released in June, but is designed to deliver a significant improvement in coding and other capabilities. Also Read: China's Zhipu posts 132% rise in annual revenue on AI boom Z.ai plans to release the model's weights, the underlying parameters that allow developers to download, modify and customise an AI system, within two weeks. The company said benchmark results showed GLM-5.3 substantially outperforming its already well-regarded predecessor and approaching, and in some tests approaching, Anthropic's Fable 5. The release comes as Chinese AI companies accelerate efforts to challenge the dominance of US developers in increasingly capable large language models. San Francisco-based Anthropic and OpenAI remain widely regarded as having some of the industry's most capable models. But Chinese companies have increasingly narrowed the performance gap, while offering models at lower costs and, in several cases, making their weights openly available to developers. Moonshot AI's Kimi, DeepSeek's latest V4 models and Alibaba Group Holding's Qwen have all scored highly in AI benchmarks. Z.ai is seeking to build on that momentum with a permissive licence intended to attract developers and users to its models. China's AI challengers gain groundThe competitive landscape is also shifting on price. DeepSeek raised prices for its V4 Flash and Pro models by as much as fourfold on Thursday. Even after the increase, the blended price of V4 Pro remains below that of GLM-5.2, according to Artificial Analysis. The third-party evaluator gives DeepSeek's top model and GLM-5.2 an identical intelligence score of 53. Both trail Kimi K3 as well as US frontier models including Claude Fable 5 and OpenAI's GPT 5.6. Also Read: After Anthropic shutdown, China's Z. ai closes frontier gap as it plans dual listing For Z.ai, the improving performance of Chinese models comes alongside a rapidly expanding commercial footprint. The company is among China's AI pioneers and became the world's first large language model maker to go public. Its market value surged to $137 billion at its peak following the release of GLM-5.2, briefly overtaking Chinese internet companies including PDD Holdings and NetEase. The valuation has since fallen to about $80 billion, although that still represents roughly a tenfold increase from its January listing in Hong Kong. Z.ai has also been building out the computing infrastructure needed to support its models. The company completed a data centre with at least 10,000 Chinese-made chips for developing and deploying its GLM models, Bloomberg reported in July. Its annual recurring revenue, predictable yearly revenue generated from developers and businesses using its services, reached $1 billion in July, according to the report. The latest GLM release adds another contender to a rapidly evolving AI race in which Chinese developers are increasingly competing not only on model performance, but also on cost, accessibility and the ability to let developers build directly on their technology.
[10]
Zhipu's GLM-5.3 Matches Fable 5 On Coding Using Only Post-Training, And Stuns Fans By Unearthing A Vulnerability All The Way From 1981
Another day brings yet another impressive model from a China-based AI lab, with Zhipu's GLM-5.3 taking the center stage today after Meta's Muse Glimmer, NVIDIA's Nemotron 3.5 Lightning, DeepSeek's V4 Pro, and Google's Gemini 3.7 Flash dominated the airwaves earlier in the week. Zhipu's GLM-5.3 is primarily focused on coding and security, with impressive gains in relevant benchmarks, such as a ~1.5x gain on DeepSWE, a 6x gain on Terminal Bench 3.0, and an apex predator status on CyberGym The GLM-5.3 is distinctive in that it sports the same 743 billion parameters of its predecessor, with all of the gains in coding and cyber security coming from Zhipu's post-training post-training protocols. As such, GLM-5.3 has a score of 66.9 on DeepSWE vs. a score of 46.2 for GLM 5.2, and 69.7 for Anthropic's Mythos-class Fable 5. On Terminal Bench 3.0, GLM-5.3 achieves a score of 28.3 vs. 4.6 for GLM-5.2 and 33.7 for Fable 5. On CyberGym, GLM-5.3 is currently leading with a score of 84.5 vs. 83.8 for Fable 5 and 83.6 for GPT-5.6 Sol. Remember, all of these gains are solely the result of Zhipu's post-training process, which is an eminently impressive feat in and of itself. Zhipu has also published a ledger, disclosing 2,436 security findings across 269 open-source projects, of which 1,097 were Critical or High severity. Interestingly, the oldest vulnerability that GLM-5.3 found dated all the way back to 1981! What's more, these security exploits had evaded discovery for an average of 26.6 years. Finally, do note that Zhipu's next model will switch to a new architecture with 2x as many parameters, which suggests that the AI lab is now seriously gunning after Fable 5's lunch. The company will release the model weights for GLM-5.3 in a few days. Follow Wccftech on Google to get more of our news coverage in your feeds.
[11]
Z.ai Delays GLM-5.3 AI Model Release After Unexpected Hacking Skills Emerge
Z.ai paused its GLM-5.3 AI model after tests showed it could handle advanced hacking tasks. The move adds to growing concerns about releasing powerful AI systems in public. China's Z.ai put its powerful GLM-5.3 AI model on hold. The decision came after tests showed that the model could handle difficult hacking tasks. The company announced that GLM 5.3 is highly capable of finding and exploiting flaws. Thus, the company is holding the public release by two weeks to strengthen security. The company trained the model specifically to identify vulnerabilities, but it's also exploiting them. For now, the company has granted limited partners access to the model in controlled settings. Z.ai is not the only company dealing with this problem. OpenAI and Anthropic have also seen their AI models show strong cyber skills during tests. These models can identify software flaws and assist with certain hacking tasks. The growing abilities have made AI safety a bigger concern. Companies now have to think about what their models can do after release, not just what they were built to do. AI can help security teams find weak spots before hackers do. The same tools can also be used for the wrong reasons. A model that can write code, find flaws, and complete tasks on its own can be risky in the wrong hands. Z.ai's decision shows why testing matters. A model may seem safe at first. More tests can reveal abilities that were not expected. The problem starts when even the developers are not sure how far a model's abilities can go. In such cases, releasing the weights could create risks that are hard to reverse. GLM-5.3 adds another example to an AI safety debate that is only getting bigger.
Share
Copy Link
Chinese AI startup Z.ai unveiled GLM-5.3, an open-weight AI model with unexpectedly strong cybersecurity capabilities that rivals Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol in vulnerability discovery. The model scored 84.5% on CyberGym but trails competitors on exploit development. Z.ai is delaying public release for two weeks to strengthen safeguards amid concerns about malicious actors exploiting advanced cyber capabilities.
Chinese AI startup Z.ai announced GLM-5.3 on Friday, an open-weight AI model that developed advanced cyber capabilities faster than the company anticipated.
1
The coding AI achieved a score of 84.5% on CyberGym, a benchmark testing vulnerability discovery and validation, slightly surpassing Z.ai's reported scores for Anthropic's Mythos 5 at 83.8% and OpenAI's GPT-5.6 Sol at 83.6%.3
The model's capabilities emerged through post-training scaling across diverse environments, with Z.ai stating that "cyber capability developed faster than we expected" as the system progressed from identifying isolated vulnerabilities to forming complete exploitation chains.2

Source: Wccftech
GLM-5.3 uses the same 743-billion-parameter base model as GLM-5.2, with improvements coming entirely from expanded reinforcement learning and post-training rather than expensive pretraining cycles.
4
The model demonstrated substantial coding improvements, jumping from 4.6 to 28.3 on Terminal-Bench 3.0 and from 46.2 to 66.9 on DeepSWE v1.1.4
Z.ai claims its GLM models have identified more than 2,400 security flaws across real-world software, including over 1,000 critical and high-severity vulnerabilities in systems like the Linux kernel and widely used VMware and Apache projects.5
While GLM-5.3 excels at vulnerability discovery, it trails significantly behind leading models on ExploitBench, which tests AI-driven cybersecurity tools' ability to convert discovered flaws into working attacks.
3
The Chinese AI model scored 54.4% on this benchmark, compared with 78% for Mythos 5 and 76.5% for GPT-5.6 Sol.2
In timed testing, GLM-5.3 completed 105 attack-development tasks in two hours and 130 in six hours, while Mythos 5 completed 181 and 247 tasks respectively.3

Source: SiliconANGLE
The performance difference highlights the dual-use nature of AI cybersecurity capabilities. Anthropic has restricted Mythos 5 access to vetted organizations precisely because systems capable of finding and exploiting software flaws can assist defenders but may also lower barriers for malicious actors.
3
OpenAI president Greg Brockman warned this week that recent incidents of AI agents autonomously hacking systems like Hugging Face represent "a watershed moment for cybersecurity" that previews how typical threat actors will evolve in coming months.1
Z.ai is implementing a staged release for GLM-5.3, delaying public availability of the open-weight AI model for approximately two weeks while conducting safety evaluations and strengthening safeguards.
5
The company introduced a trusted access program that initially limits the model's most sensitive cybersecurity functions to verified users and selected security partners in controlled settings.1
Z.ai added multiple protection layers including systems to screen risky requests, monitor the model's operations, and train it to reject malicious tasks while distinguishing harmful activity from legitimate uses like bug fixing, cybersecurity education, and authorized security testing.3
Critics note these safeguards become harder to enforce once weights are publicly released, as users can download, modify, or combine the model with external tools.
3
Z.ai acknowledged this limitation, stating it won't be able to control how people use the model after public release.5
The company is launching an "Open Source Shield" initiative to audit selected open-source projects, provide model access for cyber-defense work, and integrate code-auditing functions into its ZCode programming product.3
Related Stories
GLM-5.3's release underscores China's growing advantage in open-weight AI models despite US restrictions on access to advanced chips for training AI systems.
1
Recent months have seen powerful Chinese AI model releases including Qwen 3.8 Max from Alibaba and Kimi 3 from Moonshot AI, with Z.ai previously stating it used Chinese-made chips from Huawei to train some models.1
The company's earlier GLM-5.2 gained traction among Western developers for coding capabilities approaching leading US models but at significantly lower costs.3

Source: Axios
Vercel CEO Guillermo Rauch reported his engineers tested GLM-5.3 for scanning sites for bugs, stating "Given its lower costs, I expect this to be a boon for defensive security work. It's the new open frontier."
1
Hugging Face used GLM-5.2 to investigate a recent breach after guardrails on US frontier models declined to assist.5
AI expert Nathan Lambert noted the model "looks exceptional, with a somewhat astounding increase in scores" and represents "another step towards the inevitable proliferation of very strong cyber capabilities across the economy."1
The Trump administration is reportedly considering regulatory frameworks for open-source models as their capabilities rival closed counterparts, while the US government now reviews frontier models as part of release processes.1
5
Summarized by
Navi
[4]
24 Jun 2026•Technology

25 Jun 2026•Technology

28 Jul 2026•Technology

1
Technology

2
Policy and Regulation

3
Business and Economy
