4 Sources
[1]
ByteDance trains massive AI model in bid to rival Anthropic
ByteDance is training an AI model that could approach the size of Anthropic's most cutting-edge Mythos system, as Chinese companies continue to narrow the gap with the top US labs. The Chinese tech giant is at an early stage of training a model with as many as 10 trillion parameters -- three times larger than Moonshot's Kimi K3, the biggest Chinese model released to date, according to three people with knowledge of the matter. The ByteDance model is being pre-trained -- a stage that typically takes three to six months -- before it is fine-tuned and released if all goes well, one of the people said. The exact model size would only be determined at a later stage. Anthropic doesn't disclose the size of its models, but industry estimates say its most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion. While parameter count sets the fundamental capacity or memory limits for the models to store information, actual capability also depends on other factors such as data quality and training methods. ByteDance's efforts to train one of the world's largest AI models show Chinese labs' ambition to not only catch up but outperform their US peers in the most advanced level of AI. In the past weeks alone, Chinese models from Moonshot and Alibaba show strong performance on benchmarks, lagging behind only Anthropic's Fable 5 in certain areas. Mythos 5, Anthropic's most advanced model, is only available to approved organisations after a temporary ban in June due to security concerns. Industry insiders say multiple Chinese labs are in the process of training models of the size of Fable 5, while ByteDance is currently the most ambitious in pushing for the largest. ByteDance, the parent of viral video platform TikTok, has kept a low profile in its AI development as its models are mostly closed, unlike many of its Chinese peers. Its latest SeeDance model ranks among the most advanced globally in video generation, while its flagship consumer-facing model Doubao is the most popular in China with 324 million monthly active users. Over the past three years, ByteDance has invested in AI more aggressively than any of the other Chinese tech giants, building out its network of data centers and hiring researchers. It has doubled down on its cloud unit, Volcano Engine, which sells AI solutions to enterprises. ByteDance also has ambitions to develop custom AI chips. Seed, its model development team led by former Google DeepMind scientist Wu Yonghui, has about 2,000 members in China and overseas. The team includes core researchers, infrastructure engineers, a data labelling team and translators. Its model development has also implemented a more independent approach that does not involve "distilling" existing models from other labs, according to one of the people. This approach has been in place for more than a year, which some believe has led to its slower development versus rivals. Model distillation is the process of compressing a large, complex AI model into a smaller, faster one by training the smaller model to copy the knowledge and outputs of the larger teacher. ByteDance's management, led by founder Zhang Yiming, believes only independent development can result in a model that outperforms rivals. Zhang reiterated his stance in an internal meeting two weeks ago, where he told the Seed team to target "world-leading model capabilities" in the long run without getting too worried about falling behind in the near term, according to the person. Chinese media Latepost and The Information first reported Zhang's comments from the recent internal meeting. ByteDance did not respond to a request for comment.
[2]
ByteDance targets mega AI model nearing Anthropic's Mythos
ByteDance is training an AI model that could approach the size of Anthropic's most cutting-edge Mythos system, as Chinese companies continue to narrow the gap with the top US labs. The Chinese tech giant is at an early stage of training a model with as many as 10tn parameters -- three times larger than Moonshot's Kimi K3, the biggest Chinese model released to date, according to three people with knowledge of the matter. The ByteDance model is being pre-trained -- a stage that typically takes three to six months -- before it is fine-tuned and released if all goes well, one of the people said. The exact model size would only be determined at a later stage. Anthropic doesn't disclose the size of its models, but industry estimates say its most advanced Mythos 5 has about 8tn parameters and Fable 5 about 5tn. While parameter count sets the fundamental capacity or memory limits for the models to store information, actual capability also depends on other factors such as data quality and training methods. ByteDance's efforts to train one of the world's largest AI models show Chinese labs' ambition to not only catch up but outperform their US peers in the most advanced level of AI. In the past weeks alone, Chinese models from Moonshot and Alibaba show strong performance on benchmarks, lagging behind only Anthropic's Fable 5 in certain areas. Mythos 5, Anthropic's most advanced model, is only available to approved organisations after a temporary ban in June due to security concerns. Industry insiders say multiple Chinese labs are in the process of training models of the size of Fable 5, while ByteDance is currently the most ambitious in pushing for the largest. ByteDance, the parent of viral video platform TikTok, has kept a low profile in its AI development as its models are mostly closed, unlike many of its Chinese peers. Its latest SeeDance model ranks among the most advanced globally in video generation, while its flagship consumer-facing model Doubao is the most popular in China with 324mn monthly active users. Over the past three years, ByteDance has invested in AI more aggressively than any of the other Chinese tech giants, building out its network of data centres and hiring researchers. It has doubled down on its cloud unit, Volcano Engine, which sells AI solutions to enterprises. ByteDance also has ambitions to develop custom AI chips. Seed, its model development team led by former Google DeepMind scientist Wu Yonghui, has about 2,000 members in China and overseas. The team includes core researchers, infrastructure engineers, a data labelling team and translators. Its model development has also implemented a more independent approach that does not involve "distilling" existing models from other labs, according to one of the people. This approach has been in place for more than a year, which some believe has led to its slower development versus rivals. Model distillation is the process of compressing a large, complex AI model into a smaller, faster one by training the smaller model to copy the knowledge and outputs of the larger teacher. ByteDance's management, led by founder Zhang Yiming, believes only independent development can result in a model that outperforms rivals. Zhang reiterated his stance in an internal meeting two weeks ago, where he told the Seed team to target "world-leading model capabilities" in the long run without getting too worried about falling behind in the near term, according to the person. Chinese media Latepost and The Information first reported Zhang's comments from the recent internal meeting. ByteDance did not respond to a request for comment.
[3]
ByteDance is training a 10-trillion-parameter model to chase the frontier
The TikTok owner is building a model more than three times the size of Kimi K3, aiming to close the gap with the West's biggest systems, according to the Financial Times. ByteDance is going big, in the most literal sense. The owner of TikTok is training an AI model with around 10 trillion parameters, a scale that would put it among the largest ever built, and it is not hiding what it wants that scale for. According to the Financial Times, the model is meant to rival Anthropic's Mythos, one of the frontier systems that Chinese developers have so far struggled to match. The size is itself the statement. At roughly 10 trillion parameters, the model would be more than three times as large as Moonshot's Kimi K3, which sits among the biggest Chinese models today at about 2.8 trillion. That gap is not a rounding difference; it is the kind of leap that reorders where a model ranks against the field. It is still early, though: the project is said to be in the early pre-training phase, a stage that can take three to six months, with fine-tuning and further work to come before anything is released. Parameter count is not everything, of course. Bigger models are not automatically better, and the industry has learned that data quality, training technique and efficiency often matter as much as raw scale. Even so, committing the compute to train a model this size is a declaration in its own right, a signal that ByteDance wants to compete at the very top rather than ship a capable also-ran. It has some ground to stand on. Its Doubao assistant is already one of China's most used AI products, which gives the company both a distribution channel and a reason to want a frontier model of its own. Behind the effort sits a pointed instruction, too: founder Zhang Yiming reportedly told staff to avoid leaning on AI distillation for short-term gains, pushing them to build genuine capability rather than copy it. That warning lands amid a live controversy, since US labs have accused Chinese firms of distilling their models. Moonshot, for one, is alleged to have leaned on Anthropic's work for Kimi K3, claims China disputes and answers with its own. The fight has spilled into tooling as well, and concerns about Chinese coding tools and Claude Code show how tangled the two ecosystems have become even as they compete. ByteDance is far from the only Chinese firm scaling up. Alibaba has been pushing Qwen toward the top ranks and claims its latest is the world's number-two open-weight model, while the sector as a whole races to close the gap. The compute those ambitions demand is staggering, and Moonshot reportedly used 20,000 Nvidia chips to train Kimi K3, which hints at what a far larger model will consume. Access is becoming a strategic lever, too. China has weighed curbing overseas access to its best models, a sign that these systems are increasingly treated as national assets rather than mere products. Hanging over all of it is the chip question, because US export controls have limited China's access to the most advanced accelerators, so training a model this large tests how far Chinese firms can push with the hardware they can actually get. ByteDance does bring advantages others lack. Its enormous consumer reach through TikTok and Doubao supplies both training data and a ready place to deploy a frontier model, turning research spending into products almost immediately. There is a strategic subtext beyond the balance sheet as well, since a Chinese company matching a top Western model would be a symbolic milestone in a rivalry both governments increasingly frame in national terms. For now, the model remains a work in progress, and much can still change between an early pre-training run and a finished system. Yet the sheer size of the target shows how determined China's giants are to stand alongside the West's frontier labs rather than trail behind them.
[4]
ByteDance Is Reportedly Training a Massive New AI Model to Rival Anthropic's Mythos - Alibaba Gr Hldgs (N
ByteDance is reportedly developing a new AI model that may rival the scale of Anthropic's advanced Mythos system, in a bid to close the gap with leading U.S. AI companies. The Chinese tech giant is training an AI model with up to 10 trillion parameters, making it about three times larger than Moonshot's Kimi K3, currently China's largest released AI model. The project is still in its early stages, reported the Financial Times. ByteDance's AI model is currently in pre-training, a process that usually lasts three to six months. Its final size will be decided later, before potential fine-tuning and release if development progresses successfully, as per the report. The TikTok parent has quietly advanced its AI efforts, keeping most models private compared with other Chinese rivals. Its SeeDance video generation model is considered one of the world's most advanced, while its consumer AI assistant Doubao leads China's market with 324 million monthly active users. ByteDance founder Zhang Yiming has urged the company's AI team to focus on building a world-class model through independent development rather than chasing short-term competition. He believes long-term innovation is the key to surpassing rivals and has told the Seed team not to worry about temporary setbacks, according to the publication. ByteDance did not immediately respond to Benzinga's request for comments. China's AI Race Goes Bigger ByteDance developing a massive AI model highlights China's goal of moving beyond catching up with U.S. AI leaders to potentially surpass them. Several Chinese AI labs are reportedly developing extremely large models comparable to Fable 5, with ByteDance leading the race to build the biggest and most ambitious system. Notably, Anthropic has not revealed the exact size of its AI models, but estimates suggest its top-tier Mythos 5 model has around 8 trillion parameters, while Fable 5 may have about 5 trillion parameters. Disclaimer: This content was partially produced with the help of AI tools and was reviewed and published by Benzinga editors. Image via Shutterstock Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
Share
Copy Link
ByteDance is training an AI model with up to 10 trillion parameters, three times larger than any Chinese model released to date. The TikTok parent aims to rival Anthropic's Mythos system as Chinese companies push to close the gap with leading US AI labs through independent model development.
ByteDance is training an AI model with as many as 10 trillion parameters, positioning the Chinese tech giant to rival Anthropic's most advanced systems in what represents one of the most ambitious efforts yet to close the gap with US AI labs
1
2
. The massive new AI model would be approximately three times larger than Moonshot's Kimi K3, currently the biggest Chinese model released to date at about 2.8 trillion parameters. While Anthropic doesn't publicly disclose model sizes, industry estimates suggest its most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion parameters3
.The project remains in its early pre-training phase, a stage that typically takes three to six months before fine-tuning and potential release
4
. Three people with knowledge of the matter confirmed that the exact model size would only be determined at a later stage. Parameter count establishes the fundamental capacity for models to store information, though actual capability also depends on factors like data quality and training methods1
.ByteDance has implemented a more independent approach that does not involve model distillation from other labs, a strategy that has been in place for more than a year
2
. Model distillation compresses large AI models into smaller, faster ones by training the smaller model to copy the knowledge and outputs of the larger teacher. Some believe this independent model development approach has led to slower progress versus rivals, but ByteDance founder Zhang Yiming maintains that only independent development can result in a model that outperforms competitors1
.
Source: FT
Zhang Yiming reiterated his stance in an internal meeting two weeks ago, telling the Seed team to target world-leading model capabilities in the long run without getting too worried about falling behind in the near term, according to reports from Chinese media Latepost and The Information
3
. This directive arrives amid live controversy, as US labs have accused Chinese firms of distilling their models, with Moonshot specifically alleged to have leaned on Anthropic's work for Kimi K33
.Over the past three years, ByteDance has invested in AI more aggressively than any other Chinese tech giants, building out its network of data centers and hiring researchers
2
. The company has doubled down on its cloud unit, Volcano Engine, which sells AI solutions to enterprises, and has ambitions to develop custom AI chips. The Seed model development team, led by former Google DeepMind scientist Wu Yonghui, has about 2,000 members in China and overseas, including core researchers, infrastructure engineers, a data labeling team and translators1
.
Source: The Next Web
The TikTok parent has kept a low profile in its AI development as its models are mostly closed, unlike many Chinese peers. Its latest SeeDance model ranks among the most advanced globally in video generation, while its flagship consumer-facing model Doubao is the most popular in China with 324 million monthly active users
4
. This enormous consumer reach through TikTok and Doubao supplies both training data and a ready place to deploy frontier models, turning research spending into products almost immediately3
.Related Stories
Industry insiders say multiple Chinese labs are in the process of training models of the size of Fable 5, while ByteDance is currently the most ambitious in pushing for the largest
2
. In recent weeks, Chinese models from Moonshot and Alibaba have shown strong benchmark performance, lagging behind only Anthropic's Fable 5 in certain areas. Mythos 5, Anthropic's most advanced model, is only available to approved organizations after a temporary ban in June due to security concerns1
.Alibaba has been pushing Qwen toward the top ranks and claims its latest is the world's number-two open-weight model
3
. The compute demands for these ambitions are staggering, with Moonshot reportedly using 20,000 Nvidia chips to train Kimi K3, hinting at what a far larger model will consume. Hanging over all efforts are US chip export controls, which have limited China's access to the most advanced accelerators, making training a model this large a test of how far Chinese firms can push with available hardware3
. China has also weighed curbing overseas access to its best models, signaling these systems are increasingly treated as national assets rather than mere products, adding strategic weight to ByteDance's frontier pursuit.Summarized by
Navi
[1]
30 Jul 2026•Business and Economy

08 Dec 2024•Technology

23 Jan 2025•Technology

1
Technology

2
Technology

3
Science and Research
