ByteDance trains massive 10 trillion parameter AI model to rival Anthropic's Mythos

Reviewed byNidhi Govil

4 Sources

Share

ByteDance is training an AI model with up to 10 trillion parameters, three times larger than any Chinese model released to date. The TikTok parent aims to rival Anthropic's Mythos system as Chinese companies push to close the gap with leading US AI labs through independent model development.

ByteDance Pushes for Frontier AI with 10 Trillion Parameter Model

ByteDance is training an AI model with as many as 10 trillion parameters, positioning the Chinese tech giant to rival Anthropic's most advanced systems in what represents one of the most ambitious efforts yet to close the gap with US AI labs

1

2

. The massive new AI model would be approximately three times larger than Moonshot's Kimi K3, currently the biggest Chinese model released to date at about 2.8 trillion parameters. While Anthropic doesn't publicly disclose model sizes, industry estimates suggest its most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion parameters

3

.

The project remains in its early pre-training phase, a stage that typically takes three to six months before fine-tuning and potential release

4

. Three people with knowledge of the matter confirmed that the exact model size would only be determined at a later stage. Parameter count establishes the fundamental capacity for models to store information, though actual capability also depends on factors like data quality and training methods

1

.

Independent Model Development Strategy Sets ByteDance Apart

ByteDance has implemented a more independent approach that does not involve model distillation from other labs, a strategy that has been in place for more than a year

2

. Model distillation compresses large AI models into smaller, faster ones by training the smaller model to copy the knowledge and outputs of the larger teacher. Some believe this independent model development approach has led to slower progress versus rivals, but ByteDance founder Zhang Yiming maintains that only independent development can result in a model that outperforms competitors

1

.

Source: FT

Source: FT

Zhang Yiming reiterated his stance in an internal meeting two weeks ago, telling the Seed team to target world-leading model capabilities in the long run without getting too worried about falling behind in the near term, according to reports from Chinese media Latepost and The Information

3

. This directive arrives amid live controversy, as US labs have accused Chinese firms of distilling their models, with Moonshot specifically alleged to have leaned on Anthropic's work for Kimi K3

3

.

ByteDance's AI Infrastructure and Market Position

Over the past three years, ByteDance has invested in AI more aggressively than any other Chinese tech giants, building out its network of data centers and hiring researchers

2

. The company has doubled down on its cloud unit, Volcano Engine, which sells AI solutions to enterprises, and has ambitions to develop custom AI chips. The Seed model development team, led by former Google DeepMind scientist Wu Yonghui, has about 2,000 members in China and overseas, including core researchers, infrastructure engineers, a data labeling team and translators

1

.

Source: The Next Web

Source: The Next Web

The TikTok parent has kept a low profile in its AI development as its models are mostly closed, unlike many Chinese peers. Its latest SeeDance model ranks among the most advanced globally in video generation, while its flagship consumer-facing model Doubao is the most popular in China with 324 million monthly active users

4

. This enormous consumer reach through TikTok and Doubao supplies both training data and a ready place to deploy frontier models, turning research spending into products almost immediately

3

.

Chinese AI Labs Race to Match Western Frontier Models

Industry insiders say multiple Chinese labs are in the process of training models of the size of Fable 5, while ByteDance is currently the most ambitious in pushing for the largest

2

. In recent weeks, Chinese models from Moonshot and Alibaba have shown strong benchmark performance, lagging behind only Anthropic's Fable 5 in certain areas. Mythos 5, Anthropic's most advanced model, is only available to approved organizations after a temporary ban in June due to security concerns

1

.

Alibaba has been pushing Qwen toward the top ranks and claims its latest is the world's number-two open-weight model

3

. The compute demands for these ambitions are staggering, with Moonshot reportedly using 20,000 Nvidia chips to train Kimi K3, hinting at what a far larger model will consume. Hanging over all efforts are US chip export controls, which have limited China's access to the most advanced accelerators, making training a model this large a test of how far Chinese firms can push with available hardware

3

. China has also weighed curbing overseas access to its best models, signaling these systems are increasingly treated as national assets rather than mere products, adding strategic weight to ByteDance's frontier pursuit.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved