ByteDance reportedly plans 10 trillion total-parameter model with 30,000 GPUs

53 minutes ago 8

ByteDance is reportedly gearing up to pre-train an AI model with roughly 10 trillion parameters, a scale that would dwarf every known Chinese AI system and put the company in direct competition with the most advanced Western labs. The effort would require approximately 30,000 GPUs and an estimated 3 to 6 months of continuous pre-training.

To put 10 trillion parameters in perspective, that is more than three times the size of Moonshot AI’s Kimi K3, which sits at 2.8 trillion parameters and currently ranks among the largest models produced in China. Parameters are essentially the knobs an AI model tunes during training to learn patterns in data.

What ByteDance is actually building

The model in question uses a Mixture of Experts (MoE) architecture. Rather than activating every parameter for every query, MoE models route each input to a subset of specialized “expert” sub-networks. This makes them far more efficient to run at inference time than a dense model of the same total size.

ByteDance’s Seed AI team, which reportedly consists of around 2,000 staff members, is leading the effort. The team has prior experience scaling training runs, having previously trained models up to 175 billion parameters using approximately 12,000 GPUs. ByteDance has published research on its MegaScale infrastructure, which was designed to manage distributed training systems exceeding 10,000 GPUs.

The pre-training phase is just the first step. After the initial run, the model would undergo fine-tuning before any potential public release. ByteDance founder Zhang Yiming has reportedly directed the Seed AI team to focus on genuine technological breakthroughs rather than relying on distillation from Western AI systems.

Why ByteDance thinks it can pull this off

ByteDance has a few structural advantages that make this more than just a vanity project. The company operates TikTok, which serves over a billion users globally, and Doubao, one of the most widely used AI assistants in China. That consumer footprint generates an enormous volume of interaction data that can inform both training and fine-tuning decisions.

The elephant in the room, of course, is the US export restrictions on advanced AI chips. Washington has progressively tightened controls on the sale of high-end Nvidia GPUs and other accelerators to Chinese entities. ByteDance has not publicly detailed which specific GPU models it plans to use for this training run, but the company’s ability to marshal 30,000 of them suggests it has found a workable path through the restrictions.

The competitive landscape this reshapes

ByteDance is not the only Chinese AI lab with big ambitions. Alibaba’s Qwen team, Baidu, and startups like DeepSeek and Moonshot AI have all been pushing the boundaries of model scale. But a 10 trillion parameter model would represent a significant leap beyond what any of them have publicly announced or deployed.

The target also puts ByteDance in the conversation with Western frontier labs. Anthropic’s latest systems, the competitive benchmark ByteDance is reportedly aiming to match, represent the current ceiling of publicly known AI capability.

The 3 to 6 month pre-training window means results, or complications, could start emerging before the end of 2026.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article