On August 7, the Financial Times reported, citing three informed sources, that ByteDance has begun pre-training a large model with up to 100 trillion parameters. The project is still in its early stages, and the final scale has yet to be determined. If trained to the upper limit, it would exceed KimiK3 by more than three times, making it the largest model known among Chinese teams in terms of parameter scale. ByteDance has directly pushed the model size close to Anthropic's flagship. Although Anthropic has never disclosed the parameter count for Mythos5, the Financial Times cited industry estimates suggesting that Mythos5 has around 80 trillion parameters, while Fable5 has about 50 trillion. If ByteDance's new model achieves 100 trillion parameters, it would at least match the scale of the leading closed-source models in the U.S. This model is currently in pre-training (training the model's foundational capabilities using massive amounts of data). The Financial Times noted that this phase typically takes 3 to 6 months, followed by further post-training before a successful release. More parameters do not necessarily equate to stronger capabilities; the final performance also depends on architecture, training data, and training methods. Zhang Yiming recently made it clear internally at Seed that he opposes distilling competitor models, urging the team to accept short-term setbacks and strive to reach the world's top tier through self-research. This new model, with a maximum of 100 trillion parameters, represents ByteDance's most aggressive bet on scale to date.
All Comments