The Ali Baba Qwen team plans to launch Qwen 3.8-Flash-Next on Wednesday and position it as a preview of the next generation of Qwen 4 structures rather than a complete flagship model. At this stage, the focus of disclosure is on the structure orientation and parameter design, and performance performance has not been made public simultaneously.
1.2 billion total parameters, only 6 billion activated once
According to the pre-publishing note, the model is designed for a total of 12.5 billion parameters, but only 6 billion parameters are activated per token. The Qwen team described it as a multi-mode model and indicated that it had been released in advance to allow developers to adapt to the complete Qwen 4 series.
Based on the information available, the model probably uses a hybrid expert structure, i.e., the splitting of different capabilities into multiple submodels, and the reasoning uses only those parts relevant to the current mission. This would reduce the calculus required for a single operation while maintaining a larger model size.
Positioning is not an official flagship.
The Qwen team and Hugging Face present this version more closely to the Preview. This means that the focus of the release is not just to give a whole new generation of models, but to deliver some of the structural improvements ahead of time, allowing developers to prepare for subsequent official versions.
The designation "3.8" also reflects that it is more like a transitional version than the official landing of Qwen 4. According to the article, the team is now placing more emphasis on the direction of the new structure than on packaging the update in a complete replacement.
Benchmark tests are still not published
Qwen had not published the results of the comparison with its own Qwen 3 series or with mainstream models overseas, nor had model weights been placed on ModeScope. As a result, the potential capabilities of the outside world can only be judged on the basis of parameter information disclosed by the team, and actual performance remains to be validated.
The article mentions that the rate of release of open source models in China has continued to accelerate in recent years. Teams such as Ali Baba, DeepSeek and Moonshott are continuously rolling out downloadable, fine-tuned, locally run open weight models. If Qwen released this time as planned, an open model with an overall parameter of 12.5 billion and a activated parameter of $6 billion could further lower the threshold for developers to use high performance models.
