Add Decrypt as your preferred source to see more of our stories on Google.
In brief
Alibaba's Qwen team is set to release Qwen 3.8-Flash-Next on Wednesday, a Mixture-of-Experts model described as a preview of the Qwen4 architecture.
The team's pre-release briefing cites 125 billion total parameters with only 6 billion active per token.
Hard benchmark scores haven't been published yet, and the weights aren't live on ModelScope as of this writing.
Alibaba is set to release Qwen 3.8-Flash-Next on Wednesday, a 125-billion-parameter model that activates just 6 billion per token. The Qwen team framed it as a preview of the next-generation Qwen 4 architecture, not a finished flagship.
There is no official information on the model, but based on rumors, it will likely be a mixture-of-experts system, a design that splits the network into many specialized sub-models and lights up only the relevant ones for each task. A 125 billion parameters model thus would run with the compute bill of one that’s just 6 billion parameters.
Parameters are basically all the dials a model can tweak. The more parameters, the more capable a model is and the more computing power it will require. A mixture of experts makes it possible for an extremely powerful model to only activate what it needs to provide the best output without wasting resources.
Qwen 3.8 Flash Next is releasing Tomorrow. 125B paramters +51B N-gram and 6B active. Its based on the next generation Qwen 4 architecture.
Alibaba’s Qwen team does describe the model as multimodal and built on the upcoming Qwen 4 architecture, and says it shipped the early build so developers can prepare for the full family.
Why the "3.8" isn't the news
Alibaba has put out a teaser, though the team calls it a preview. The plan is to ship the architecture improvements now, ahead of the complete Qwen 4 rollout. Hugging Face, where the weights also live, also describes it as "a preview of the Qwen 4 architecture."
Hard benchmarks haven't landed yet. Qwen hasn't published side-by-side scores against its own Qwen 3 line or Western rivals, so the 125 billion and 6 billion paramater figures are not verified, and we can only speculate on its performance.
China's open-weight cadence has been relentless. A mysterious free model, Ox Alpha, recently beat Anthropic’s Fable on certain coding benchmarks, with no known builder behind it. Alibaba, DeepSeek, and Moonshot have all shipped capable weights anyone can download, fine-tune, and run.
Open weights let developers build without sending data to a closed API, and they undercut the cost of hosted models. That's why an open 125 billion-parameter model with 6 billion active parameters matters: It puts near-frontier capability on commodity hardware.
But we’ll have to wait for the numbers. Qwen hadn't posted benchmark scores for the release as of this writing.
Daily Debrief Newsletter
Start every day with the top news stories right now, plus original features, a podcast, videos and more.