Voltar às notícias
ProdutoInstruções AI Understanding

Decrypt reports Alibaba plans Qwen 3.8-Flash-Next preview of Qwen 4 architecture

Decrypt reports that Alibaba’s Qwen team plans to release Qwen 3.8-Flash-Next as a multimodal preview of the forthcoming Qwen 4 architecture. The report says the model is described as having 125 billion total parameters and 6 billion active per token, but those specifications, its performance, and its availability…

Por 6 min read
AI-generated editorial illustration accompanying Decrypt reports Alibaba plans Qwen 3.8-Flash-Next preview of Qwen 4 architecture
A versão curta

Decrypt reports that Alibaba’s Qwen team plans to release Qwen 3.8-Flash-Next as a multimodal preview of the forthcoming Qwen 4 architecture. The report says the model is described as having 125 billion total parameters and 6 billion active per token, but those specifications, its performance, and its availability…

O que aconteceu

Decrypt reports that Alibaba’s Qwen team plans to release Qwen 3.8-Flash-Next on Wednesday as an early preview of the next-generation Qwen 4 architecture. The report describes the model as multimodal and says the team intended the early release to help developers prepare for the broader Qwen 4 family. Decrypt also reports that benchmark scores had not been published and that the weights were not live on ModelScope when the article was written.

Decrypt reports that Alibaba’s Qwen team planned to release Qwen 3.8-Flash-Next on Wednesday. The article characterizes it as a preview rather than a completed flagship model, and says the team presented it as an early look at the architecture expected to underpin Qwen 4. According to Decrypt, the model is multimodal, although the supplied report does not specify which input or output modalities are supported or describe any demonstrated use case. The description therefore identifies the intended role of the release, while leaving the practical behavior of the preview unspecified in the supplied report.

The article says the pre-release briefing cited 125 billion total parameters and 6 billion active parameters per token. Decrypt explains that this configuration would be consistent with a mixture-of-experts design, in which only selected portions of a larger network are used for an individual request. However, the report also says that the 125-billion and 6-billion figures were not verified in the supplied material. A separate post from an account identified as AiBattle mentioned an additional 51 billion n-gram parameters, but that post is not sufficient evidence for treating the figure as established fact. Those figures describe the reported configuration, not a confirmed measurement of how many parameters would be used in every request or how the model would perform.

Decrypt says Alibaba’s Qwen team framed the early build as a way for developers to prepare for the full Qwen 4 family. At the time of writing, the report says hard benchmark results had not been published, there were no side-by-side scores against the Qwen 3 line or Western competitors, and the weights were not live on ModelScope. The article also says Hugging Face described the model as a preview of the Qwen 4 architecture, but the supplied material does not establish a public download, a final release package, or the model’s licensing terms. No official Alibaba announcement or independent technical evaluation is included in the source. These gaps leave the status of the preview, its public availability, and the meaning of the reported figures unresolved in the source material.

Leia a fonte primária: decrypt.co

Por que isso importa

If the reported specifications are accurate, Qwen 3.8-Flash-Next would represent a large model designed to use substantially less computation for each token than its total parameter count suggests. Decrypt presents the release as part of the continuing expansion of open-weight models that developers can download, fine-tune, and run without sending data to a closed API. The practical significance remains uncertain until the model, licensing, benchmarks, and hardware requirements are available.

The reported design matters because a model can have a large total parameter count while activating only a smaller subset for each token. In principle, that can reduce the computation required for individual requests compared with a dense model of the same total size. Decrypt uses this reasoning to explain why Qwen 3.8-Flash-Next could be significant to developers seeking capable systems with lower serving costs. That implication is architectural, not a demonstrated result: the source provides no measured latency, energy use, memory requirement, or cost comparison. The source therefore supplies a reason to examine the architecture, but not evidence that the efficiency implied by it has been achieved in deployment.

Decrypt places the reported release within the wider movement toward open-weight AI models. The article says open weights can allow developers to download, fine-tune, and run a model locally or on their own infrastructure instead of sending data to a closed application-programming interface. That can offer organizations more control over data handling and customization, but the supplied report does not establish that Qwen 3.8-Flash-Next is available under an open license, that it can run on ordinary hardware, or that its operating costs would be lower in practice. The distinction between the general possibilities of open weights and the specific status of this model remains important in interpreting the report.

The potential public and industry impact therefore depends on evidence that is not yet available. A genuinely efficient model with strong capabilities could widen access to advanced language and multimodal tools, especially for developers who cannot afford the largest hosted systems. It could also intensify competition among open-weight model providers. Decrypt’s suggestion that the model may offer near-frontier capability on commodity hardware is not supported by benchmark results in the article, so it should be treated as an unverified possibility rather than a finding. The report also provides no safety assessment, privacy analysis, or information about restrictions on downstream use. For now, those possible effects remain dependent on the missing release information and evaluation evidence described above.

O que assistir a seguir

The key tests are whether Alibaba releases the model as described, publishes technical documentation and licensing terms, and provides reproducible comparisons with earlier Qwen models and other systems. Readers should also watch for independent evaluations of quality, speed, memory use, safety behavior, and multimodal performance. The supplied report does not establish whether the model was released on schedule, where it can be obtained, or whether the reported architecture delivers the claimed efficiency.

The first question is whether Alibaba releases Qwen 3.8-Flash-Next as scheduled and makes usable weights or an access method available. The supplied report does not confirm a completed release. A meaningful follow-up should identify the exact model files, license, supported modalities, context limits, hardware requirements, and whether the final system differs from the preview described by Decrypt. The release status and package details will determine what can actually be assessed.

The second question is performance. Qwen should ideally publish evaluation methods and results covering general language tasks, coding, reasoning, multimodal inputs, latency, memory use, and cost. Independent evaluators should test the model under comparable conditions rather than relying on parameter counts. Until those results appear, the reported architecture cannot establish that Qwen 3.8-Flash-Next matches or exceeds Qwen 3 models or systems from other providers. The available evidence therefore remains limited to the description of the planned preview.

Safety and deployment details also require scrutiny. The source does not describe safeguards, training data, geographic availability, usage restrictions, or procedures for reporting harmful behavior. Developers considering local deployment will need to assess not only inference efficiency but also model reliability, refusal behavior, privacy implications, and the risks of fine-tuning. The most important unknown is whether this preview represents a consequential technical improvement or mainly an early architectural signal ahead of Qwen 4. Those questions remain open until the release and supporting documentation can be examined.

Guias e questionários relacionados

ChatGPT e LLMModelos de IA explicadosTransformadoresTreinamento de IATeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossário
Achou isso útil?