Retour aux Actualités
ProduitBriefing AI Understanding

Qwen présente Qwen3.8-Flash-Next comme un aperçu ouvert de sa prochaine architecture

Qwen décrit Qwen3.8-Flash-Next comme un modèle multimodal de mélange d'experts avec 125 milliards de jetons au total et 6 milliards de paramètres actifs, et indique qu'il donne un aperçu de l'architecture prévue pour Qwen4.

5 min readRead the linked source
Source-page capture accompanying Qwen presents Qwen3.8-Flash-Next as an open-weight preview of its next architecture
Référence sourceSource enregistrée
Éditeur
qwen.ai
Lien source
qwen.aihttps://qwen.ai/blog?id=qwen3.8-flash-next
Type de source
Source liée : le statut de source principale n'a pas été établi.
Également cité

Histoire révisée pour la dernière fois

ContexteComprenez cela en 60 secondes

Commencez ici

Termes clés

Poids
Une valeur numérique apprise qui met à l'échelle les signaux transitant par un réseau neuronal.
Mélange d'experts (MoE)
Une architecture avec des sous-réseaux spécialisés où seuls des experts sélectionnés s'exécutent par entrée.
Mémoire (mémoire de l'agent)
Contexte stocké qu'un agent IA utilise au fil des étapes ou des sessions pour améliorer la continuité.
Testez-vousQuiz sur les modèles d'IA expliqués

Ce qui a changé depuis la publication

  1. Première publication
  2. Decrypt adds reporting on the same planned Qwen 3.8-Flash-Next preview covered by the canonical update. It says Alibaba’s Qwen team planned a Wednesday release, describes the model as a multimodal preview of Qwen 4, and reports that benchmark scores and live weights were not yet available. The reported 125-billion total and 6-billion active-parameter figures remain unverified in the supplied material.
  3. This materially advances the eligible Qwen3.8-Flash continuing event: Investing.com reports that Alibaba has released the production Qwen3.8-Flash model with downloadable weights, QwenCloud pricing, reported benchmark scores and cost claims, while also releasing the Qwen3.8-Flash-Next preview linked to its planned Qwen4 architecture.
  4. Blockchain.News materially advances the same Qwen3.8-Flash-Next release event represented by the canonical update. It reports the release of 176-billion-parameter open weights, describes Gated DeltaNet and Qwen Sparse Attention for contexts up to 1 million tokens, cites Alibaba’s throughput figures, and reports NVIDIA validation on GB300 NVL72. These details are not independently confirmed in the provided source.
  5. The Qwen primary source materially adds architectural and testing details to the existing Qwen3.8-Flash-Next release entry: it describes the model as a multimodal MoE preview for Qwen4, gives 125 billion total and 6 billion active parameters, and identifies 72.5-gigabyte and 78.9-gigabyte quantized versions tested on DGX Spark.
Source video from qwen.ai · shown with attribution.

Que s'est-il passé

Qwen’s source presents Qwen3.8-Flash-Next as an open- multimodal mixture-of-experts model and an early preview of the architecture used in Qwen4. The source says the model contains 125 billion tokens in total but activates 6 billion, a design it associates with a substantial performance boost. It also references quantized versions tested on NVIDIA’s DGX Spark system.

The Qwen page introduces Qwen3.8-Flash-Next as “another open weights model from Qwen.” It describes the system as a multimodal MoE model and says it serves as an early preview of the architecture used in Qwen4. That makes the model itself the central development, while the Qwen4 reference provides forward-looking context about the company’s model family. The source does not provide a formal launch date or a detailed release announcement beyond this description.

The source gives two scale figures: 125 billion total tokens and 6 billion active parameters. It presents the difference between those figures as the reason the model receives a “pretty big performance boost.” That is a claim made by the source, not a result independently demonstrated in the supplied material. No benchmark table, comparison model, test protocol, response-time measurement or quality score is included, so the practical meaning of the claimed boost remains unverified here.

The source also says the model has been tried on a DGX Spark using Unsloth quantized versions. It specifically mentions a 72.5-gigabyte UD-IQ1_S model and a 78.9-gigabyte UD-Q2_K_XL model. These details indicate that at least some compressed or quantized forms are being examined on that computing platform. The source does not state whether those files are officially distributed by Qwen, what precision tradeoffs they make, or whether they are suitable for other hardware.

The author says exploration is continuing and identifies one preferred result from an xhigh reasoning-effort configuration of the UD-Q2_K_XL version. The source also refers to generated pelican images, including a pelican-riding-a-bicycle example. These are anecdotal demonstrations rather than a controlled evaluation. The material supplied does not establish the model’s image quality, reasoning reliability, multimodal coverage, reproducibility or general availability.

Détails de la source: qwen.ai ↗

Pourquoi c'est important

The announcement offers an early indication of Qwen’s next model architecture while emphasizing a large gap between total and active parameters. If the source’s description is accurate, the model may be relevant to developers evaluating open- systems that seek to combine broad model capacity with lower active computation. However, the source provides no independent benchmarks or deployment data.

The announcement matters because it connects an available open- model with the architecture Qwen says will inform Qwen4. Open weights can give developers and researchers an object they can inspect, adapt or run within their own environments, but the source does not specify the legal terms governing those uses. The practical significance therefore depends on licensing and documentation that are not included here.

The model’s stated architecture highlights a recurring engineering tradeoff: a system can have a large total parameter count while activating a much smaller subset for a particular request. In this source, Qwen presents 125 billion total parameters and 6 billion active parameters as a way to pursue capacity with a lower active workload. Whether that translates into lower cost, faster responses or better quality cannot be concluded from the source alone.

The multimodal description could make the model relevant beyond text-only applications. Yet “multimodal” is not further defined in the supplied material. There is no inventory of input or output types, no description of supported languages or media, and no evidence about how the model handles real-world images, complex instructions or safety-sensitive content. The generated pelican examples show only that visual outputs were attempted in the reported exploration.

The quantized versions are also practically important because they make the model’s size and hardware demands part of the story. The source identifies files measured at 72.5 and 78.9 gigabytes and says they were tested on a DGX Spark. It does not say whether ordinary developers can run them, how much memory they require in operation, or how quality changes under quantization. Those omissions limit what can responsibly be inferred about accessibility.

Interactive Mechanism

Mécanisme interactif : comment cela fonctionne réellement

Explorez de manière interactive la technologie sous-jacente à ce développement.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Vérification de concept interactive+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Que regarder ensuite

Important unknowns include the model’s release date, license, supported modalities, benchmark results, hardware requirements, safety evaluations and access conditions. Further testing will be needed to determine whether the reported active-parameter design delivers consistent advantages across practical workloads, and whether the quantized versions preserve the model’s capabilities.

The first priority is confirmation of the formal release terms. Readers and developers need to know whether Qwen3.8-Flash-Next is officially downloadable, which weights and formats are provided, what license applies, and whether use is permitted commercially or only for research. None of those conditions is stated in the supplied source.

Independent evaluations should test the source’s performance claim. Useful follow-up would compare the 6-billion-active-parameter configuration with other open- models on multimodal understanding, generation, reasoning, coding and long-context tasks, while reporting hardware, quantization settings and latency. The present source offers no such controlled comparison.

The model’s safety and reliability profile also remains unknown. No red-team findings, refusal testing, hallucination measurements, privacy analysis or misuse safeguards appear in the material. Before the system is used in consequential settings, developers would need evidence about failure modes across both text and visual inputs, along with clear guidance on human review.

Finally, the Qwen4 connection should be treated as an architectural preview rather than a promise about a future release. The source says the model previews architecture used in Qwen4, but it gives no schedule, specifications or assurance that the final family will match this model. Continued testing may clarify whether the reported quantized configurations and active-parameter design are durable features or exploratory choices.

Guides et quiz associés

Modèles d'IA expliquésTransformateursFormation IAChatGPT et LLMTestez ce que vous savez : essayez un quiz gratuit sur l'IARecherchez un terme d'IA dans notre glossaireSuivez le suivi des versions du modèle AI

Mises à jour et corrections

Cette histoire canonique est mise à jour lorsque l’événement en développement change matériellement. Son URL et sa date de publication originale ne changent jamais.

  • The Qwen primary source materially adds architectural and testing details to the existing Qwen3.8-Flash-Next release entry: it describes the model as a multimodal MoE preview for Qwen4, gives 125 billion total and 6 billion active parameters, and identifies 72.5-gigabyte and 78.9-gigabyte quantized versions tested on DGX Spark.
  • Blockchain.News materially advances the same Qwen3.8-Flash-Next release event represented by the canonical update. It reports the release of 176-billion-parameter open weights, describes Gated DeltaNet and Qwen Sparse Attention for contexts up to 1 million tokens, cites Alibaba’s throughput figures, and reports NVIDIA validation on GB300 NVL72. These details are not independently confirmed in the provided source.
  • This materially advances the eligible Qwen3.8-Flash continuing event: Investing.com reports that Alibaba has released the production Qwen3.8-Flash model with downloadable weights, QwenCloud pricing, reported benchmark scores and cost claims, while also releasing the Qwen3.8-Flash-Next preview linked to its planned Qwen4 architecture.
  • Decrypt adds reporting on the same planned Qwen 3.8-Flash-Next preview covered by the canonical update. It says Alibaba’s Qwen team planned a Wednesday release, describes the model as a multimodal preview of Qwen 4, and reports that benchmark scores and live weights were not yet available. The reported 125-billion total and 6-billion active-parameter figures remain unverified in the supplied material.
Voir le journal des corrections publiques
Vous avez trouvé cela utile ?