뉴스로 돌아가기
제품AI Understanding 브리핑

Qwen는 Qwen3.8-Flash-Next를 다음 아키텍처의 개방형 미리보기로 제시합니다.

Qwen는 Qwen3.8-Flash-Next를 총 1,250억 개의 토큰과 60억 개의 활성 매개변수를 갖춘 다중 모드 전문가 혼합 모델로 설명하고 Qwen4를 위해 계획된 아키텍처를 미리 보여 준다고 말합니다.

5 min readRead the linked source
Source-page capture accompanying Qwen presents Qwen3.8-Flash-Next as an open-weight preview of its next architecture
소스 참조녹음된 소스
출판사
qwen.ai
소스 링크
qwen.aihttps://qwen.ai/blog?id=qwen3.8-flash-next
소스 유형
연결된 소스 — 기본 소스 상태가 설정되지 않았습니다.
또한 인용됨

마지막으로 수정된 스토리

맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

무게
신경망을 통과하는 신호의 크기를 조정하는 학습된 숫자 값입니다.
전문가 혼합(MoE)
입력당 선택된 전문가만 실행하는 특수 하위 네트워크가 있는 아키텍처입니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

출간 이후 달라진 점

  1. 처음 출판됨
  2. Decrypt adds reporting on the same planned Qwen 3.8-Flash-Next preview covered by the canonical update. It says Alibaba’s Qwen team planned a Wednesday release, describes the model as a multimodal preview of Qwen 4, and reports that benchmark scores and live weights were not yet available. The reported 125-billion total and 6-billion active-parameter figures remain unverified in the supplied material.
  3. This materially advances the eligible Qwen3.8-Flash continuing event: Investing.com reports that Alibaba has released the production Qwen3.8-Flash model with downloadable weights, QwenCloud pricing, reported benchmark scores and cost claims, while also releasing the Qwen3.8-Flash-Next preview linked to its planned Qwen4 architecture.
  4. Blockchain.News materially advances the same Qwen3.8-Flash-Next release event represented by the canonical update. It reports the release of 176-billion-parameter open weights, describes Gated DeltaNet and Qwen Sparse Attention for contexts up to 1 million tokens, cites Alibaba’s throughput figures, and reports NVIDIA validation on GB300 NVL72. These details are not independently confirmed in the provided source.
  5. The Qwen primary source materially adds architectural and testing details to the existing Qwen3.8-Flash-Next release entry: it describes the model as a multimodal MoE preview for Qwen4, gives 125 billion total and 6 billion active parameters, and identifies 72.5-gigabyte and 78.9-gigabyte quantized versions tested on DGX Spark.
Source video from qwen.ai · shown with attribution.

무슨 일이 일어났나요?

Qwen’s source presents Qwen3.8-Flash-Next as an open- multimodal mixture-of-experts model and an early preview of the architecture used in Qwen4. The source says the model contains 125 billion tokens in total but activates 6 billion, a design it associates with a substantial performance boost. It also references quantized versions tested on NVIDIA’s DGX Spark system.

The Qwen page introduces Qwen3.8-Flash-Next as “another open weights model from Qwen.” It describes the system as a multimodal MoE model and says it serves as an early preview of the architecture used in Qwen4. That makes the model itself the central development, while the Qwen4 reference provides forward-looking context about the company’s model family. The source does not provide a formal launch date or a detailed release announcement beyond this description.

The source gives two scale figures: 125 billion total tokens and 6 billion active parameters. It presents the difference between those figures as the reason the model receives a “pretty big performance boost.” That is a claim made by the source, not a result independently demonstrated in the supplied material. No benchmark table, comparison model, test protocol, response-time measurement or quality score is included, so the practical meaning of the claimed boost remains unverified here.

The source also says the model has been tried on a DGX Spark using Unsloth quantized versions. It specifically mentions a 72.5-gigabyte UD-IQ1_S model and a 78.9-gigabyte UD-Q2_K_XL model. These details indicate that at least some compressed or quantized forms are being examined on that computing platform. The source does not state whether those files are officially distributed by Qwen, what precision tradeoffs they make, or whether they are suitable for other hardware.

The author says exploration is continuing and identifies one preferred result from an xhigh reasoning-effort configuration of the UD-Q2_K_XL version. The source also refers to generated pelican images, including a pelican-riding-a-bicycle example. These are anecdotal demonstrations rather than a controlled evaluation. The material supplied does not establish the model’s image quality, reasoning reliability, multimodal coverage, reproducibility or general availability.

소스 세부정보: qwen.ai ↗

왜 중요한가요?

The announcement offers an early indication of Qwen’s next model architecture while emphasizing a large gap between total and active parameters. If the source’s description is accurate, the model may be relevant to developers evaluating open- systems that seek to combine broad model capacity with lower active computation. However, the source provides no independent benchmarks or deployment data.

The announcement matters because it connects an available open- model with the architecture Qwen says will inform Qwen4. Open weights can give developers and researchers an object they can inspect, adapt or run within their own environments, but the source does not specify the legal terms governing those uses. The practical significance therefore depends on licensing and documentation that are not included here.

The model’s stated architecture highlights a recurring engineering tradeoff: a system can have a large total parameter count while activating a much smaller subset for a particular request. In this source, Qwen presents 125 billion total parameters and 6 billion active parameters as a way to pursue capacity with a lower active workload. Whether that translates into lower cost, faster responses or better quality cannot be concluded from the source alone.

The multimodal description could make the model relevant beyond text-only applications. Yet “multimodal” is not further defined in the supplied material. There is no inventory of input or output types, no description of supported languages or media, and no evidence about how the model handles real-world images, complex instructions or safety-sensitive content. The generated pelican examples show only that visual outputs were attempted in the reported exploration.

The quantized versions are also practically important because they make the model’s size and hardware demands part of the story. The source identifies files measured at 72.5 and 78.9 gigabytes and says they were tested on a DGX Spark. It does not say whether ordinary developers can run them, how much memory they require in operation, or how quality changes under quantization. Those omissions limit what can responsibly be inferred about accessibility.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

Important unknowns include the model’s release date, license, supported modalities, benchmark results, hardware requirements, safety evaluations and access conditions. Further testing will be needed to determine whether the reported active-parameter design delivers consistent advantages across practical workloads, and whether the quantized versions preserve the model’s capabilities.

The first priority is confirmation of the formal release terms. Readers and developers need to know whether Qwen3.8-Flash-Next is officially downloadable, which weights and formats are provided, what license applies, and whether use is permitted commercially or only for research. None of those conditions is stated in the supplied source.

Independent evaluations should test the source’s performance claim. Useful follow-up would compare the 6-billion-active-parameter configuration with other open- models on multimodal understanding, generation, reasoning, coding and long-context tasks, while reporting hardware, quantization settings and latency. The present source offers no such controlled comparison.

The model’s safety and reliability profile also remains unknown. No red-team findings, refusal testing, hallucination measurements, privacy analysis or misuse safeguards appear in the material. Before the system is used in consequential settings, developers would need evidence about failure modes across both text and visual inputs, along with clear guidance on human review.

Finally, the Qwen4 connection should be treated as an architectural preview rather than a promise about a future release. The source says the model previews architecture used in Qwen4, but it gives no schedule, specifications or assurance that the final family will match this model. Continued testing may clarify whether the reported quantized configurations and active-parameter design are durable features or exploratory choices.

관련 가이드 및 퀴즈

AI 모델 설명트랜스포머AI 트레이닝ChatGPT와 LLM알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.

업데이트 및 수정

이 정식 스토리는 진행 중인 이벤트가 실질적으로 변경될 때 업데이트됩니다. URL과 원래 출판 날짜는 절대 변경되지 않습니다.

  • The Qwen primary source materially adds architectural and testing details to the existing Qwen3.8-Flash-Next release entry: it describes the model as a multimodal MoE preview for Qwen4, gives 125 billion total and 6 billion active parameters, and identifies 72.5-gigabyte and 78.9-gigabyte quantized versions tested on DGX Spark.
  • Blockchain.News materially advances the same Qwen3.8-Flash-Next release event represented by the canonical update. It reports the release of 176-billion-parameter open weights, describes Gated DeltaNet and Qwen Sparse Attention for contexts up to 1 million tokens, cites Alibaba’s throughput figures, and reports NVIDIA validation on GB300 NVL72. These details are not independently confirmed in the provided source.
  • This materially advances the eligible Qwen3.8-Flash continuing event: Investing.com reports that Alibaba has released the production Qwen3.8-Flash model with downloadable weights, QwenCloud pricing, reported benchmark scores and cost claims, while also releasing the Qwen3.8-Flash-Next preview linked to its planned Qwen4 architecture.
  • Decrypt adds reporting on the same planned Qwen 3.8-Flash-Next preview covered by the canonical update. It says Alibaba’s Qwen team planned a Wednesday release, describes the model as a multimodal preview of Qwen 4, and reports that benchmark scores and live weights were not yet available. The reported 125-billion total and 6-billion active-parameter figures remain unverified in the supplied material.
공개 수정 로그 보기
이것이 유용하다고 생각하시나요?