뉴스로 돌아가기
제품AI Understanding 브리핑

Alibaba, 다중 모드 에이전트 작업을 위한 Qwen3.8-Omni-Flash 출시

Alibaba는 1백만 개의 컨텍스트 토큰을 지원하는 기본 옴니모달 모델인 Qwen3.8-Omni-Flash를 출시했습니다. 이 모델은 상당히 낮은 API 비용으로 Gemini 3.8 Flash에 필적하는 오디오 및 비디오 성능을 달성한다고 주장합니다.

5 min readRead the linked source
Source-provided image accompanying Alibaba releases Qwen3.8-Omni-Flash for multimodal agent tasks
소스 참조녹음된 소스
출판사
gigazine.net
소스 링크
gigazine.nethttps://gigazine.net/gsc_news/en/20260918-qwen-3-8-omni-flash/
소스 유형
연결된 소스 — 기본 소스 상태가 설정되지 않았습니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

API(애플리케이션 프로그래밍 인터페이스)
한 소프트웨어 시스템이 다른 시스템에 요청을 보내고 응답을 받는 구조화된 방식입니다.
컨텍스트 창
언어 모델이 한 번에 처리할 수 있는 최대 입력 토큰 양입니다.
벤치마크
모델 성능을 측정하고 비교하는 데 사용되는 표준화된 테스트 또는 데이터 세트입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

Alibaba released Qwen3.8-Omni-Flash, a new model in its Qwen series designed for omnimodal processing. The model supports a 1 million token and targets workflows such as video editing, music video production, and real-time conversation. According to Gigazine, the model outperforms its predecessor, Qwen3.5-Omni-Plus, by over 25% on average across 29 benchmarks. Alibaba claims the new model achieves audio and video performance close to or surpassing Gemini 3.8 Flash, while reducing API costs for voice input by over 98% and voice/video input by over 93% compared to the previous version. The release includes extensions to Qwen-MM-Plugins and an open-source harness called Qwen-Live Harness, though the latter's GitHub page was inaccessible at the time of reporting.

Alibaba has added Qwen3.8-Omni-Flash to its Qwen series, describing it as a next-generation native omnimodal model. The model is designed to handle a wide range of workflows, including video editing, music video production, film production, audiovisual summarization, and real-time conversation. It supports a of 1 million tokens, which is a substantial increase that allows for the processing of long-form audio and video files.

According to Gigazine, the model delivers superior results compared to its predecessor, Qwen3.5-Omni-Plus. Across 29 benchmarks, the average score is reported to be more than 25% higher. Specific improvements are noted in core functions such as understanding long audio, audio/video inference, captioning, and multi-speaker recognition. For instance, it achieved a score of 82.7 in LongAudioSpan and 89.7 in AliMeeting, a for transcribing multi-person, multi-channel conference audio in Chinese.

Alibaba claims that Qwen3.8-Omni-Flash scales data, context, and agent environments to achieve audio and video performance close to that of Gemini 3.8 Flash, with overall audio performance surpassing it. The company highlights significant cost reductions, stating that the API cost per hour for voice input is over 98% lower, and for voice and video input, it is over 93% lower than the previous model. These cost claims are central to the model's value proposition for enterprise users.

To support the model's capabilities, Alibaba has extended Qwen-MM-Plugins to add on-demand perception, tool usage, and workflow execution capabilities for long audio and video. The company also announced the open-sourcing of Qwen-Live Harness, a comprehensive harness designed based on the Qwen3.8-Omni-Flash-Realtime API. However, Gigazine notes that at the time of writing, the Qwen-Live Harness GitHub page was not accessible and returned a 404 error, indicating potential delays or issues with the public release of this component.

소스 세부정보: gigazine.net ↗

왜 중요한가요?

This release is significant because it addresses the high cost and complexity of processing long-form audio and video data in AI agents. By claiming near-frontier performance at a fraction of the cost of previous models, Alibaba is positioning Qwen3.8-Omni-Flash as a practical tool for enterprise applications that require real-time multimodal reasoning. The shift from treating audio and video as simple perceptual inputs to core media for agent reasoning could accelerate the integration of AI into real-world productivity scenarios, such as automated media production and complex conversational agents. However, the specific comparisons to Gemini 3.8 Flash are based on Alibaba's claims and have not been independently verified by third-party evaluators.

The release of Qwen3.8-Omni-Flash is significant for the AI industry because it targets a specific pain point: the high cost and technical complexity of processing long-form multimodal data. By claiming near-frontier performance at a drastically reduced cost, Alibaba is challenging the status quo of multimodal AI pricing and accessibility. This could lead to a broader adoption of AI agents in industries that rely heavily on audio and video data, such as media production, customer service, and real-time translation.

The model's ability to handle 1 million tokens of context is a major technical advancement, as it allows for the processing of much longer audio and video files without segmentation. This is crucial for applications like film production or long-form meeting transcription, where context continuity is essential. The improvement in multi-speaker recognition and long audio understanding suggests that the model is better suited for complex, real-world scenarios than previous iterations.

However, the claims of outperforming or matching Gemini 3.8 Flash are based on Alibaba's internal benchmarks and have not been independently verified. This is a common issue in the AI industry, where companies often rely on self-reported metrics. Independent evaluations will be necessary to confirm the true performance and cost-effectiveness of the model. Until then, enterprises should approach the claims with caution and conduct their own testing before making significant investments.

The open-sourcing of Qwen-Live Harness is a positive step for the developer community, as it provides a framework for building applications on top of the model. However, the inaccessibility of the GitHub page at the time of reporting is a red flag that could delay adoption. Developers will need to wait for the harness to be fully available and stable before they can start building production-ready applications.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

Developers should monitor the availability and stability of the Qwen-Live Harness, as the source notes its GitHub page returned a 404 error. Additionally, independent benchmarks comparing Qwen3.8-Omni-Flash against other frontier models like Gemini 3.8 Flash will be crucial to validate Alibaba's performance claims. The practical impact on API pricing for multimodal tasks will also be a key factor for enterprises considering adoption.

The availability and stability of the Qwen-Live Harness will be a key factor in the model's adoption. If the GitHub page remains inaccessible or if the harness has significant bugs, it could hinder developers from building applications on top of Qwen3.8-Omni-Flash. Monitoring the GitHub repository and any updates from Alibaba will be essential.

Independent benchmarks comparing Qwen3.8-Omni-Flash against other frontier models, such as Gemini 3.8 Flash and GPT-4o, will be crucial to validate Alibaba's performance claims. Third-party evaluations will provide a more objective view of the model's capabilities and help enterprises make informed decisions about adoption.

The practical impact on API pricing for multimodal tasks will also be a key factor. If the cost reductions claimed by Alibaba are realized, it could lead to a significant shift in the market, making multimodal AI more accessible to smaller companies and developers. Monitoring the actual API costs and comparing them to competitors will be important.

The model's performance in real-world scenarios, such as video editing and real-time conversation, will be a key indicator of its practical utility. User feedback and case studies from early adopters will provide valuable insights into the model's strengths and limitations.

관련 가이드 및 퀴즈

AI 모델 설명AI 에이전트AI란 무엇인가?알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?