Back to News
ProductAI Understanding briefing

Alibaba releases Qwen3.8-Omni-Flash for multimodal agent tasks

Alibaba has released Qwen3.8-Omni-Flash, a native omnimodal model supporting 1 million tokens of context, which it claims achieves audio and video performance comparable to Gemini 3.8 Flash at significantly lower API costs.

5 min readRead the linked source
Source-provided image accompanying Alibaba releases Qwen3.8-Omni-Flash for multimodal agent tasks
Source referenceSource recorded
Publisher
gigazine.net
Source link
gigazine.nethttps://gigazine.net/gsc_news/en/20260918-qwen-3-8-omni-flash/
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Start here

Key terms

API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Context Window
The maximum amount of input tokens a language model can process at once.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Test yourselfAI Models Explained Quiz

What happened

Alibaba released Qwen3.8-Omni-Flash, a new model in its Qwen series designed for omnimodal processing. The model supports a 1 million token and targets workflows such as video editing, music video production, and real-time conversation. According to Gigazine, the model outperforms its predecessor, Qwen3.5-Omni-Plus, by over 25% on average across 29 benchmarks. Alibaba claims the new model achieves audio and video performance close to or surpassing Gemini 3.8 Flash, while reducing API costs for voice input by over 98% and voice/video input by over 93% compared to the previous version. The release includes extensions to Qwen-MM-Plugins and an open-source harness called Qwen-Live Harness, though the latter's GitHub page was inaccessible at the time of reporting.

Alibaba has added Qwen3.8-Omni-Flash to its Qwen series, describing it as a next-generation native omnimodal model. The model is designed to handle a wide range of workflows, including video editing, music video production, film production, audiovisual summarization, and real-time conversation. It supports a of 1 million tokens, which is a substantial increase that allows for the processing of long-form audio and video files.

According to Gigazine, the model delivers superior results compared to its predecessor, Qwen3.5-Omni-Plus. Across 29 benchmarks, the average score is reported to be more than 25% higher. Specific improvements are noted in core functions such as understanding long audio, audio/video inference, captioning, and multi-speaker recognition. For instance, it achieved a score of 82.7 in LongAudioSpan and 89.7 in AliMeeting, a for transcribing multi-person, multi-channel conference audio in Chinese.

Alibaba claims that Qwen3.8-Omni-Flash scales data, context, and agent environments to achieve audio and video performance close to that of Gemini 3.8 Flash, with overall audio performance surpassing it. The company highlights significant cost reductions, stating that the API cost per hour for voice input is over 98% lower, and for voice and video input, it is over 93% lower than the previous model. These cost claims are central to the model's value proposition for enterprise users.

To support the model's capabilities, Alibaba has extended Qwen-MM-Plugins to add on-demand perception, tool usage, and workflow execution capabilities for long audio and video. The company also announced the open-sourcing of Qwen-Live Harness, a comprehensive harness designed based on the Qwen3.8-Omni-Flash-Realtime API. However, Gigazine notes that at the time of writing, the Qwen-Live Harness GitHub page was not accessible and returned a 404 error, indicating potential delays or issues with the public release of this component.

Source details: gigazine.net

Why it matters

This release is significant because it addresses the high cost and complexity of processing long-form audio and video data in AI agents. By claiming near-frontier performance at a fraction of the cost of previous models, Alibaba is positioning Qwen3.8-Omni-Flash as a practical tool for enterprise applications that require real-time multimodal reasoning. The shift from treating audio and video as simple perceptual inputs to core media for agent reasoning could accelerate the integration of AI into real-world productivity scenarios, such as automated media production and complex conversational agents. However, the specific comparisons to Gemini 3.8 Flash are based on Alibaba's claims and have not been independently verified by third-party evaluators.

The release of Qwen3.8-Omni-Flash is significant for the AI industry because it targets a specific pain point: the high cost and technical complexity of processing long-form multimodal data. By claiming near-frontier performance at a drastically reduced cost, Alibaba is challenging the status quo of multimodal AI pricing and accessibility. This could lead to a broader adoption of AI agents in industries that rely heavily on audio and video data, such as media production, customer service, and real-time translation.

The model's ability to handle 1 million tokens of context is a major technical advancement, as it allows for the processing of much longer audio and video files without segmentation. This is crucial for applications like film production or long-form meeting transcription, where context continuity is essential. The improvement in multi-speaker recognition and long audio understanding suggests that the model is better suited for complex, real-world scenarios than previous iterations.

However, the claims of outperforming or matching Gemini 3.8 Flash are based on Alibaba's internal benchmarks and have not been independently verified. This is a common issue in the AI industry, where companies often rely on self-reported metrics. Independent evaluations will be necessary to confirm the true performance and cost-effectiveness of the model. Until then, enterprises should approach the claims with caution and conduct their own testing before making significant investments.

The open-sourcing of Qwen-Live Harness is a positive step for the developer community, as it provides a framework for building applications on top of the model. However, the inaccessibility of the GitHub page at the time of reporting is a red flag that could delay adoption. Developers will need to wait for the harness to be fully available and stable before they can start building production-ready applications.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interactive Concept Check+10 Points
AI Models Explained Quiz

What is the best response when AI Models Explained makes a mistake in production?

What to watch next

Developers should monitor the availability and stability of the Qwen-Live Harness, as the source notes its GitHub page returned a 404 error. Additionally, independent benchmarks comparing Qwen3.8-Omni-Flash against other frontier models like Gemini 3.8 Flash will be crucial to validate Alibaba's performance claims. The practical impact on API pricing for multimodal tasks will also be a key factor for enterprises considering adoption.

The availability and stability of the Qwen-Live Harness will be a key factor in the model's adoption. If the GitHub page remains inaccessible or if the harness has significant bugs, it could hinder developers from building applications on top of Qwen3.8-Omni-Flash. Monitoring the GitHub repository and any updates from Alibaba will be essential.

Independent benchmarks comparing Qwen3.8-Omni-Flash against other frontier models, such as Gemini 3.8 Flash and GPT-4o, will be crucial to validate Alibaba's performance claims. Third-party evaluations will provide a more objective view of the model's capabilities and help enterprises make informed decisions about adoption.

The practical impact on API pricing for multimodal tasks will also be a key factor. If the cost reductions claimed by Alibaba are realized, it could lead to a significant shift in the market, making multimodal AI more accessible to smaller companies and developers. Monitoring the actual API costs and comparing them to competitors will be important.

The model's performance in real-world scenarios, such as video editing and real-time conversation, will be a key indicator of its practical utility. User feedback and case studies from early adopters will provide valuable insights into the model's strengths and limitations.

Related guides & quizzes

AI Models ExplainedAI AgentsWhat is AI?Test what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?