Komawa Labarai
SamfuraAI Understanding takaitaccen bayani

Alibaba ya saki Qwen3.8-Omni-Flash don ayyuka na wakili na multimodal

Alibaba ya saki Qwen3.8-Omni-Flash, samfurin omnimodal na asali wanda ke tallafawa alamomin mahallin miliyan 1, wanda ta yi iƙirarin samun aikin sauti da bidiyo kwatankwacin Gemini 3.8 Flash a ƙananan farashin API.

5 min readRead the linked source
Source-provided image accompanying Alibaba releases Qwen3.8-Omni-Flash for multimodal agent tasks
Tushen tusheAn rubuta tushen tushe
Mawallafi
gigazine.net
Tushen hanyar haɗin gwiwa
gigazine.nethttps://gigazine.net/gsc_news/en/20260918-qwen-3-8-omni-flash/
Nau'in tushe
Tushen da aka haɗa - ba a kafa matsayin tushen farko ba.
MaganaFahimtar wannan a cikin daƙiƙa 60

Fara a nan

Mabuɗin sharuddan

API (Tsarin Tsare-tsare na Aikace-aikacen)
Hanyar da aka tsara don tsarin software ɗaya don aika buƙatun zuwa da karɓar amsa daga wani tsarin.
Tagar yanayi
Matsakaicin adadin shigarwar alamun da samfurin harshe zai iya aiwatarwa a lokaci ɗaya.
Alamar alama
Daidaitaccen gwaji ko saitin bayanai da aka yi amfani da shi don aunawa da kwatanta aikin ƙira.
Gwada kankaAI Model An Bayyana Tambayoyi

Me ya faru

Alibaba released Qwen3.8-Omni-Flash, a new model in its Qwen series designed for omnimodal processing. The model supports a 1 million token and targets workflows such as video editing, music video production, and real-time conversation. According to Gigazine, the model outperforms its predecessor, Qwen3.5-Omni-Plus, by over 25% on average across 29 benchmarks. Alibaba claims the new model achieves audio and video performance close to or surpassing Gemini 3.8 Flash, while reducing API costs for voice input by over 98% and voice/video input by over 93% compared to the previous version. The release includes extensions to Qwen-MM-Plugins and an open-source harness called Qwen-Live Harness, though the latter's GitHub page was inaccessible at the time of reporting.

Alibaba has added Qwen3.8-Omni-Flash to its Qwen series, describing it as a next-generation native omnimodal model. The model is designed to handle a wide range of workflows, including video editing, music video production, film production, audiovisual summarization, and real-time conversation. It supports a of 1 million tokens, which is a substantial increase that allows for the processing of long-form audio and video files.

According to Gigazine, the model delivers superior results compared to its predecessor, Qwen3.5-Omni-Plus. Across 29 benchmarks, the average score is reported to be more than 25% higher. Specific improvements are noted in core functions such as understanding long audio, audio/video inference, captioning, and multi-speaker recognition. For instance, it achieved a score of 82.7 in LongAudioSpan and 89.7 in AliMeeting, a for transcribing multi-person, multi-channel conference audio in Chinese.

Alibaba claims that Qwen3.8-Omni-Flash scales data, context, and agent environments to achieve audio and video performance close to that of Gemini 3.8 Flash, with overall audio performance surpassing it. The company highlights significant cost reductions, stating that the API cost per hour for voice input is over 98% lower, and for voice and video input, it is over 93% lower than the previous model. These cost claims are central to the model's value proposition for enterprise users.

To support the model's capabilities, Alibaba has extended Qwen-MM-Plugins to add on-demand perception, tool usage, and workflow execution capabilities for long audio and video. The company also announced the open-sourcing of Qwen-Live Harness, a comprehensive harness designed based on the Qwen3.8-Omni-Flash-Realtime API. However, Gigazine notes that at the time of writing, the Qwen-Live Harness GitHub page was not accessible and returned a 404 error, indicating potential delays or issues with the public release of this component.

Bayanan tushe: gigazine.net ↗

Me ya sa yake da mahimmanci

This release is significant because it addresses the high cost and complexity of processing long-form audio and video data in AI agents. By claiming near-frontier performance at a fraction of the cost of previous models, Alibaba is positioning Qwen3.8-Omni-Flash as a practical tool for enterprise applications that require real-time multimodal reasoning. The shift from treating audio and video as simple perceptual inputs to core media for agent reasoning could accelerate the integration of AI into real-world productivity scenarios, such as automated media production and complex conversational agents. However, the specific comparisons to Gemini 3.8 Flash are based on Alibaba's claims and have not been independently verified by third-party evaluators.

The release of Qwen3.8-Omni-Flash is significant for the AI industry because it targets a specific pain point: the high cost and technical complexity of processing long-form multimodal data. By claiming near-frontier performance at a drastically reduced cost, Alibaba is challenging the status quo of multimodal AI pricing and accessibility. This could lead to a broader adoption of AI agents in industries that rely heavily on audio and video data, such as media production, customer service, and real-time translation.

The model's ability to handle 1 million tokens of context is a major technical advancement, as it allows for the processing of much longer audio and video files without segmentation. This is crucial for applications like film production or long-form meeting transcription, where context continuity is essential. The improvement in multi-speaker recognition and long audio understanding suggests that the model is better suited for complex, real-world scenarios than previous iterations.

However, the claims of outperforming or matching Gemini 3.8 Flash are based on Alibaba's internal benchmarks and have not been independently verified. This is a common issue in the AI industry, where companies often rely on self-reported metrics. Independent evaluations will be necessary to confirm the true performance and cost-effectiveness of the model. Until then, enterprises should approach the claims with caution and conduct their own testing before making significant investments.

The open-sourcing of Qwen-Live Harness is a positive step for the developer community, as it provides a framework for building applications on top of the model. However, the inaccessibility of the GitHub page at the time of reporting is a red flag that could delay adoption. Developers will need to wait for the harness to be fully available and stable before they can start building production-ready applications.

Interactive Mechanism

Ingantacciyar hanyar sadarwa: Yadda A zahiri yake Aiki

Bincika fasahar da ke bayan wannan ci gaban ta hanyar mu'amala.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Duba ra'ayi na hulɗa+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Abin kallo na gaba

Developers should monitor the availability and stability of the Qwen-Live Harness, as the source notes its GitHub page returned a 404 error. Additionally, independent benchmarks comparing Qwen3.8-Omni-Flash against other frontier models like Gemini 3.8 Flash will be crucial to validate Alibaba's performance claims. The practical impact on API pricing for multimodal tasks will also be a key factor for enterprises considering adoption.

The availability and stability of the Qwen-Live Harness will be a key factor in the model's adoption. If the GitHub page remains inaccessible or if the harness has significant bugs, it could hinder developers from building applications on top of Qwen3.8-Omni-Flash. Monitoring the GitHub repository and any updates from Alibaba will be essential.

Independent benchmarks comparing Qwen3.8-Omni-Flash against other frontier models, such as Gemini 3.8 Flash and GPT-4o, will be crucial to validate Alibaba's performance claims. Third-party evaluations will provide a more objective view of the model's capabilities and help enterprises make informed decisions about adoption.

The practical impact on API pricing for multimodal tasks will also be a key factor. If the cost reductions claimed by Alibaba are realized, it could lead to a significant shift in the market, making multimodal AI more accessible to smaller companies and developers. Monitoring the actual API costs and comparing them to competitors will be important.

The model's performance in real-world scenarios, such as video editing and real-time conversation, will be a key indicator of its practical utility. User feedback and case studies from early adopters will provide valuable insights into the model's strengths and limitations.

Jagorori masu alaƙa & tambayoyin tambayoyi

AI Model ya bayyanaWakilan AIMenene AI?Gwada abin da kuka sani - gwada gwajin AI kyautaNemo kalmar AI a cikin ƙamus ɗin muBi samfurin AI na sakin tracker
An sami wannan yana da amfani?