Ku laabo Warka
AlaabtaAI Understanding warbixin kooban

Alibaba waxay sii daysay Qwen3.8-Omni-Flash ee hawlaha wakiilka qaab-dhismeedka badan

Alibaba waxa ay sii daysay Qwen3.8-Omni-Flash, oo ah moodal omnimodal ah oo taageera 1 milyan oo calaamado macnaha guud ah, kaas oo ay ku andacoonayso in uu gaadho waxqabadka maqalka iyo muuqaalka ah ee u dhigma Gemini 3.8 Flash oo si weyn hoos ugu dhacay kharashaadka API.

5 min readRead the linked source
Source-provided image accompanying Alibaba releases Qwen3.8-Omni-Flash for multimodal agent tasks
Xigasho SourceIsha la duubay
Daabacaha
gigazine.net
Xidhiidhka isha
gigazine.nethttps://gigazine.net/gsc_news/en/20260918-qwen-3-8-omni-flash/
Nooca isha
Isha ku xidhan — heerka isha aasaasiga ah lama damin.
Dulucda sheekadaKu fahan tan 60 ilbiriqsi gudahood

Halkan ka bilow

Qodobbada muhiimka ah

API (Interface Programming Interface)
Habka habaysan ee hal nidaam software si uu codsiyada ugu diro ugana helo jawaabaha nidaam kale.
Daaqadda macnaha
Qadarka ugu badan ee calaamadaha gelinta ee qaabka luqadda ayaa socodsiin kara hal mar.
Benchmark
Tijaabo la habeeyey ama kayd xogeed oo loo isticmaalo in lagu cabbiro laguna barbar dhigo waxqabadka moodeelka.
Is tijaabiMoodooyinka AI Kedis La Sharaxay

Maxaa dhacay

Alibaba released Qwen3.8-Omni-Flash, a new model in its Qwen series designed for omnimodal processing. The model supports a 1 million token and targets workflows such as video editing, music video production, and real-time conversation. According to Gigazine, the model outperforms its predecessor, Qwen3.5-Omni-Plus, by over 25% on average across 29 benchmarks. Alibaba claims the new model achieves audio and video performance close to or surpassing Gemini 3.8 Flash, while reducing API costs for voice input by over 98% and voice/video input by over 93% compared to the previous version. The release includes extensions to Qwen-MM-Plugins and an open-source harness called Qwen-Live Harness, though the latter's GitHub page was inaccessible at the time of reporting.

Alibaba has added Qwen3.8-Omni-Flash to its Qwen series, describing it as a next-generation native omnimodal model. The model is designed to handle a wide range of workflows, including video editing, music video production, film production, audiovisual summarization, and real-time conversation. It supports a of 1 million tokens, which is a substantial increase that allows for the processing of long-form audio and video files.

According to Gigazine, the model delivers superior results compared to its predecessor, Qwen3.5-Omni-Plus. Across 29 benchmarks, the average score is reported to be more than 25% higher. Specific improvements are noted in core functions such as understanding long audio, audio/video inference, captioning, and multi-speaker recognition. For instance, it achieved a score of 82.7 in LongAudioSpan and 89.7 in AliMeeting, a for transcribing multi-person, multi-channel conference audio in Chinese.

Alibaba claims that Qwen3.8-Omni-Flash scales data, context, and agent environments to achieve audio and video performance close to that of Gemini 3.8 Flash, with overall audio performance surpassing it. The company highlights significant cost reductions, stating that the API cost per hour for voice input is over 98% lower, and for voice and video input, it is over 93% lower than the previous model. These cost claims are central to the model's value proposition for enterprise users.

To support the model's capabilities, Alibaba has extended Qwen-MM-Plugins to add on-demand perception, tool usage, and workflow execution capabilities for long audio and video. The company also announced the open-sourcing of Qwen-Live Harness, a comprehensive harness designed based on the Qwen3.8-Omni-Flash-Realtime API. However, Gigazine notes that at the time of writing, the Qwen-Live Harness GitHub page was not accessible and returned a 404 error, indicating potential delays or issues with the public release of this component.

Faahfaahinta isha: gigazine.net ↗

Maxay muhiim u tahay

This release is significant because it addresses the high cost and complexity of processing long-form audio and video data in AI agents. By claiming near-frontier performance at a fraction of the cost of previous models, Alibaba is positioning Qwen3.8-Omni-Flash as a practical tool for enterprise applications that require real-time multimodal reasoning. The shift from treating audio and video as simple perceptual inputs to core media for agent reasoning could accelerate the integration of AI into real-world productivity scenarios, such as automated media production and complex conversational agents. However, the specific comparisons to Gemini 3.8 Flash are based on Alibaba's claims and have not been independently verified by third-party evaluators.

The release of Qwen3.8-Omni-Flash is significant for the AI industry because it targets a specific pain point: the high cost and technical complexity of processing long-form multimodal data. By claiming near-frontier performance at a drastically reduced cost, Alibaba is challenging the status quo of multimodal AI pricing and accessibility. This could lead to a broader adoption of AI agents in industries that rely heavily on audio and video data, such as media production, customer service, and real-time translation.

The model's ability to handle 1 million tokens of context is a major technical advancement, as it allows for the processing of much longer audio and video files without segmentation. This is crucial for applications like film production or long-form meeting transcription, where context continuity is essential. The improvement in multi-speaker recognition and long audio understanding suggests that the model is better suited for complex, real-world scenarios than previous iterations.

However, the claims of outperforming or matching Gemini 3.8 Flash are based on Alibaba's internal benchmarks and have not been independently verified. This is a common issue in the AI industry, where companies often rely on self-reported metrics. Independent evaluations will be necessary to confirm the true performance and cost-effectiveness of the model. Until then, enterprises should approach the claims with caution and conduct their own testing before making significant investments.

The open-sourcing of Qwen-Live Harness is a positive step for the developer community, as it provides a framework for building applications on top of the model. However, the inaccessibility of the GitHub page at the time of reporting is a red flag that could delay adoption. Developers will need to wait for the harness to be fully available and stable before they can start building production-ready applications.

Interactive Mechanism

Farsamaynta Is-dhexgalka: Sida Dhabta Ay U Shaqeyso

U baadh tignoolajiyada hoose ee ka dambeeya horumarkan si isdhexgal leh.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Hubinta Fikradda Is-dhexgalka+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Maxaa la daawan doona xiga

Developers should monitor the availability and stability of the Qwen-Live Harness, as the source notes its GitHub page returned a 404 error. Additionally, independent benchmarks comparing Qwen3.8-Omni-Flash against other frontier models like Gemini 3.8 Flash will be crucial to validate Alibaba's performance claims. The practical impact on API pricing for multimodal tasks will also be a key factor for enterprises considering adoption.

The availability and stability of the Qwen-Live Harness will be a key factor in the model's adoption. If the GitHub page remains inaccessible or if the harness has significant bugs, it could hinder developers from building applications on top of Qwen3.8-Omni-Flash. Monitoring the GitHub repository and any updates from Alibaba will be essential.

Independent benchmarks comparing Qwen3.8-Omni-Flash against other frontier models, such as Gemini 3.8 Flash and GPT-4o, will be crucial to validate Alibaba's performance claims. Third-party evaluations will provide a more objective view of the model's capabilities and help enterprises make informed decisions about adoption.

The practical impact on API pricing for multimodal tasks will also be a key factor. If the cost reductions claimed by Alibaba are realized, it could lead to a significant shift in the market, making multimodal AI more accessible to smaller companies and developers. Monitoring the actual API costs and comparing them to competitors will be important.

The model's performance in real-world scenarios, such as video editing and real-time conversation, will be a key indicator of its practical utility. User feedback and case studies from early adopters will provide valuable insights into the model's strengths and limitations.

Tilmaamaha la xidhiidha & su'aalaha

Moodooyinka AI ayaa la sharaxayWakiilada AIWaa maxay AI?Tijaabi waxaad taqaan - isku day kedis AI oo bilaash ahKa raadi erey AI qaamuuskeenaRaac qaabka AI raadraaca sii deynta
Tan faa'iido ma u heshay?