Back to News
ProductAI Understanding briefing

Alibaba launches Qwen‑Audio 3.1 with up to 95% price cut on voice APIs

Alibaba’s Qwen‑Audio 3.1 family adds five voice models and slashes API fees – up to 95% for speech‑to‑text, about 85% for real‑time dialogue and roughly 70% for text‑to‑speech – positioning the Chinese cloud provider as a low‑cost rival to Google and OpenAI.

4 min readRead the linked source
Source-provided image accompanying Alibaba launches Qwen‑Audio 3.1 with up to 95% price cut on voice APIs
Source referenceSource recorded
Publisher
pasqualepillitteri.it
Source link
pasqualepillitteri.ithttps://pasqualepillitteri.it/en/news/18492/alibaba-qwen-audio-3-1-voice-api-price-cut
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Start here

Key terms

API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Large Language Model (LLM)
A language model trained on massive text corpora to generate and analyze text.
Inference
The runtime phase where a trained model generates predictions or outputs.
Test yourselfAI Models Explained Quiz

What happened

Alibaba introduced the Qwen‑Audio 3.1 suite, a five‑model voice‑AI package, and announced steep reductions in its cloud‑based voice API pricing.

Alibaba’s cloud division announced the Qwen‑Audio 3.1 family, comprising five distinct voice models that cover speech recognition with speaker identification (ASR‑Next), real‑time interactive dialogue, and text‑to‑speech synthesis with built‑in sound effects (TTS‑Next). The models are listed on the Qwen Cloud Model Studio platform under identifiers such as qwen-audio-3.1-asr-flash-filetrans.

Alongside the launch, Alibaba disclosed percentage‑based price reductions for its voice APIs: up to 95% for speech‑to‑text, roughly 85% for real‑time dialogue, and about 70% for text‑to‑speech. The source does not provide absolute dollar figures or the exact date when the new rates become effective.

The announcement follows recent Alibaba milestones, including the unveiling of the Qwen‑4 large language model and the Zhenwu V900 AI chip, suggesting that internal hardware efficiencies may be enabling the aggressive pricing strategy.

The Qwen‑Audio 3.1 suite claims support for 30 languages and 16 Chinese dialects in transcription, while synthesis coverage is limited to 16 languages. No open‑weight release is mentioned; the service remains a pay‑as‑you‑go cloud offering.

Source details: pasqualepillitteri.it ↗

Why it matters

Voice AI pricing is a decisive factor for enterprises that process large volumes of audio, such as call‑center transcription, meeting‑recording services, and voice‑assistant platforms. By cutting ASR fees up to 95% and offering comparable discounts for real‑time dialogue and TTS, Alibaba directly challenges the cost structures of Google’s Gemini and OpenAI’s voice offerings, potentially reshaping vendor selection for cost‑sensitive customers. The move also signals Alibaba’s broader strategy to leverage its own AI chips and recent large‑scale model releases to drive down costs, a tactic that could pressure competitors to revisit their pricing or accelerate feature rollouts. However, the actual impact hinges on the undisclosed baseline prices, language coverage, and quality of the models, especially for non‑Chinese languages.

Cost is the primary lever for enterprises that ingest hours of audio daily. A 95% reduction in ASR fees could make Alibaba the default provider for large‑scale transcription projects, provided the accuracy meets enterprise standards.

By bundling speaker recognition and pre‑mixed audio synthesis, Alibaba adds functional breadth that may reduce the need for multiple third‑party services, simplifying integration pipelines.

The price cuts could force Google and OpenAI to either lower their rates or accelerate feature differentiation, potentially benefiting end‑users through more competitive pricing or improved model capabilities.

Uncertainty remains around the baseline pricing, language‑specific quality, and latency. Without transparent benchmarks, customers must conduct their own comparative tests before committing.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Model Parameter Size:8B Parameters
VRAM Required5.5 GBGPU memory footprint
Target HardwareMacBook / Single GPUDeployment tier
Privacy100% Air-GappedLocal device capability
Core takeaway: Small, quantized models (3B–8B) now run directly inside smartphones and laptops with complete data privacy, while mammoth 400B+ models remain the domain of datacenter clusters.
Interactive Concept Check+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

What to watch next

Key indicators to monitor include the official dollar price list on Qwen Cloud, real‑world accuracy benchmarks across the 30 supported transcription languages, and any competitive pricing or feature responses from Google and OpenAI. Adoption rates in enterprise transcription and voice‑assistant markets will reveal whether the price cuts translate into measurable market share gains.

Release of the official price list in USD or RMB on the Qwen Cloud dashboard, which will allow precise cost‑per‑hour calculations.

Independent accuracy and latency benchmarks across the claimed language set, especially for English and other high‑volume languages.

Responses from Google (Gemini) and OpenAI (Voice) regarding pricing or feature updates, which could indicate a broader market shift.

Adoption metrics from enterprise customers, such as call‑center operators or meeting‑recording platforms, that will signal whether the price cuts translate into real‑world market share.

Related guides & quizzes

AI Models ExplainedAI AgentsFuture of AIAI EthicsTest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI model release tracker
Found this useful?