Back to News
ProductAI Understanding briefing

Sarvam AI launches Saaras V4 speech-to-text model for enterprise

Sarvam AI has released Saaras V4, a speech-to-text model supporting English and 22 Indian languages, featuring low latency and specific optimizations for noisy and telephony audio environments.

6 min readRead the linked source
Source-page capture accompanying Sarvam AI launches Saaras V4 speech-to-text model for enterprise
Source referenceSource recorded
Publisher
enterpriseai.economictimes.indiatimes.com
Source link
enterpriseai.economictimes.indiatimes.comhttps://enterpriseai.economictimes.indiatimes.com/news/industry/sarvam-ai-launches-saaras-v4-to-strengthen-enterprise-speech-applications/134501023
Source type
Linked source β€” primary-source status has not been established.
ContextUnderstand this in 60 seconds

Start here

Key terms

API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Large Language Model (LLM)
A language model trained on massive text corpora to generate and analyze text.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Test yourselfAI Models Explained Quiz

What happened

Sarvam AI launched Saaras V4, a new speech-to-text model designed for enterprise applications. The model supports English and 22 Indian languages and is built to handle accents, dialects, and mixed-language conversations. It utilizes a 3-billion-parameter hybrid state-space model decoder trained in-house. The system generates five types of output directly from audio, including normalized transcripts, verbatim transcripts, code-mixed text, transliteration, and English translation. Sarvam reported that the model achieved the lowest word error rate across seven global English speech datasets and specific error rates on the IndicVoices dataset. The launch follows the recent release of Sarvam's Vision 2.1 document processing model.

Sarvam AI, an Indian AI startup, announced the launch of Saaras V4, its latest speech-to-text model. The company stated in a post on X that the model was built with specific attention to noise, accents, dialects, and mixed-language speech, describing it as its most capable speech-to-text model yet. The model is designed to support English and 22 Indian languages, targeting enterprise workflows such as customer service calls, voice agents, and telephony where speech quality varies.

Technically, Saaras V4 combines a speech audio encoder with a 3-billion-parameter hybrid state-space model decoder. Sarvam indicated that this decoder was trained from scratch in-house. The model is capable of generating five distinct types of output directly from audio input without the need for a separate large language model. These outputs include a normalized transcript, a verbatim transcript, code-mixed text, transliteration of Indic speech into the Latin script, and an English translation.

Sarvam reported performance metrics indicating that Saaras V4 achieved the lowest word error rate across seven global English speech datasets, which cover six international accents and an Indian English . For Indian languages, the company reported an error rate of 2.9 percent across the top 10 Indian languages and 5.22 percent across all 22 official Indian languages on the IndicVoices dataset. The model also supports automatic language identification.

The system includes a feature called keyterm prompting, which allows developers to provide up to 50 specific words or phrases, such as brand names or technical terminology, to improve recognition accuracy for those terms. Sarvam stated that the model is optimized for real-world audio conditions, including noisy and compressed recordings, as well as 8kHz telephony audio. The company reported a streaming time-to-first-token of below 150 milliseconds and the ability to process long-form recordings in less than a second per multi-minute recording.

This launch follows the recent release of Sarvam's Vision 2.1 model, which focuses on document intelligence. Vision 2.1 processes complex multi-page tables, forms, and handwritten documents across 22 official Indian languages and English. Sarvam has priced its document digitization API at 1.50 Indian Rupees per page. The combined release of Saaras V4 and Vision 2.1 reflects the company's broader focus on the data layer of enterprise AI, converting unstructured information into formats that downstream AI systems can process.

Source details: enterpriseai.economictimes.indiatimes.com β†—

Why it matters

This launch expands Sarvam AI's enterprise AI stack by providing a specialized tool for converting unstructured speech data into processable formats. The model's specific optimizations for noisy environments, 8kHz telephony audio, and mixed-language speech address practical challenges in customer service and voice agent workflows. By offering direct outputs like transliteration and translation without requiring a separate large language model, it simplifies integration for developers. The reported low latency and competitive error rates suggest a focus on production-ready performance for real-world Indian and global English speech scenarios.

The release of Saaras V4 is significant because it addresses specific technical challenges in speech recognition for the Indian market and global English accents. By handling mixed-language conversations and noisy environments, the model targets practical enterprise use cases where standard speech-to-text systems often fail. This positions Sarvam AI as a provider of specialized infrastructure for AI applications that rely on accurate speech data.

The ability to generate multiple output types, including transliteration and translation, directly from the audio encoder-decoder stack reduces the computational overhead and complexity for developers. This is a practical implication for enterprises looking to deploy voice agents or process customer service calls, as it streamlines the pipeline from raw audio to actionable text data without requiring additional large language model calls for basic transcription and translation tasks.

Sarvam's focus on the 'data layer' of enterprise AI suggests a strategic shift towards providing foundational tools that other AI systems can build upon. By improving the accuracy and speed of converting unstructured speech and document data, Sarvam aims to enhance the performance of downstream AI applications. This approach is relevant for organizations seeking to integrate AI into existing workflows that involve significant amounts of unstructured data.

The reported performance metrics, such as the lowest word error rate on specific benchmarks and low latency, are important for assessing the model's viability in production environments. However, these claims are based on Sarvam's own reporting and have not been independently verified by third-party benchmarks. The specific optimization for 8kHz telephony audio is particularly relevant for industries like banking and telecommunications, where call quality is often lower than in consumer-grade recordings.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:πŸ›‘οΈ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language modelβ€”it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interactive Concept Check+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

What to watch next

Developers should monitor the availability of the Saaras V4 API and specific pricing structures, which are not detailed in the source. Independent verification of the reported word error rates and latency metrics on diverse, real-world datasets will be crucial for assessing its actual performance. Additionally, the integration of Saaras V4 with Sarvam's Vision 2.1 model to create a unified data layer for both speech and document processing is a key development to track.

The source does not provide specific pricing or access details for the Saaras V4 API. While the Vision 2.1 API is priced at 1.50 Indian Rupees per page, the cost structure for speech-to-text services, including per-minute or per-token pricing, is not documented. Developers and enterprises should monitor Sarvam's official channels for API availability and pricing information.

Independent verification of the reported word error rates and latency metrics is necessary to confirm the model's performance in real-world scenarios. The benchmarks cited by Sarvam, such as the IndicVoices dataset and global English speech datasets, should be cross-referenced with third-party evaluations to ensure the claims are accurate and representative of diverse speech patterns.

The integration of Saaras V4 with Sarvam's Vision 2.1 model is a key area to watch. The company's strategy of building a comprehensive data layer for enterprise AI could lead to new products or services that combine speech and document processing. This could have significant implications for industries that rely on both types of unstructured data, such as insurance, healthcare, and finance.

The competitive landscape for speech-to-text models in the Indian market is evolving. The launch of Saaras V4 adds to the number of specialized models available for Indian languages and mixed-language speech. Monitoring how this model performs against competitors in terms of accuracy, latency, and cost will be important for enterprises making technology decisions.

Related guides & quizzes

AI Models ExplainedAI AgentsWhat is AI?Test what you know β€” try a free AI quizLook up an AI term in our glossaryFollow the AI model release tracker
Found this useful?