返回新聞
產品展示AI Understanding 簡報

Sarvam launches Vision 2.1 for 22 Indian languages

Sarvam AI has released Sarvam Vision 2.1, a document understanding model capable of extracting structured data from forms, tables, and handwritten text in English and 22 Indian languages.

5 min readRead the linked source
Source-provided image accompanying Sarvam launches Vision 2.1 for 22 Indian languages
來源參考來源記錄
出版商
cnbctv18.com
來源連結
cnbctv18.comhttps://www.cnbctv18.com/technology/sarvam-launches-new-ai-model-to-read-documents-in-22-indian-languages-19998111.htm
來源類型
連結來源-主要來源狀態尚未確定。
背景60 秒內了解這一點

從這裡開始

關鍵術語

API(應用程式介面)
一種軟體系統向另一個系統發送請求並接收回應的結構化方式。
OCR(光學字元辨識)
將圖像或掃描中的文字轉換為機器可讀文字的技術。
視覺語言模型 (VLM)
聯合處理視覺和文字訊息的多模態模型。
測試一下自己AI 模型解釋測驗

發生了什麼事

Sarvam AI launched Sarvam Vision 2.1, an updated AI model designed to read and understand documents in English and 22 Indian languages. The release focuses on extracting structured digital data from complex sources, including multi-page tables, forms, and handwritten text. According to CNBC TV18, the company reported that Vision 2.1 achieved a score of 87.3 on the olmOCR-Bench and 87.39 on its own Indic benchmark, which Sarvam described as state-of-the-art performance. The model was trained using a mix of real-world and synthetic data, followed by supervised fine-tuning and reinforcement learning with verifiable rewards to reduce hallucinations and inconsistent results identified in user feedback for the previous version. Sarvam also stated that it has optimized the underlying technology to lower the cost of running the model and has optimized its API for large-scale commercial use.

Sarvam AI, an Indian AI company, has launched Sarvam Vision 2.1, a new model specifically designed to read and understand documents in English and 22 Indian languages. The primary function of the model is to extract information from various document types, including forms, tables, and handwritten text, converting them into structured digital data. This release follows the initial launch of Sarvam Vision in February 2026, which focused on basic text reading and table understanding.

According to CNBC TV18, Sarvam reported that Vision 2.1 scored 87.3 on the olmOCR-Bench and 87.39 on its proprietary Indic benchmark. The company characterized these results as state-of-the-art performance. The new version is designed to handle more complex documents than its predecessor, such as tables spanning multiple pages and forms requiring extraction into specific fields. A key new capability is the recognition of handwritten text in Indian languages, aimed at digitizing documents that are difficult for conventional text-reading software to process.

The company stated that user feedback on the first version indicated issues with hallucinations and inconsistent results. To address this, Sarvam trained Vision 2.1 using a combination of real-world and synthetic data, including handwritten and printed forms, artificially filled-in forms from the web, and handwritten content from videos. The training process included supervised fine-tuning and reinforcement learning using verifiable rewards, a technique intended to improve performance based on checkable outputs.

Sarvam also announced improvements to the technology used to run the model, allowing it to be offered at a lower price than initially planned. The company's API, which enables integration into other software, has been optimized for large-scale commercial use. The model utilizes a combination of a vision-language model and additional systems to understand document layout and reading order, which the company claims improves accuracy over processing entire pages independently.

來源詳情: cnbctv18.com ↗

為什麼這很重要

This launch addresses a significant gap in AI capabilities for non-English, low-resource languages, specifically within the Indian market. By enabling accurate extraction of structured data from handwritten and complex documents in 22 languages, the model facilitates the digitization of records that conventional software cannot process. This has practical implications for sectors such as banking, government services, and logistics in India, where document-heavy workflows often rely on manual data entry. The reported reduction in operational costs and the focus on reducing hallucinations suggest a move toward more reliable, enterprise-grade deployment for multilingual document processing.

The launch of Sarvam Vision 2.1 is significant because it targets a specific and high-volume need in the Indian market: the digitization of multilingual documents, particularly those containing handwritten text. Many administrative, financial, and legal records in India are maintained in regional languages and often in handwritten form, making them inaccessible to standard OCR tools that are predominantly optimized for English and printed text.

By supporting 22 Indian languages, the model has the potential to streamline workflows in sectors such as banking, insurance, and government services, where manual data entry is a bottleneck. The reported reduction in hallucinations and inconsistent results is crucial for enterprise adoption, as reliability is a primary concern for automated document processing. The use of reinforcement learning with verifiable rewards suggests a methodological shift toward more robust and accurate outputs.

The optimization for lower running costs and large-scale API use indicates a commercial strategy focused on accessibility and integration. This could lower the barrier to entry for smaller businesses and public sector entities looking to digitize their records. The focus on structured data extraction rather than just text recognition aligns with practical business needs for data analysis and automation.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
互動式概念檢查+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

接下來看什麼

Independent verification of the reported benchmark scores and real-world accuracy in diverse document scenarios. The actual pricing structure and access conditions for the API, which are currently not detailed in the source. Adoption rates among Indian enterprises and government bodies for digitizing legacy handwritten records. Continued performance tracking on saturated benchmarks like OmniDocBench as the company shifts focus to real-world workflow testing.

Independent verification of the benchmark scores is necessary, as the reported metrics are self-reported by Sarvam. Third-party evaluations on standardized tests like olmOCR-Bench and OmniDocBench will provide a more objective measure of the model's performance relative to competitors.

The specific pricing and access conditions for the Sarvam Vision 2.1 API are not detailed in the source. Clarification on whether the model is available via public cloud, on-premise, or hybrid deployment, and the associated costs, will be important for potential users.

Real-world performance in diverse document scenarios, particularly with varied handwriting styles and document layouts, will determine the model's practical utility. The company's plan to track performance on real-world document workflows is a positive step, but independent case studies will be needed to confirm its effectiveness.

Adoption by major Indian enterprises and government agencies will be a key indicator of the model's impact. Partnerships with banks, insurance companies, and public sector units could accelerate the digitization of legacy records and demonstrate the model's scalability.

相關指引和測驗

人工智慧模型解釋什麼是人工智慧?AI 的未來測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語
覺得有用嗎?