What happened
Indian AI startup Sarvam has released Vision 2.1, an updated model designed to process complex document types including forms, multi-page tables, and handwritten text. The model supports English and 22 Indian languages, aiming to convert unstructured physical or digital documents into structured data for downstream software applications.
Sarvam Vision 2.1 is an evolution of the company's first-generation vision model launched in February 2026. The new iteration is specifically engineered to handle complex document structures, such as forms requiring field-level identification and tables that span multiple pages.
The model supports English and 22 Indian languages. Sarvam claims the model can recognize both printed and handwritten text, a feature intended to assist in the digitization of records that are currently difficult to process using standard OCR technology.
To improve reliability, Sarvam utilized a training regimen involving real and simulated data, followed by supervised fine-tuning and with verifiable rewards. This approach is intended to minimize hallucinations and ensure consistent data extraction.
Sarvam reported performance scores of 87.3 on the olmOCR-Bench and 87.39 on its proprietary Indic benchmark, characterizing these results as state-of-the-art for document understanding tasks in the region.
Why it matters
Vision 2.1 addresses significant challenges in digitizing documents within India's diverse linguistic landscape, where traditional optical character recognition (OCR) often struggles with handwriting variations and complex layouts. By focusing on reducing hallucinations and improving extraction consistency, Sarvam aims to provide a more reliable tool for organizations managing large volumes of records. The model's ability to handle multi-page tables and field-specific form extraction represents a practical advancement for automated data entry and digital transformation workflows in the region.
The primary value of Vision 2.1 lies in its ability to bridge the gap between physical, multilingual records and structured digital data. Many organizations in India rely on handwritten forms and complex paper-based documentation that are not easily machine-readable.
By reducing the error rates and hallucinations noted by users of the previous version, Sarvam is attempting to make AI-driven document processing viable for sensitive workflows where data accuracy is critical.
The model's support for 22 Indian languages is a strategic focus for the company, which aims to build AI infrastructure tailored to India's specific multilingual requirements rather than relying solely on models trained primarily on English-language datasets.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
crm_get_transaction(id='4092').In AI, what are a model's "parameters"?
What to watch next
Users should monitor the practical deployment of Vision 2.1 in real-world enterprise environments to see if the reported benchmark scores translate to consistent performance across varied document qualities. It remains unknown how Sarvam plans to offer access to this model, including specific pricing models, API availability, or whether it will be integrated into existing enterprise platforms. Future updates may clarify the extent of its 'state-of-the-art' claims compared to broader global document-processing models.
The company has not disclosed specific details regarding the commercial availability, pricing, or technical access requirements for Vision 2.1.
Independent verification of the model's performance on diverse, real-world document samples—beyond the reported benchmark scores—will be necessary to determine its utility in professional settings.
Future developments may include integration capabilities with existing enterprise software suites, which would be a key indicator of the model's practical adoption potential.