Back to News
ProductAI Understanding briefing

Sarvam AI unveils Vision 2.1 with expanded Indian‑language OCR and production‑ready inference

Sarvam AI announced Vision 2.1, a vision‑language model that adds structured data extraction, complex table handling and handwritten text recognition for 22 Indian languages, plus a new Indic OCR benchmark.

4 min readRead the linked source
Source-provided image accompanying Sarvam AI unveils Vision 2.1 with expanded Indian‑language OCR and production‑ready inference
Source referenceSource recorded
Publisher
storyboard18.com
Source link
storyboard18.comhttps://www.storyboard18.com/digital/sarvam-ai-launches-vision-2-1-to-boost-document-intelligence-indian-language-ocr-111459.htm
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Start here

Key terms

OCR (Optical Character Recognition)
Technology that converts text in images or scans into machine-readable text.
Inference
The runtime phase where a trained model generates predictions or outputs.
API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Test yourselfAI Models Explained Quiz

What happened

Sarvam AI launched Vision 2.1, an updated vision‑language model aimed at document intelligence. The model now supports key‑value extraction from forms, multi‑page table processing, and handwritten text recognition in India’s 22 official languages. Sarvam also released the Sarvam Indic OCR Bench, a 6,909‑sample benchmark covering newspapers, brochures, textbooks and historic texts, and made it publicly available on Hugging Face. The company reported an 87.39 % overall accuracy on this benchmark and claimed frontier performance on global benchmarks such as olmOCR‑Bench and OmniDocBench. Internally, the model was trained on synthetic and real‑world data, fine‑tuned with supervised learning, and further refined using reinforcement learning with verifiable rewards (RLVR). New architectural components—a semantic layout parser and a pointer reading‑order network—were added to improve document structure handling. In response to feedback on the prior version, Sarmam optimized the stack to lower serving costs and make the API more suitable for production workloads.

Sarvam AI’s Vision 2.1 model is positioned as a document‑intelligence engine capable of processing both individual pages and whole documents. It extracts key‑value pairs from forms, parses complex multi‑page tables, and recognises handwritten text in Indic scripts, a capability not widely available in existing OCR tools.

The company built its training data from a mix of synthetic generation and real‑world samples, then applied supervised fine‑tuning followed by reinforcement learning with verifiable rewards (RLVR). New components—a semantic layout parser and a pointer reading‑order network—were integrated to improve handling of document layout and reading order.

Alongside the model, Sarvam released the Sarvam Indic OCR Bench, containing 6,909 samples (6,609 in Indian languages, 300 in English) drawn from diverse sources spanning the 19th century to present day. The benchmark focuses on character‑ and word‑level accuracy, and the model achieved 87.39 % overall accuracy on it.

Sarvam claims the model also attains frontier results on global benchmarks such as olmOCR‑Bench and OmniDocBench, though independent verification is pending. The stack was re‑engineered to reduce serving costs and make the API more suitable for production workloads.

Source details: storyboard18.com ↗

Why it matters

Vision 2.1 addresses a critical gap in Indian‑language document automation, where existing OCR solutions often struggle with diverse scripts, complex tables, and handwritten content. By delivering higher accuracy across 22 official languages and offering a public benchmark, Sarvam provides a concrete baseline for developers and enterprises seeking to automate data extraction from regional documents. The production‑focused improvements lower the cost barrier for businesses that need to process large volumes of forms, invoices, or archival material, potentially accelerating digitisation efforts in sectors such as banking, government, and education. Moreover, the open‑source benchmark invites community validation, fostering transparency and encouraging further research on multilingual OCR. However, pricing, licensing terms, and exact availability of the Vision 2.1 API remain undisclosed, leaving enterprises to await further details before committing to deployment.

The ability to reliably extract structured data from Indian‑language documents can reduce manual data entry costs and error rates for businesses operating in multilingual environments, a significant operational advantage in markets like India where diverse scripts are common.

By publishing the Indic OCR Bench on Hugging Face, Sarvam encourages open evaluation, which can accelerate research and improve the robustness of OCR solutions for low‑resource languages.

Production‑grade optimisations lower the total cost of ownership, making large‑scale deployment financially viable for enterprises that need to process high volumes of forms, invoices, or archival records.

The model’s support for handwritten text recognition expands use cases to include legacy paperwork, field notes, and handwritten applications, areas traditionally difficult for automated systems.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interactive Concept Check+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

What to watch next

Future announcements from Sarvam regarding pricing, cloud‑hosting options, or on‑premise licensing for Vision 2.1; adoption metrics from early enterprise pilots; community results on the Indic OCR Bench that could reveal strengths or weaknesses of the model; and any follow‑up releases that expand language coverage or add new document‑processing features.

Pricing and licensing details for Vision 2.1, which will determine accessibility for startups versus large enterprises.

Performance results from independent third‑party evaluations on the Indic OCR Bench, which could confirm or challenge Sarvam’s reported accuracy.

Adoption case studies from sectors such as banking, government, or education that illustrate real‑world impact and ROI.

Potential extensions to additional Indian languages or scripts, and any future model iterations that further improve accuracy or reduce latency.

Related guides & quizzes

AI Models ExplainedAI TrainingFuture of AITest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI model release tracker
Found this useful?