Înapoi la Știri
ProdusAI Understanding briefing

Sarvam AI dezvăluie Vision 2.1 cu OCR extins în limba indiană și inferență pregătită pentru producție

Sarvam AI a anunțat Vision 2.1, un model de limbaj de viziune care adaugă extragerea de date structurată, gestionarea complexă a tabelelor și recunoașterea textului scris de mână pentru 22 de limbi indiene, plus un nou benchmark Indic OCR.

4 min readRead the linked source
Source-provided image accompanying Sarvam AI unveils Vision 2.1 with expanded Indian‑language OCR and production‑ready inference
Referință la sursăSursa înregistrată
Editor
storyboard18.com
Link sursă
storyboard18.comhttps://www.storyboard18.com/digital/sarvam-ai-launches-vision-2-1-to-boost-document-intelligence-indian-language-ocr-111459.htm
Tip sursă
Sursă conectată — starea sursei primare nu a fost stabilită.
ContextÎnțelege asta în 60 de secunde

Începeți de aici

Termeni cheie

OCR (recunoaștere optică a caracterelor)
Tehnologie care convertește textul în imagini sau scanează în text care poate fi citit de mașină.
Inferență
Faza de rulare în care un model antrenat generează predicții sau rezultate.
API (Interfață de programare a aplicației)
O modalitate structurată pentru ca un sistem software să trimită cereri și să primească răspunsuri de la un alt sistem.
Testează-teTest explicativ pentru modelele AI

Ce sa întâmplat

Sarvam AI launched Vision 2.1, an updated vision‑language model aimed at document intelligence. The model now supports key‑value extraction from forms, multi‑page table processing, and handwritten text recognition in India’s 22 official languages. Sarvam also released the Sarvam Indic OCR Bench, a 6,909‑sample benchmark covering newspapers, brochures, textbooks and historic texts, and made it publicly available on Hugging Face. The company reported an 87.39 % overall accuracy on this benchmark and claimed frontier performance on global benchmarks such as olmOCR‑Bench and OmniDocBench. Internally, the model was trained on synthetic and real‑world data, fine‑tuned with supervised learning, and further refined using reinforcement learning with verifiable rewards (RLVR). New architectural components—a semantic layout parser and a pointer reading‑order network—were added to improve document structure handling. In response to feedback on the prior version, Sarmam optimized the stack to lower serving costs and make the API more suitable for production workloads.

Sarvam AI’s Vision 2.1 model is positioned as a document‑intelligence engine capable of processing both individual pages and whole documents. It extracts key‑value pairs from forms, parses complex multi‑page tables, and recognises handwritten text in Indic scripts, a capability not widely available in existing OCR tools.

The company built its training data from a mix of synthetic generation and real‑world samples, then applied supervised fine‑tuning followed by reinforcement learning with verifiable rewards (RLVR). New components—a semantic layout parser and a pointer reading‑order network—were integrated to improve handling of document layout and reading order.

Alongside the model, Sarvam released the Sarvam Indic OCR Bench, containing 6,909 samples (6,609 in Indian languages, 300 in English) drawn from diverse sources spanning the 19th century to present day. The benchmark focuses on character‑ and word‑level accuracy, and the model achieved 87.39 % overall accuracy on it.

Sarvam claims the model also attains frontier results on global benchmarks such as olmOCR‑Bench and OmniDocBench, though independent verification is pending. The stack was re‑engineered to reduce serving costs and make the API more suitable for production workloads.

Detalii sursa: storyboard18.com ↗

De ce contează

Vision 2.1 addresses a critical gap in Indian‑language document automation, where existing OCR solutions often struggle with diverse scripts, complex tables, and handwritten content. By delivering higher accuracy across 22 official languages and offering a public benchmark, Sarvam provides a concrete baseline for developers and enterprises seeking to automate data extraction from regional documents. The production‑focused improvements lower the cost barrier for businesses that need to process large volumes of forms, invoices, or archival material, potentially accelerating digitisation efforts in sectors such as banking, government, and education. Moreover, the open‑source benchmark invites community validation, fostering transparency and encouraging further research on multilingual OCR. However, pricing, licensing terms, and exact availability of the Vision 2.1 API remain undisclosed, leaving enterprises to await further details before committing to deployment.

The ability to reliably extract structured data from Indian‑language documents can reduce manual data entry costs and error rates for businesses operating in multilingual environments, a significant operational advantage in markets like India where diverse scripts are common.

By publishing the Indic OCR Bench on Hugging Face, Sarvam encourages open evaluation, which can accelerate research and improve the robustness of OCR solutions for low‑resource languages.

Production‑grade optimisations lower the total cost of ownership, making large‑scale deployment financially viable for enterprises that need to process high volumes of forms, invoices, or archival records.

The model’s support for handwritten text recognition expands use cases to include legacy paperwork, field notes, and handwritten applications, areas traditionally difficult for automated systems.

Interactive Mechanism

Mecanism interactiv: cum funcționează de fapt

Explorați tehnologia care stau la baza acestei dezvoltări în mod interactiv.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Verificare interactivă a conceptului+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Ce să urmărești în continuare

Future announcements from Sarvam regarding pricing, cloud‑hosting options, or on‑premise licensing for Vision 2.1; adoption metrics from early enterprise pilots; community results on the Indic OCR Bench that could reveal strengths or weaknesses of the model; and any follow‑up releases that expand language coverage or add new document‑processing features.

Pricing and licensing details for Vision 2.1, which will determine accessibility for startups versus large enterprises.

Performance results from independent third‑party evaluations on the Indic OCR Bench, which could confirm or challenge Sarvam’s reported accuracy.

Adoption case studies from sectors such as banking, government, or education that illustrate real‑world impact and ROI.

Potential extensions to additional Indian languages or scripts, and any future model iterations that further improve accuracy or reduce latency.

Ghiduri și chestionare conexe

Modelele AI explicateAntrenament AIViitorul IATestați ceea ce știți — încercați un test AI gratuitCăutați un termen AI în glosarul nostruUrmați instrumentul de urmărire a lansării modelului AI
Ai găsit asta util?