From Scratch to SOTA: Training a 3B State-Space Vision Model for 1.4 Billion People
Krishna Prasad Srinivasan
- When
- Tuesday, June 303:20 PM – 3:40 PM · 20 min
- Where
- Track 2San Francisco, CA · imported from ai.engineer's public schedule feed
About this session
India has 22 official languages. Across those languages live over a billion people whose knowledge is locked inside scanned images in scripts that most frontier models perform poorly. The problem is dire - until now, there wasn't even a comprehensive benchmark to measure Indic OCR performance, let alone training data at scale. When Sarvam AI set out to solve this, we had to build the infrastructure before the model, creating the first ground-truth benchmark for Indic document intelligence. In this talk, Krishna Srinivasan, who led the Vision Models team to build India's first sovereign VLM from scratch, will walk through the end-to-end engineering lifecycle. We will cover: (a) Architecture: Why we chose a 3B-parameter state-space architecture over transformer baselines to handle high-resolution visual inputs with minimal memory overhead and faster inference. (b) Training Pipeline: The exact recipe we used: starting with text-only pre-training, moving to continual pre-training with text and images, followed by SFT. Finally, we'll cover the advances we made in implementing large-scale RL with Verifiable Rewards for visual tasks in just 3 days using deterministic character-level reward signals. (c) Compute Efficiency: How we trained a frontier-competitive multimodal model with extreme capital efficiency, optimizing distributed training and GPU cluster management to punch far above our compute class. (d) Agentic Workflows: How this model powers Sarvam Akshar, a first-of-its-kind agentic document intelligence workbench featuring visual grounding and automated proofreading loops. The results speak for themselves: Sarvam Vision achieves best-in-class global scores (84.3% on olmOCR-Bench, 93.28% on OmniDocBench) and dominates Indic OCR. Attendees will learn the blueprint for compute-efficient multimodal training, and deploying state-space VLMs for population-scale enterprise workloads.
Speaker
Head of Vision Models, Sarvam
Krishna Prasad Srinivasan is a Head of Vision Models at Sarvam, where he led a lean team to train Sarvam Vision, India's first sovereign VLM: a 3B state-space model that topped global OCR benchmarks at launch and led the Indic OCR Bench across 22 languages. He now leads the vision vertical's models, research, and product. Previously, he was Tech Lead for AI at Microsoft Research, where he built multilingual copilots for education and developed Indic translation models that outperformed commercial systems. Before that, Krishna was a researcher at Harvard, where he engineered a novel OCR architecture using contrastive learning that outperformed industry benchmarks on complex multilingual documents.
More in Vision & OCR
- The State of VisionTuesday, June 30 · 10:45 AM – 11:05 AM · Track 2
- Building the Document Context Layer for AI AgentsTuesday, June 30 · 11:10 AM – 11:30 AM · Track 2
- Skill issue: stop deploying vision language models, use them with Skills to build e2e vision apps on edgeTuesday, June 30 · 11:40 AM – 12:00 PM · Track 2
- Modality Misalignment and Originality Attribution in Short-Form Video: A Multi-Agent Approach at Platform ScaleTuesday, June 30 · 12:05 PM – 12:25 PM · Track 2
- From Ingestion to Agents: How Leading AI Teams Build on Document IntelligenceTuesday, June 30 · 1:30 PM – 1:50 PM · Track 2
For developers: this programme is open data — JSON, iCal, schedule XML and an MCP endpoint.Show endpointsHide
- JSONEvery published session and speaker, in one request./aie-worldsfair-2026-import/feed.json
- iCalSubscribe in Google, Apple or Outlook Calendar./aie-worldsfair-2026-import/feed.ics
- Schedule XMLfrab / pentabarf — the format conference apps import./aie-worldsfair-2026-import/feed.xml
- MCP + RESTPoint Claude at the programme. OpenAPI 3.1 included./agents
No key, no signup, CORS open. Everything here is generated from the same data the organisers edit.