June 29 – July 2, 2026 · San Francisco, CA · imported from ai.engineer's public schedule feed

AI Engineer World's Fair 2026 — unofficial import demo

Unofficial demo. This programme was imported from the AI Engineer World's Fair's own public schedule feed to show vibeboard at real conference scale. Not affiliated with, or endorsed by, the organisers.

All sessions
Vision & OCRSponsor Session

From Scratch to SOTA: Training a 3B State-Space Vision Model for 1.4 Billion People

Krishna Prasad Srinivasan

When
Tuesday, June 303:20 PM – 3:40 PM · 20 min
Where
Track 2San Francisco, CA · imported from ai.engineer's public schedule feed
Google Calendar

About this session

India has 22 official languages. Across those languages live over a billion people whose knowledge is locked inside scanned images in scripts that most frontier models perform poorly. The problem is dire - until now, there wasn't even a comprehensive benchmark to measure Indic OCR performance, let alone training data at scale. When Sarvam AI set out to solve this, we had to build the infrastructure before the model, creating the first ground-truth benchmark for Indic document intelligence. In this talk, Krishna Srinivasan, who led the Vision Models team to build India's first sovereign VLM from scratch, will walk through the end-to-end engineering lifecycle. We will cover: (a) Architecture: Why we chose a 3B-parameter state-space architecture over transformer baselines to handle high-resolution visual inputs with minimal memory overhead and faster inference. (b) Training Pipeline: The exact recipe we used: starting with text-only pre-training, moving to continual pre-training with text and images, followed by SFT. Finally, we'll cover the advances we made in implementing large-scale RL with Verifiable Rewards for visual tasks in just 3 days using deterministic character-level reward signals. (c) Compute Efficiency: How we trained a frontier-competitive multimodal model with extreme capital efficiency, optimizing distributed training and GPU cluster management to punch far above our compute class. (d) Agentic Workflows: How this model powers Sarvam Akshar, a first-of-its-kind agentic document intelligence workbench featuring visual grounding and automated proofreading loops. The results speak for themselves: Sarvam Vision achieves best-in-class global scores (84.3% on olmOCR-Bench, 93.28% on OmniDocBench) and dominates Indic OCR. Attendees will learn the blueprint for compute-efficient multimodal training, and deploying state-space VLMs for population-scale enterprise workloads.

Speaker

Krishna Prasad Srinivasan
Krishna Prasad Srinivasan

Head of Vision Models, Sarvam

Krishna Prasad Srinivasan is a Head of Vision Models at Sarvam, where he led a lean team to train Sarvam Vision, India's first sovereign VLM: a 3B state-space model that topped global OCR benchmarks at launch and led the Indic OCR Bench across 22 languages. He now leads the vision vertical's models, research, and product. Previously, he was Tech Lead for AI at Microsoft Research, where he built multilingual copilots for education and developed Indic translation models that outperformed commercial systems. Before that, Krishna was a researcher at Harvard, where he engineered a novel OCR architecture using contrastive learning that outperformed industry benchmarks on complex multilingual documents.

More in Vision & OCR