OCR training runs

Reading scanned Sanskrit books: an open vision-language model fine-tuned for page OCR, and the bulk OCR jobs that turn scans into corpus text.

Headline score: Median letter error rate · lower is better The share of letters read wrong on 54 held-out book pages, median over pages: 0.0096 means 0.96% of letters.Letter error rate: the share of letters an OCR model reads wrong, as a fraction (0.0096 means 0.96% of letters). Lower is better.

Live now

Progress

Lower is better on the headline.

Best Median letter error rate over time

Each dot is a scored run, placed at the day it ended; the line is the best score so far. Numbered lines are this track's milestones, listed below.

Milestones
  1. 5 Oct 2026 · OCR-VLM pilot: first OCR run
  2. 6 Oct 2026 · OCR-VLM v1 beats Google Vision on Sanskrit pages

Model size and Median letter error rate

Parameters (log scale) against Median letter error rate: this track's runs. Hover a point for its name.

All OCR runs

Published · 4 runs

All OCR runs. Column headers sort the table.
Data
OCR pilot test
ocr_pilot100k_test
6 Oct 2026 — — 0.57 $0.52 — —
OCR-VLM v1
ocrvlm_v1_q35
5 Oct 2026 — OCR page labels 7.33 $9.53 Winner 0.0096
OCR-VLM pilot
ocrvlm_pilot_q35
4 Oct 2026 — OCR page labels 1.37 $1.80 Keep —
OCR pilot 100k Running
ocr_pilot100k
— — — — $21.63 — —

Cost “—” means no cloud bill (our own GPU at home) or a cost that was not recorded; “Free” is free cloud compute. GPU-h is GPU hours.