Best open OCR models

How to pick an open OCR model for scans and PDFs: output format, resolution, license and size. Picks with facts from the model cards.

OCR models turn page images into text. The current open generation is vision-language models fine-tuned for documents, which read layout, tables, and equations and return markdown instead of loose text. Picking one comes down to the output you need, throughput at page volume, and license. The catalog below lists the rest.

What matters when picking an OCR model

Output format. Plain text, markdown, or text with layout boxes are different jobs. Check whether the card's example prompt produces what your pipeline consumes. DeepSeek-OCR's own example prompt is "Convert the document to markdown."

Volume. OCR is usually batch work over thousands or millions of pages. Throughput per GPU-hour matters more than single-page latency, and a small model that is accurate enough wins on cost. Batching in an inference server like vLLM, which both cards below document, is how you get throughput.

Hard content. Tables, math, handwriting, and multi-column layouts separate the models. Ask for a benchmark on your own pages, because vendor numbers use the vendor's test set.

Resolution modes. Some models trade image resolution against token count. DeepSeek-OCR documents Tiny and Gundam modes, set through base size, image size, and crop settings.

Precision. BF16 is the reference. Quantized builds such as FP8 raise throughput on supported GPUs. See the VRAM calculator.

License. Both picks below are permissive, per their cards.

Picks by situation

Small and fast: DeepSeek-OCR. The catalog lists about 3.3B parameters, and the card lists an MIT license. The paper, "DeepSeek-OCR: Contexts Optical Compression," explores compressing text into vision tokens. The card points to the GitHub repository for inference acceleration and PDF processing.

Whole-PDF pipelines: olmOCR-2-7B-1025. Ai2's card says it is fine-tuned from Qwen2.5-VL-7B-Instruct and trained further with GRPO reinforcement learning to improve math equations, tables, and other tricky OCR cases. The card is licensed under Apache 2.0 and reports olmOCR-Bench scores when used with the olmOCR toolkit, which renders, rotates, and retries pages and runs efficiently through vLLM. The catalog shows about 8.3B parameters.

Highest throughput on the same model: olmOCR-2-7B-1025-FP8. The BF16 card says Ai2 recommends the FP8 version for all practical purposes except further fine-tuning.

One 80 GB GPU: either model leaves plenty of room, so use the extra memory for larger batches rather than a larger model.

Most permissive license: DeepSeek-OCR is MIT and olmOCR-2 is Apache 2.0.

Serving notes

Render PDF pages at a consistent resolution before sending them, and retry pages that fail rather than rerunning the whole file. Keep the GPU busy with a queue of pages; an idle card is the most expensive part of an OCR job.

Aquanode rents GPUs by the hour, which suits batch OCR: start a card, run the queue, and stop. See pricing. Compare two models on a sample of your own pages before you commit a whole archive.

Open models for OCR and document parsing

All 27 models in the catalog for this task, grouped by size.

Under 3B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
GLM-OCR1.3B3.0 GB–RTX 4070 Super$0.121/hr
surya-ocr-2686M1.5 GB–RTX 4070 Super$0.121/hr
GOT-OCR2_0716M1.6 GB–RTX 4070 Super$0.121/hr
HunyuanOCR1.1B2.5 GB–RTX 4070 Super$0.121/hr
LightOnOCR-2-1B1.0B2.2 GB–RTX 4070 Super$0.121/hr
MinerU2.5-Pro-2604-1.2B1.2B2.6 GB–RTX 4070 Super$0.121/hr
granite-docling-258M258M0.6 GB–RTX 4070 Super$0.121/hr
trocr-base-printed333M1.5 GB–V100$0.088/hr
trocr-base-handwritten333M1.5 GB–V100$0.088/hr
GOT-OCR-2.0-hf561M1.3 GB–RTX 4070 Super$0.121/hr
Qari-OCR-v0.3-VL-2B-Instruct2.2B4.9 GB–V100$0.088/hr
NVIDIA-Nemotron-Parse-v1.1957M4.3 GB–V100$0.088/hr
OvisOCR2853M1.9 GB–RTX 4070 Super$0.121/hr
MinerU2.5-2509-1.2B1.2B2.6 GB–RTX 4070 Super$0.121/hr
typhoon-ocr1.5-2b2.1B4.8 GB–RTX 4070 Super$0.121/hr
LightOnOCR-1B-10251.2B2.6 GB–RTX 4070 Super$0.121/hr
TeleOCR1.4B3.2 GB–RTX 4070 Super$0.121/hr
LightOnOCR-2-1B-bbox-soup1.0B2.2 GB–RTX 4070 Super$0.121/hr

3B to 10B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Unlimited-OCR3.3B7.5 GB–RTX 4070 Super$0.121/hr
chandra-ocr-25.3B11.8 GB–RTX 4070 Super$0.121/hr
DeepSeek-OCR3.3B7.5 GBMITRTX 4070 Super$0.121/hr
DeepSeek-OCR-23.4B7.6 GB–RTX 4070 Super$0.121/hr
dots.mocr3.0B6.8 GB–RTX 4070 Super$0.121/hr
typhoon-ocr-3b3.8B8.4 GB–RTX 4070 Super$0.121/hr
dots.ocr3.0B6.8 GB–RTX 4070 Super$0.121/hr
olmOCR-2-7B-10258.3B18.5 GB–RTX A5000$0.176/hr

10B to 40B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Infinity-Parser2-Pro35.1B78.5 GB–A100$1.21/hr

VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

Sources

Updated 2026-10-07.

More ways to choose a model

By GPU memory:

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.