Best open OCR models
How to pick an open OCR model for scans and PDFs: output format, resolution, license and size. Picks with facts from the model cards.
OCR models turn page images into text. The current open generation is vision-language models fine-tuned for documents, which read layout, tables, and equations and return markdown instead of loose text. Picking one comes down to the output you need, throughput at page volume, and license. The catalog below lists the rest.
What matters when picking an OCR model
Output format. Plain text, markdown, or text with layout boxes are different jobs. Check whether the card's example prompt produces what your pipeline consumes. DeepSeek-OCR's own example prompt is "Convert the document to markdown."
Volume. OCR is usually batch work over thousands or millions of pages. Throughput per GPU-hour matters more than single-page latency, and a small model that is accurate enough wins on cost. Batching in an inference server like vLLM, which both cards below document, is how you get throughput.
Hard content. Tables, math, handwriting, and multi-column layouts separate the models. Ask for a benchmark on your own pages, because vendor numbers use the vendor's test set.
Resolution modes. Some models trade image resolution against token count. DeepSeek-OCR documents Tiny and Gundam modes, set through base size, image size, and crop settings.
Precision. BF16 is the reference. Quantized builds such as FP8 raise throughput on supported GPUs. See the VRAM calculator.
License. Both picks below are permissive, per their cards.
Picks by situation
Small and fast: DeepSeek-OCR. The catalog lists about 3.3B parameters, and the card lists an MIT license. The paper, "DeepSeek-OCR: Contexts Optical Compression," explores compressing text into vision tokens. The card points to the GitHub repository for inference acceleration and PDF processing.
Whole-PDF pipelines: olmOCR-2-7B-1025. Ai2's card says it is fine-tuned from Qwen2.5-VL-7B-Instruct and trained further with GRPO reinforcement learning to improve math equations, tables, and other tricky OCR cases. The card is licensed under Apache 2.0 and reports olmOCR-Bench scores when used with the olmOCR toolkit, which renders, rotates, and retries pages and runs efficiently through vLLM. The catalog shows about 8.3B parameters.
Highest throughput on the same model: olmOCR-2-7B-1025-FP8. The BF16 card says Ai2 recommends the FP8 version for all practical purposes except further fine-tuning.
One 80 GB GPU: either model leaves plenty of room, so use the extra memory for larger batches rather than a larger model.
Most permissive license: DeepSeek-OCR is MIT and olmOCR-2 is Apache 2.0.
Serving notes
Render PDF pages at a consistent resolution before sending them, and retry pages that fail rather than rerunning the whole file. Keep the GPU busy with a queue of pages; an idle card is the most expensive part of an OCR job.
Aquanode rents GPUs by the hour, which suits batch OCR: start a card, run the queue, and stop. See pricing. Compare two models on a sample of your own pages before you commit a whole archive.
Open models for OCR and document parsing
All 27 models in the catalog for this task, grouped by size.
Under 3B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| GLM-OCR | 1.3B | 3.0 GB | – | RTX 4070 Super | $0.121/hr |
| surya-ocr-2 | 686M | 1.5 GB | – | RTX 4070 Super | $0.121/hr |
| GOT-OCR2_0 | 716M | 1.6 GB | – | RTX 4070 Super | $0.121/hr |
| HunyuanOCR | 1.1B | 2.5 GB | – | RTX 4070 Super | $0.121/hr |
| LightOnOCR-2-1B | 1.0B | 2.2 GB | – | RTX 4070 Super | $0.121/hr |
| MinerU2.5-Pro-2604-1.2B | 1.2B | 2.6 GB | – | RTX 4070 Super | $0.121/hr |
| granite-docling-258M | 258M | 0.6 GB | – | RTX 4070 Super | $0.121/hr |
| trocr-base-printed | 333M | 1.5 GB | – | V100 | $0.088/hr |
| trocr-base-handwritten | 333M | 1.5 GB | – | V100 | $0.088/hr |
| GOT-OCR-2.0-hf | 561M | 1.3 GB | – | RTX 4070 Super | $0.121/hr |
| Qari-OCR-v0.3-VL-2B-Instruct | 2.2B | 4.9 GB | – | V100 | $0.088/hr |
| NVIDIA-Nemotron-Parse-v1.1 | 957M | 4.3 GB | – | V100 | $0.088/hr |
| OvisOCR2 | 853M | 1.9 GB | – | RTX 4070 Super | $0.121/hr |
| MinerU2.5-2509-1.2B | 1.2B | 2.6 GB | – | RTX 4070 Super | $0.121/hr |
| typhoon-ocr1.5-2b | 2.1B | 4.8 GB | – | RTX 4070 Super | $0.121/hr |
| LightOnOCR-1B-1025 | 1.2B | 2.6 GB | – | RTX 4070 Super | $0.121/hr |
| TeleOCR | 1.4B | 3.2 GB | – | RTX 4070 Super | $0.121/hr |
| LightOnOCR-2-1B-bbox-soup | 1.0B | 2.2 GB | – | RTX 4070 Super | $0.121/hr |
3B to 10B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Unlimited-OCR | 3.3B | 7.5 GB | – | RTX 4070 Super | $0.121/hr |
| chandra-ocr-2 | 5.3B | 11.8 GB | – | RTX 4070 Super | $0.121/hr |
| DeepSeek-OCR | 3.3B | 7.5 GB | MIT | RTX 4070 Super | $0.121/hr |
| DeepSeek-OCR-2 | 3.4B | 7.6 GB | – | RTX 4070 Super | $0.121/hr |
| dots.mocr | 3.0B | 6.8 GB | – | RTX 4070 Super | $0.121/hr |
| typhoon-ocr-3b | 3.8B | 8.4 GB | – | RTX 4070 Super | $0.121/hr |
| dots.ocr | 3.0B | 6.8 GB | – | RTX 4070 Super | $0.121/hr |
| olmOCR-2-7B-1025 | 8.3B | 18.5 GB | – | RTX A5000 | $0.176/hr |
10B to 40B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Infinity-Parser2-Pro | 35.1B | 78.5 GB | – | A100 | $1.21/hr |
VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
Sources
Updated 2026-10-07.
More ways to choose a model
- Best open models for coding
- Best open models for reasoning
- Best open models for chat and assistants
- Best open models for vision-language
- Best open models for speech-to-text
- Best open models for text-to-speech
- Best open models for image generation
- Best open models for video generation
- Best open models for embeddings
- Best open models for translation
By GPU memory: