How to run Olmo 3 with vLLM and SGLang
Engine support, chat template, Think versus Instruct behavior and recommended sampling for running Ai2 Olmo 3 models.
This guide covers serving the Olmo 3 family from Ai2.
Variants
The family has Base (Olmo-3-7B), Think (7B and 32B) and Instruct (7B and 32B) models. Pick Think when you want reasoning traces and Instruct for direct answers.
Engine support
- The cards state support for vLLM and SGLang through OpenAI-compatible APIs. The Think card shows
vllm serve "allenai/Olmo-3-7B-Think". - Transformers 4.57.0 or newer is required.
- The cards do not document llama.cpp or Ollama.
Chat template
- Turns use
<|im_start|>and<|im_end|>. - The Think model has a default system message identifying it as Olmo, built by Ai2, with a December 2024 date cutoff.
- The Instruct model's default system prompt reads "You are a helpful function-calling AI assistant. You do not currently have access to any functions."
- Think models emit reasoning inside
<think>tags before the answer, so parse or strip them in your client.
Sampling and length
The cards recommend temperature 0.6, top_p 0.95 and max tokens 32,768.
The base 7B card states a 65,536 token context window. The cards document no YaRN or rope scaling flags, so do not set them without checking the config.
Pitfalls
- The Base model has no safety filtering and is not a chat model.
- Weights are Apache 2.0 and intended for research and educational use under Ai2's Responsible Use Guidelines.
Launch commands by Olmo 3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Olmo-3-7B-Instruct (7.3B)
Needs about 16.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve allenai/Olmo-3-7B-Instruct --tensor-parallel-size 1SGLang
sglang serve --model-path allenai/Olmo-3-7B-Instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Olmo-3-1125-32B (32.2B)
Needs about 72.0 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.
vLLM
vllm serve allenai/Olmo-3-1125-32B --tensor-parallel-size 1SGLang
sglang serve --model-path allenai/Olmo-3-1125-32B --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/allenai/Olmo-3-7B-Instruct
- https://huggingface.co/allenai/Olmo-3-7B-Think
- https://huggingface.co/allenai/Olmo-3-1025-7B
Updated 2026-10-07.