How to run Olmo 3 with vLLM and SGLang

Engine support, chat template, Think versus Instruct behavior and recommended sampling for running Ai2 Olmo 3 models.

This guide covers serving the Olmo 3 family from Ai2.

Variants

The family has Base (Olmo-3-7B), Think (7B and 32B) and Instruct (7B and 32B) models. Pick Think when you want reasoning traces and Instruct for direct answers.

Engine support

  • The cards state support for vLLM and SGLang through OpenAI-compatible APIs. The Think card shows vllm serve "allenai/Olmo-3-7B-Think".
  • Transformers 4.57.0 or newer is required.
  • The cards do not document llama.cpp or Ollama.

Chat template

  • Turns use <|im_start|> and <|im_end|>.
  • The Think model has a default system message identifying it as Olmo, built by Ai2, with a December 2024 date cutoff.
  • The Instruct model's default system prompt reads "You are a helpful function-calling AI assistant. You do not currently have access to any functions."
  • Think models emit reasoning inside <think> tags before the answer, so parse or strip them in your client.

Sampling and length

The cards recommend temperature 0.6, top_p 0.95 and max tokens 32,768.

The base 7B card states a 65,536 token context window. The cards document no YaRN or rope scaling flags, so do not set them without checking the config.

Pitfalls

  • The Base model has no safety filtering and is not a chat model.
  • Weights are Apache 2.0 and intended for research and educational use under Ai2's Responsible Use Guidelines.

Launch commands by Olmo 3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Olmo-3-7B-Instruct (7.3B)

Needs about 16.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve allenai/Olmo-3-7B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path allenai/Olmo-3-7B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Olmo-3-1125-32B (32.2B)

Needs about 72.0 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve allenai/Olmo-3-1125-32B --tensor-parallel-size 1

SGLang

sglang serve --model-path allenai/Olmo-3-1125-32B --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.