How to run Granite 3.3 with vLLM and SGLang
Engine support, chat template, the thinking toggle and 128K context notes for running IBM Granite 3.3 instruct models.
This guide covers serving the Granite 3 family, using the Granite 3.3 8B cards as the reference.
Engine support
- The instruct card lists PyTorch, Transformers and Accelerate for local use, and states compatibility with vLLM and SGLang.
- It does not give minimum engine versions, and it does not document llama.cpp or Ollama.
- Weights are BF16.
Chat template and thinking
- Use
tokenizer.apply_chat_template()withroleandcontentmessages. - Reasoning is switched on with a
thinkingparameter. When enabled, the model wraps its reasoning in<think></think>and the final answer in<response></response>. A server that returns raw text needs those tags parsed or stripped client side. - Leave thinking off for plain instruction following.
Context length
The cards state a 128K token context window. They document no YaRN or rope scaling flags.
Good fits
The card lists summarization, classification, extraction, question answering, code, RAG and long-document tasks. Supported languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch and Chinese.
Pitfalls
- The
-basecheckpoint is a plain completion model (with Fill-in-the-Middle code support), not a chat model. Serve the instruct checkpoint for chat. - Weights are Apache 2.0.
Launch commands by Granite 3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
granite-3.0-1b-a400m-instruct (1.3B)
Needs about 3.0 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve ibm-granite/granite-3.0-1b-a400m-instruct --tensor-parallel-size 1SGLang
sglang serve --model-path ibm-granite/granite-3.0-1b-a400m-instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
granite-3.3-2b-instruct (2.5B)
Needs about 5.7 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve ibm-granite/granite-3.3-2b-instruct --tensor-parallel-size 1SGLang
sglang serve --model-path ibm-granite/granite-3.3-2b-instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
granite-3.0-8b-instruct (8.2B)
Needs about 18.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.
vLLM
vllm serve ibm-granite/granite-3.0-8b-instruct --tensor-parallel-size 1SGLang
sglang serve --model-path ibm-granite/granite-3.0-8b-instruct --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/ibm-granite/granite-3.3-8b-instruct
- https://huggingface.co/ibm-granite/granite-3.3-8b-base
Updated 2026-10-07.