How to run Granite 3.3 with vLLM and SGLang

Engine support, chat template, the thinking toggle and 128K context notes for running IBM Granite 3.3 instruct models.

This guide covers serving the Granite 3 family, using the Granite 3.3 8B cards as the reference.

Engine support

  • The instruct card lists PyTorch, Transformers and Accelerate for local use, and states compatibility with vLLM and SGLang.
  • It does not give minimum engine versions, and it does not document llama.cpp or Ollama.
  • Weights are BF16.

Chat template and thinking

  • Use tokenizer.apply_chat_template() with role and content messages.
  • Reasoning is switched on with a thinking parameter. When enabled, the model wraps its reasoning in <think></think> and the final answer in <response></response>. A server that returns raw text needs those tags parsed or stripped client side.
  • Leave thinking off for plain instruction following.

Context length

The cards state a 128K token context window. They document no YaRN or rope scaling flags.

Good fits

The card lists summarization, classification, extraction, question answering, code, RAG and long-document tasks. Supported languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch and Chinese.

Pitfalls

  • The -base checkpoint is a plain completion model (with Fill-in-the-Middle code support), not a chat model. Serve the instruct checkpoint for chat.
  • Weights are Apache 2.0.

Launch commands by Granite 3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

granite-3.0-1b-a400m-instruct (1.3B)

Needs about 3.0 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve ibm-granite/granite-3.0-1b-a400m-instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-3.0-1b-a400m-instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-3.3-2b-instruct (2.5B)

Needs about 5.7 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve ibm-granite/granite-3.3-2b-instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-3.3-2b-instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-3.0-8b-instruct (8.2B)

Needs about 18.3 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve ibm-granite/granite-3.0-8b-instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-3.0-8b-instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.