How to run Granite 4.0 with vLLM and Transformers

Engine notes, chat template, tool calling and context length for running IBM Granite 4.0 hybrid Mamba2 MoE and dense models.

This guide covers serving the Granite 4.0 family. Memory sizing and launch commands for each size appear below.

Architecture and engines

The H Small card describes a hybrid decoder-only design: 4 attention layers and 36 Mamba2 layers, with a mixture-of-experts of 72 experts and 10 active per token (32B total, 9B active). The family also includes Micro Dense (3B), H Micro Dense (3B), H Tiny MoE (7B) and H Small MoE (32B).

  • The card shows Transformers (AutoModelForCausalLM, device_map="auto") and vllm serve, which exposes an OpenAI-compatible API on port 8000.
  • The card does not state minimum engine versions, and it does not document SGLang, llama.cpp or Ollama usage. Check each engine's own release notes first.

Chat template and tools

  • Turns use role tokens such as <|start_of_role|> and <|end_of_role|>.
  • A default system prompt guiding toward professional, accurate and safe answers was added in October 2025.
  • Tool calling follows the OpenAI function schema. The model emits JSON inside <tool_call> tags.
  • Multilingual quality outside English may need few-shot examples. Supported languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch and Chinese.

Long context

The sequence length is 128K tokens. The base card says the H Small model uses NoPE (no position embeddings) rather than RoPE, unlike the Micro Dense variant, so there is no YaRN or rope scaling setting documented for it. Context length is limited by memory, not a scaling flag.

Pitfalls

  • Pick the right checkpoint: the instruct models are the chat ones, the -base models are for completion and Fill-in-the-Middle code tasks.
  • Weights are Apache 2.0.

Launch commands by Granite 4 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

granite-4.0-350m (352M)

Needs about 0.8 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve ibm-granite/granite-4.0-350m --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.0-350m --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.0-1b (1.6B)

Needs about 3.6 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve ibm-granite/granite-4.0-1b --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.0-1b --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.0-h-micro (3.2B)

Needs about 7.1 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve ibm-granite/granite-4.0-h-micro --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.0-h-micro --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.1-3b (3.4B)

Needs about 7.6 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve ibm-granite/granite-4.1-3b --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.1-3b --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.2-3b (3.7B)

Needs about 8.2 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve ibm-granite/granite-4.2-3b --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.2-3b --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.0-h-tiny (6.9B)

Needs about 15.5 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve ibm-granite/granite-4.0-h-tiny --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.0-h-tiny --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.1-8b (8.8B)

Needs about 19.7 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve ibm-granite/granite-4.1-8b --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.1-8b --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.1-30b (28.9B)

Needs about 64.5 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve ibm-granite/granite-4.1-30b --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.1-30b --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.2-30b (29.3B)

Needs about 65.4 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve ibm-granite/granite-4.2-30b --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.2-30b --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

granite-4.0-h-small (32.2B)

Needs about 72.0 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve ibm-granite/granite-4.0-h-small --tensor-parallel-size 1

SGLang

sglang serve --model-path ibm-granite/granite-4.0-h-small --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.