How to run Gemma 2 with vLLM, SGLang and llama.cpp

Run Gemma 2 with vLLM, SGLang or llama.cpp: the turn-based chat template, bfloat16 weights, and engine support listed on the model card.

Gemma 2 documentation is short. This page covers what the 9B instruction-tuned card and vLLM's docs state. See the Gemma 2 hub for sizes and hardware.

Engine support

  • vLLM and SGLang: both are shown on the card with OpenAI-compatible servers. vLLM maps Gemma 2 to Gemma2ForCausalLM.
  • llama.cpp and Ollama: the card links quantized builds (GGUF) for both.
  • Transformers: update to the latest release with pip install -U transformers; the card gives no specific minimum.

Chat template

  • Turns are delimited by <start_of_turn> and <end_of_turn>, with the role (user or model) after the opening delimiter.
  • The card describes only user and model roles, so fold any system instructions into the first user turn.
  • Use the instruction-tuned (-it) checkpoint for chat; the pre-trained one has no chat behavior.

Dtype

Weights are native bfloat16. The card notes that upcasting to float32 gives no precision gain, so serve in bfloat16.

Not stated

The card does not state a context length or rope settings, so check max_position_embeddings in the checkpoint's config before setting a length flag.

Launch commands by Gemma 2 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

gemma-2b-it (2.5B)

Needs about 5.6 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve google/gemma-2b-it --tensor-parallel-size 1

SGLang

sglang serve --model-path google/gemma-2b-it --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

gemma-2-2b-it (2.6B)

Needs about 5.8 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve google/gemma-2-2b-it --tensor-parallel-size 1

SGLang

sglang serve --model-path google/gemma-2-2b-it --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

gemma-2-9b-it (9.2B)

Needs about 20.7 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve google/gemma-2-9b-it --tensor-parallel-size 1

SGLang

sglang serve --model-path google/gemma-2-9b-it --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

gemma-2-27b-it (27.2B)

Needs about 60.9 GB of VRAM at BF16. Cheapest live fit: A100 at $1.21/hr.

vLLM

vllm serve google/gemma-2-27b-it --tensor-parallel-size 1

SGLang

sglang serve --model-path google/gemma-2-27b-it --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.