How to run Gemma 4 with vLLM, SGLang and llama.cpp
Run Gemma 4 with vLLM, SGLang or llama.cpp: system role, the thinking toggle, image and audio ordering, and engine notes from the model card.
Gemma 4 ships in five variants (E2B, E4B, 12B Unified, 26B A4B MoE and 31B Dense). See the Gemma 4 hub for per-size hardware.
Engine support
- vLLM: the model card serves it with
vllm serve, and vLLM's supported-models list maps Gemma 4 toGemma4ForCausalLMandGemma4ForConditionalGeneration. - SGLang: the card launches it with
python3 -m sglang.launch_server. - llama.cpp: the Hugging Face launch post lists full image and text support with an OpenAI-compatible API. Quantized builds (GGUF) are linked from the card.
- Ollama: the card points to quantized builds for compatible apps. No minimum version is stated in the sources.
Chat template and thinking
- The
systemrole is natively supported. - Thinking is switched on by putting the
<|think|>token at the start of the system prompt (the Hugging Face post exposes this asenable_thinking=Truein the template). - With thinking on, the model emits its reasoning in a
<|channel>thoughtblock before the final answer, so strip that block if your client does not.
Multimodal inputs
- All sizes take text and images. Put image content before the text in the prompt.
- Audio goes after the text. The card lists audio for E2B, E4B and 12B, while vLLM's docs say audio is supported only on E2B and E4B, so test audio on your engine version.
- vLLM notes the model does not ingest video directly and converts it internally.
Context length
Context windows range from 128K tokens on the smallest models to 256K on the medium and large ones. The sources give no rope or YaRN flags, so none are needed for the stated window.
Launch commands by Gemma 4 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
No Gemma 4 model of a size we can compute is in the catalog yet. See the Gemma 4 model list for what is published.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/google/gemma-4-31B-it
- https://huggingface.co/blog/gemma4
- https://docs.vllm.ai/en/latest/models/supported_models.html
Updated 2026-10-07.