How to run Gemma 3 with vLLM, SGLang and llama.cpp
Run Gemma 3 with vLLM, SGLang or llama.cpp: context length, bfloat16, system prompts, image handling and pan and scan, from the model card.
Gemma 3 comes in 1B, 4B, 12B and 27B. See the Gemma 3 hub for per-size hardware.
Engine support
- vLLM and SGLang are both listed on the 27B-it card. vLLM maps Gemma 3 to
Gemma3ForCausalLM(text) andGemma3ForConditionalGeneration(multimodal). - llama.cpp: GGUF builds are linked from the Hugging Face launch post.
- Ollama: not covered in the sources opened, so no version is stated.
- Transformers: v4.50.0 or later per the card.
Chat template
- Use the instruction-tuned (
-it) checkpoints with the built-in chat template. - The card's examples use a
systemmessage followed by ausermessage, and the launch post says short system prompts work well.
Context and attention
- 128K tokens for 4B, 12B and 27B; 32K for 1B. Maximum output is 8,192 tokens.
- The architecture interleaves local sliding-window layers (1,024-token window) with global layers at a 5 to 1 ratio. The sources give no rope scaling flags.
Multimodal inputs and pitfalls
- The 1B model is text only; 4B and up accept images.
- Images are normalized to 896 x 896 and encoded as 256 tokens each. Pan and scan cropping helps non-square or high-resolution images.
- Use
bfloat16: the launch post warns other dtypes may degrade quality.
Launch commands by Gemma 3 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
gemma-3-270m (268M)
Needs about 0.6 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve google/gemma-3-270m --tensor-parallel-size 1SGLang
sglang serve --model-path google/gemma-3-270m --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
gemma-3-1b-it (1000M)
Needs about 2.2 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.
vLLM
vllm serve google/gemma-3-1b-it --tensor-parallel-size 1SGLang
sglang serve --model-path google/gemma-3-1b-it --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/google/gemma-3-27b-it
- https://huggingface.co/blog/gemma3
- https://docs.vllm.ai/en/latest/models/supported_models.html
Updated 2026-10-07.