How to run Gemma 3 with vLLM, SGLang and llama.cpp

Run Gemma 3 with vLLM, SGLang or llama.cpp: context length, bfloat16, system prompts, image handling and pan and scan, from the model card.

Gemma 3 comes in 1B, 4B, 12B and 27B. See the Gemma 3 hub for per-size hardware.

Engine support

  • vLLM and SGLang are both listed on the 27B-it card. vLLM maps Gemma 3 to Gemma3ForCausalLM (text) and Gemma3ForConditionalGeneration (multimodal).
  • llama.cpp: GGUF builds are linked from the Hugging Face launch post.
  • Ollama: not covered in the sources opened, so no version is stated.
  • Transformers: v4.50.0 or later per the card.

Chat template

  • Use the instruction-tuned (-it) checkpoints with the built-in chat template.
  • The card's examples use a system message followed by a user message, and the launch post says short system prompts work well.

Context and attention

  • 128K tokens for 4B, 12B and 27B; 32K for 1B. Maximum output is 8,192 tokens.
  • The architecture interleaves local sliding-window layers (1,024-token window) with global layers at a 5 to 1 ratio. The sources give no rope scaling flags.

Multimodal inputs and pitfalls

  • The 1B model is text only; 4B and up accept images.
  • Images are normalized to 896 x 896 and encoded as 256 tokens each. Pan and scan cropping helps non-square or high-resolution images.
  • Use bfloat16: the launch post warns other dtypes may degrade quality.

Launch commands by Gemma 3 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

gemma-3-270m (268M)

Needs about 0.6 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve google/gemma-3-270m --tensor-parallel-size 1

SGLang

sglang serve --model-path google/gemma-3-270m --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

gemma-3-1b-it (1000M)

Needs about 2.2 GB of VRAM at BF16. Cheapest live fit: RTX 4070 Super at $0.121/hr.

vLLM

vllm serve google/gemma-3-1b-it --tensor-parallel-size 1

SGLang

sglang serve --model-path google/gemma-3-1b-it --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.