How to run Llama 4 with vLLM and SGLang
Engine support, Transformers version, tool-call parser, multimodal input and context notes for running Llama 4 Scout and Maverick.
This guide covers serving the Llama 4 family. Sizing and launch commands per model appear below.
Engine support
The Llama 4 Scout Instruct card documents vLLM (OpenAI-compatible server) and SGLang, and says both handle multimodal input of text plus images. It requires Transformers 4.51.0 or higher. The card does not cover llama.cpp or Ollama, so check those projects directly.
Multimodal input
- Input is multilingual text and images, with output as text and code.
- The card says the model was tested with up to 5 input images. Meta tells developers to run their own testing beyond that.
Context length
The card lists a 10M token context window for Scout and 1M for Maverick. Set the server context limit (for example --max-model-len) to what your workload needs rather than the full window.
Tool calling
vLLM documents the llama4_pythonic parser for all Llama 4 models, used with its tool_chat_template_llama4_pythonic.jinja template. Unlike Llama 3, vLLM states that parallel tool calls are supported for Llama 4.
Quantization
The card provides BF16 weights and says int4 quantization is available on the fly. FP8 weights are listed for Maverick (see FP8).
Launch commands by Llama 4 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
No Llama 4 model of a size we can compute is in the catalog yet. See the Llama 4 model list for what is published.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
- https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct
- https://docs.vllm.ai/en/latest/features/tool_calling.html
Updated 2026-10-07.