How to run Llama 3.1 with vLLM, SGLang and Ollama

Engine support, Transformers version, tool-call parser, 128k context and quantized runtimes for running Llama 3.1 8B, 70B and 405B.

This guide covers serving the Llama 3.1 family. Sizing and launch commands per model appear below.

Engine support

  • The 8B Instruct card requires Transformers 4.43.0 or newer.
  • It names vLLM and SGLang as deployment frameworks with OpenAI-compatible APIs.
  • It notes quantized versions for llama.cpp, Ollama and LM Studio (see GGUF).

Context

The card lists a 128k token context length and Grouped-Query Attention. It does not give rope scaling flags; for background see rope scaling, and set your server context limit to what you need.

Chat template and tools

The model ships a chat template, and tool calling goes through tokenizer.apply_chat_template() with a tools parameter. In vLLM use the llama3_json parser with the 3.1 JSON template (--tool-call-parser, --chat-template).

Known pitfalls

  • vLLM states parallel tool calls are not supported for Llama 3.
  • The model may serialize array parameters as strings instead of proper arrays.

Launch commands by Llama 3.1 size

One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).

Llama-3.1-8B-Instruct (8.0B)

Needs about 17.9 GB of VRAM at BF16. Cheapest live fit: RTX A5000 at $0.176/hr.

vLLM

vllm serve meta-llama/Llama-3.1-8B-Instruct --tensor-parallel-size 1

SGLang

sglang serve --model-path meta-llama/Llama-3.1-8B-Instruct --tp 1

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Ollama

ollama run llama3.1:8b

Verified against Ollama's own library listing. Source

Llama-3.1-70B-Instruct (70.6B)

Needs about 158 GB of VRAM at BF16. Cheapest live fit: 7× RTX A5000 at $1.23/hr.

vLLM

vllm serve meta-llama/Llama-3.1-70B-Instruct --tensor-parallel-size 7

SGLang

sglang serve --model-path meta-llama/Llama-3.1-70B-Instruct --tp 7

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Llama-3.1-405B (405.9B)

Needs about 907 GB of VRAM at BF16. No live GPU fit right now.

vLLM

vllm serve meta-llama/Llama-3.1-405B

SGLang

sglang serve --model-path meta-llama/Llama-3.1-405B

Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.

Launch on Aquanode

Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.

Sources

Updated 2026-10-07.

Related

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.