How to run Llama 2 with vLLM, SGLang and llama.cpp
Gated access, the [INST] and <<SYS>> chat format and the 4K context limit for serving Llama 2 chat models.
This guide covers serving the Llama 2 family. Exact memory sizing and launch commands for each size appear below.
Access
The Hugging Face weights are gated. You must accept Meta's license and share contact information before you can download them, so serving needs an authenticated token.
Chat format
The chat models expect a specific prompt layout: INST and <<SYS>> tags plus BOS and EOS tokens. The card recommends calling strip() on inputs to avoid double spaces. Compare your rendered prompt against this layout when outputs look off.
Context length
The card lists a 4,096 token context. Plan prompts and generation to fit in 4K, and do not assume rope scaling extends it without a checkpoint that was trained for it.
Pitfalls
- The card says testing was conducted in English only, and it is not suited to non-English applications without extra work.
- The card tells developers to run safety testing before deployment, since outputs cannot be predicted in advance.
- The card does not state minimum versions for vLLM, SGLang, llama.cpp or Ollama, so check each engine's own release notes. For llama.cpp, convert to GGUF if you do not already have a GGUF file.
Launch commands by Llama 2 size
One model per size. VRAM is for the precision each is published in; the GPU count comes from the cheapest live fit, and you set it with --tensor-parallel-size (vLLM) or --tp (SGLang).
Llama-2-7b-hf (6.7B)
Needs about 15.1 GB of VRAM at F16. Cheapest live fit: V100 at $0.088/hr.
vLLM
vllm serve meta-llama/Llama-2-7b-hf --tensor-parallel-size 1SGLang
sglang serve --model-path meta-llama/Llama-2-7b-hf --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Llama-2-13b-chat-hf (13.0B)
Needs about 29.1 GB of VRAM at F16. Cheapest live fit: V100 at $0.187/hr.
vLLM
vllm serve meta-llama/Llama-2-13b-chat-hf --tensor-parallel-size 1SGLang
sglang serve --model-path meta-llama/Llama-2-13b-chat-hf --tp 1Generic example, not from the model's own docs: adjust flags (quantization, context length, parallelism) for your setup.
Launch on Aquanode
Aquanode sells GPU pods billed per second, not a hosted inference API. Rent a pod sized to the model above, open a terminal on it or save a command as a startup script, then connect to the endpoint it serves.
Sources
Updated 2026-10-07.