Llama 1 models
1 Llama 1 model from Meta on Hugging Face, from 6.7B to 6.7B parameters, published by Meta in F16, with a 2K-token context. The smallest official model, llama-7b, needs about 15.1 GB of VRAM at its published precision; the cheapest live fit is V100 at $0.088/hr.
Part of the Llama series · Next generation: Llama 2
Pick a size
One row per official Llama 1 size: the VRAM it needs at each precision, the cheapest GPU that holds it today, and the KV cache for a 32K-token context.
| Model | Parameters | Native VRAM | FP8 VRAM | INT4 VRAM | Live GPU fit (native) | Est. $/hr | KV cache at 32K |
|---|---|---|---|---|---|---|---|
| llama-7b | 6.7B | 15.1 GB | 7.5 GB | 3.8 GB | V100 | $0.088/hr | 16.0 GB |
VRAM is weights times a flat 1.2 overhead; the KV cache is a separate per-model figure at 16-bit, shown where the architecture is published. See the methodology.
Official models (1)
VRAM is for the precision the model is published in: the weights times a flat 1.2 overhead, with the KV cache not included. See the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
More from Meta
- Llama 3.2 (15 models)
- Llama 3.1 (40 models)
- Llama 3 (19 models)
- Llama 2 (8 models)
- Llama 3.3 (4 models)
- Code Llama (3 models)
- Llama Guard (4 models)