Yi-1 models
5 Yi-1 models from 01.AI on Hugging Face, from 6.1B to 34.4B parameters, published by 01.AI in BF16, with contexts from 4K to 200,000 tokens. The smallest official model, Yi-6B, needs about 13.5 GB of VRAM at its published precision; the cheapest live fit is RTX A4000 at $0.167/hr.
Pick a size
One row per official Yi-1 size: the VRAM it needs at each precision, the cheapest GPU that holds it today, and the KV cache for a 32K-token context.
| Model | Parameters | Native VRAM | FP8 VRAM | INT4 VRAM | Live GPU fit (native) | Est. $/hr | KV cache at 32K |
|---|---|---|---|---|---|---|---|
| Yi-6B | 6.1B | 13.5 GB | 6.8 GB | 3.4 GB | RTX A4000 | $0.167/hr | 2.0 GB |
| Yi-34B-Chat | 34.4B | 76.9 GB | 38.4 GB | 19.2 GB | A100 | $1.21/hr | 7.5 GB |
VRAM is weights times a flat 1.2 overhead; the KV cache is a separate per-model figure at 16-bit, shown where the architecture is published. See the methodology.
Official models (5)
| Model | Parameters | Published as | VRAM needed | Live GPU fit | Est. $/hr |
|---|---|---|---|---|---|
| Yi-6B | 6.1B | BF16 | 13.5 GB | RTX A4000 | $0.167/hr |
| Yi-6B-Chat | 6.1B | BF16 | 13.5 GB | RTX A4000 | $0.167/hr |
| Yi-6B-200K | 6.1B | BF16 | 13.5 GB | RTX A4000 | $0.167/hr |
| Yi-34B | 34.4B | BF16 | 76.9 GB | A100 | $1.21/hr |
| Yi-34B-Chat | 34.4B | BF16 | 76.9 GB | A100 | $1.21/hr |
VRAM is for the precision the model is published in: the weights times a flat 1.2 overhead, with the KV cache not included. See the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
More from 01.AI
- Yi-1.5 (6 models)