SmolLM-1 models

5 SmolLM-1 models from Hugging Face on Hugging Face, from 135M to 1.7B parameters, published by Hugging Face in F32 and BF16, with a 2K-token context. The smallest official model, SmolLM-135M-Instruct, needs about 0.3 GB of VRAM at its published precision; the cheapest live fit is RTX 4070 Super at $0.121/hr.

Part of the SmolLM series · Next generation: SmolLM2

Pick a size

One row per official SmolLM-1 size: the VRAM it needs at each precision, the cheapest GPU that holds it today, and the KV cache for a 32K-token context.

ModelParametersNative VRAMFP8 VRAMINT4 VRAMLive GPU fit (native)Est. $/hrKV cache at 32K
SmolLM-135M135M0.6 GB0.2 GB0.1 GBV100$0.088/hr0.70 GB
SmolLM-360M-Instruct362M0.8 GB0.4 GB0.2 GBRTX 4070 Super$0.121/hr1.3 GB
SmolLM-1.7B1.7B7.7 GB1.9 GB1.0 GBV100$0.088/hr6.0 GB

VRAM is weights times a flat 1.2 overhead; the KV cache is a separate per-model figure at 16-bit, shown where the architecture is published. See the methodology.

Official models (4)

ModelParametersPublished asVRAM neededLive GPU fitEst. $/hr
SmolLM-135M135MF320.6 GBV100$0.088/hr
SmolLM-135M-Instruct135MBF160.3 GBRTX 4070 Super$0.121/hr
SmolLM-360M-Instruct362MBF160.8 GBRTX 4070 Super$0.121/hr
SmolLM-1.7B1.7BF327.7 GBV100$0.088/hr

Fine-tunes and community models (1)

ModelParametersPublished asVRAM neededLive GPU fitEst. $/hr
SmolLM-135M-Instruct-FP32135MF320.6 GBV100$0.088/hr

VRAM is for the precision the model is published in: the weights times a flat 1.2 overhead, with the KV cache not included. See the methodology. The fit shown is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

More from Hugging Face

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.