Best open image generation models
How to choose an open text-to-image model: step count, text rendering, parameters, and license for commercial use, with picks by situation.
For image generation, license and sampling steps decide more than raw quality. Several top open models are free to try but restrict commercial use, and the number of denoising steps sets how many images a GPU makes per hour.
What matters for this task
Sampling steps. A diffusion model removes noise over many steps. Base models often use 30 to 50; distilled "turbo" or "schnell" variants cut this to a handful, which multiplies throughput. The card states the step count (or NFEs, number of function evaluations) the model was designed for.
Parameter count. Image models run from under 1B to over 20B parameters in the catalog. Larger models usually follow prompts better but need more memory per image. See VRAM and quantization for ways to shrink them.
Text in images. Rendering legible words (signs, posters, UI) separates models sharply. Check whether the card claims bilingual or multilingual text rendering.
Resolution. Models are trained at a native resolution; going far beyond it degrades results.
Editing. Some releases handle editing from reference images as well as generation.
License. Check commercial terms before building a product. This is the most common reason to rule a model out.
Picks by situation
Fast and openly licensed. Z-Image-Turbo (Tongyi-MAI, Apache 2.0) is a distilled 6B model that the card says needs only 8 NFEs, with accurate English and Chinese text rendering and sub-second latency on H800-class GPUs.
Fewest steps, Apache 2.0. FLUX.1 schnell (Black Forest Labs, about 11.9B parameters in the catalog) is licensed Apache 2.0 per the Hugging Face metadata and the FLUX repository. The Black Forest Labs repository lists it as a text-to-image model.
Highest quality from the FLUX line, non-commercial. FLUX.1 dev (about 11.9B parameters) uses the FLUX.1-dev Non-Commercial License per the Black Forest Labs repository, which also points to a separate paid license for commercial use. Good for research and evaluation, not for a product without that license.
Generation plus editing in one model. Qwen-Image-2.1 (Qwen) is described as a unified text-to-image and editing model with a 7B-parameter visual generation component. Its card sets 40 inference steps in the example code, and the license is the Qwen Research License, so check terms for commercial use.
Mature ecosystem and small size. Stable Diffusion XL base 1.0 (Stability AI, about 2.57B parameters in the catalog) is licensed CreativeML Open RAIL++-M and has the largest set of community add-ons. It is a two-stage design with an optional refiner model.
Licenses at a glance
Apache 2.0: Z-Image-Turbo, FLUX.1 schnell. Non-commercial: FLUX.1 dev. Qwen Research License: Qwen-Image-2.1. CreativeML Open RAIL++-M: SDXL base.
Running these on Aquanode
Rent a GPU by the hour and run the pipeline with diffusers or the vendor code. See pricing and the VRAM calculator. The table below lists all text-to-image models in the catalog.
Open models for image generation
All 40 models in the catalog for this task, grouped by size.
Under 3B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| stable-diffusion-xl-base-1.0 | 2.6B | 11.5 GB | OpenRAIL++ | V100 | $0.088/hr |
| dreamshaper-7 | 860M | 3.8 GB | – | V100 | $0.088/hr |
| sdxl-turbo | 2.6B | 11.5 GB | – | V100 | $0.088/hr |
| stable-diffusion-v1-4 | 860M | 3.8 GB | – | V100 | $0.088/hr |
| RealVisXL_V5.0 | 2.6B | 11.5 GB | OpenRAIL++ | V100 | $0.088/hr |
| sd-turbo | 866M | 3.9 GB | – | V100 | $0.088/hr |
| playground-v2.5-1024px-aesthetic | 2.6B | 11.5 GB | – | V100 | $0.088/hr |
| dreamshaper-8 | 860M | 3.8 GB | – | V100 | $0.088/hr |
| stable-diffusion-3.5-medium | 2.5B | 5.5 GB | – | RTX 4070 Super | $0.121/hr |
| LCM_Dreamshaper_v7 | 860M | 3.8 GB | – | V100 | $0.088/hr |
| Photon_v1 | 860M | 3.8 GB | – | V100 | $0.088/hr |
| stable-diffusion-3-medium-diffusers | 2.1B | 4.7 GB | – | V100 | $0.088/hr |
| NSFW-GEN-ANIME-v2 | 2.6B | 11.5 GB | – | V100 | $0.088/hr |
| LUSTIFY-v2.0 | 2.6B | 11.5 GB | – | V100 | $0.088/hr |
| LUSTIFY-v2.0 | 2.6B | 11.5 GB | – | V100 | $0.088/hr |
| SSD-1B | 1.3B | 6.0 GB | – | V100 | $0.088/hr |
| pornmasterPro_noobV3VAE | 2.6B | 11.5 GB | – | V100 | $0.088/hr |
| PixelModel-v6 | 155M | 0.7 GB | – | V100 | $0.088/hr |
| Counterfeit-V2.5 | 860M | 3.8 GB | – | V100 | $0.088/hr |
| majicMIX_realistic_v7 | 860M | 3.8 GB | – | V100 | $0.088/hr |
3B to 10B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Z-Image-Turbo | 6.2B | 27.5 GB | – | V100 | $0.187/hr |
| Qwen-Image-2.1 | 7.1B | 15.9 GB | – | RTX A4000 | $0.167/hr |
| stable-diffusion-3.5-large | 8.1B | 18.2 GB | – | RTX A5000 | $0.176/hr |
| Chroma1-HD | 8.9B | 19.9 GB | – | RTX A5000 | $0.176/hr |
| Z-Image | 6.2B | 13.8 GB | – | RTX A4000 | $0.167/hr |
| Chroma1-Base | 8.9B | 19.9 GB | – | RTX A5000 | $0.176/hr |
| FIBO | 8.3B | 18.5 GB | – | RTX A5000 | $0.176/hr |
| Juggernaut-Z-Image | 6.2B | 13.8 GB | – | V100 | $0.088/hr |
| Fibo-1.5 | 8.3B | 18.5 GB | – | RTX A5000 | $0.176/hr |
| Ming-Image-0.1-Design | 6.2B | 13.8 GB | – | RTX A4000 | $0.167/hr |
10B to 40B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| FLUX.1-dev | 11.9B | 26.6 GB | FLUX.1 [dev] Non-Commercial License | RTX A6000 | $0.363/hr |
| FLUX.1-schnell | 11.9B | 26.6 GB | Apache 2.0 | RTX A6000 | $0.363/hr |
| Qwen-Image | 20.4B | 45.7 GB | Apache 2.0 | RTX A6000 | $0.363/hr |
| HiDream-I1-Fast | 17.1B | 38.2 GB | – | RTX A6000 | $0.363/hr |
| Qwen-Image-2512 | 20.4B | 45.7 GB | – | RTX A6000 | $0.363/hr |
| Krea-2-Turbo | 12.8B | 28.7 GB | – | RTX A6000 | $0.363/hr |
| Krea-2-Raw | 12.8B | 28.7 GB | Custom license | RTX A6000 | $0.363/hr |
| FLUX.1-Krea-dev | 11.9B | 26.6 GB | – | RTX A6000 | $0.363/hr |
| HiDream-I1-Full | 17.1B | 38.2 GB | MIT | RTX A6000 | $0.363/hr |
40B to 150B parameters
| Model | Parameters | VRAM needed | License | Cheapest live fit | Est. $/hr |
|---|---|---|---|---|---|
| Cosmos3-Super-Text2Image-4Step | 64.0B | 143 GB | – | RTX A5000 × 6 | $1.06/hr |
VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.
Sources
- https://huggingface.co/Tongyi-MAI/Z-Image-Turbo
- https://huggingface.co/Qwen/Qwen-Image-2.1
- https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0
- https://huggingface.co/api/models/black-forest-labs/FLUX.1-schnell
- https://github.com/black-forest-labs/flux
Updated 2026-10-07.
More ways to choose a model
- Best open models for coding
- Best open models for reasoning
- Best open models for chat and assistants
- Best open models for vision-language
- Best open models for OCR and document parsing
- Best open models for speech-to-text
- Best open models for text-to-speech
- Best open models for video generation
- Best open models for embeddings
- Best open models for translation
By GPU memory: