Best open video generation models

How to choose an open text-to-video or image-to-video model: resolution, frames per second, clip length, audio, and license, with picks by situation.

Video generation is the most memory-hungry common GPU task, so choose by resolution, clip length and frame rate first, then by license. A model that makes five seconds at 720p needs a very different machine than one that makes a short clip at low resolution.

What matters for this task

Resolution and frame rate. Each model is trained for particular sizes (for example 480p or 720p) and a fixed frames-per-second. Cards say whether other sizes are supported; some say they are not.

Clip length. Length is frames divided by fps. A 5 second clip at 24 fps is 120 frames, and memory and time grow with frame count. Most open models make clips of a few seconds.

Text-to-video, image-to-video, or both. Image-to-video animates a still you supply, which gives far more control. Some models do both from one checkpoint.

Audio. Most open models make silent video. A few generate synchronized sound.

Architecture and size. Mixture-of-experts video models carry more total parameters than they use per step. Check the card for whether a single GPU is enough or multi-GPU is expected. See mixture of experts.

License. Several popular video models use custom licenses with usage terms. Read them.

Picks by situation

One GPU, 720p, Apache 2.0. Wan2.2 TI2V 5B (Wan-AI, Apache 2.0, about 5B parameters) handles both text-to-video and image-to-video at 720p and 24 fps (1280 by 704). The card says it can run on a single consumer GPU such as the 4090.

Highest quality from the Wan line. Wan2.2 T2V A14B (Apache 2.0, about 14.3B parameters in the catalog) uses a mixture-of-experts design and generates 5 second videos at 480p and 720p. The card's single-GPU command is listed for a GPU with at least 80GB, with offload options if memory runs out.

Video with synchronized audio. LTX-2 (Lightricks, about 18.9B parameters in the catalog) is a DiT-based model that generates video and audio together in a single model. Its card example generates 121 frames at 24 fps and lists spatial and temporal upscalers for higher resolution or frame rate. It uses the LTX-2 Community License Agreement, so read the terms.

Small, fixed-format clips. CogVideoX-5b (Z.ai, about 5.6B parameters) produces 6 second clips at 8 fps and 720 by 480, and the card says other resolutions are not supported. It has its own CogVideoX license.

Licenses at a glance

Apache 2.0: both Wan2.2 models. LTX-2 Community License Agreement: LTX-2. CogVideoX License: CogVideoX-5b.

Running these on Aquanode

Rent a GPU by the hour for the memory your resolution needs. See pricing and the VRAM calculator. The table below lists all video generation models in the catalog.

Open models for video generation

All 31 models in the catalog for this task, grouped by size.

Under 3B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
LTX-Video1.9B8.6 GB–V100$0.088/hr
Wan2.1-T2V-1.3B-Diffusers1.4B6.3 GB–V100$0.088/hr
stable-video-diffusion-img2vid-xt1.5B6.8 GB–V100$0.088/hr
stable-video-diffusion-img2vid1.5B6.8 GB–V100$0.088/hr
Wan2.1-T2V-1.3B1.4B6.3 GB–V100$0.088/hr
CogVideoX-2b1.7B3.8 GB–V100$0.088/hr
text-to-video-ms-1.7b1.4B6.3 GB–V100$0.088/hr
LTX-Video-0.9.51.9B4.3 GB–RTX 4070 Super$0.121/hr
Wan2.1-VACE-1.3B-diffusers2.2B9.6 GB–V100$0.088/hr
stable-virtual-camera1.3B5.7 GB–V100$0.088/hr

3B to 10B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Wan2.2-TI2V-5B-Diffusers5.0B22.4 GBApache 2.0RTX A5000$0.176/hr
FastWan2.2-TI2V-5B-FullAttn-Diffusers5.0B11.2 GB–RTX 4070 Super$0.121/hr
CogVideoX-5b5.6B12.5 GB–RTX A4000$0.167/hr
CogVideoX-5b-I2V5.6B12.6 GB–RTX A4000$0.167/hr

10B to 40B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
LTX-218.9B42.2 GBCustom licenseRTX A6000$0.363/hr
Wan2.2-T2V-A14B-Diffusers14.3B63.9 GBApache 2.0A100$1.21/hr
Wan2.2-I2V-A14B-Diffusers14.3B63.9 GBApache 2.0A100$1.21/hr
Wan2.1-I2V-14B-720P-Diffusers16.4B73.3 GB–A100$1.21/hr
Wan2.1-T2V-14B-Diffusers14.3B63.9 GB–A100$1.21/hr
Wan2.1-I2V-14B-480P-Diffusers16.4B73.3 GB–A100$1.21/hr
LTX-2.5-Diffusers19.0B42.4 GB–RTX A6000$0.363/hr
Wan2.1-I2V-14B-480P16.4B73.3 GB–A100$1.21/hr
Wan2.1-T2V-14B14.3B63.9 GB–A100$1.21/hr
Wan2.1-I2V-14B-720P16.4B73.3 GB–A100$1.21/hr
Wan2.2-S2V-14B16.3B36.4 GB–RTX A6000$0.363/hr
lingbot-world-fast18.5B82.9 GB–RTX PRO 6000$1.38/hr
Cosmos-H-Surgical15.2B33.9 GB–RTX A6000$0.363/hr
Wan2.1-VACE-14B17.3B77.5 GB–A100$1.21/hr
mochi-1-preview10.0B44.8 GBApache 2.0RTX A6000$0.363/hr
Kandinsky-6.0-Pro-5s-Diffusers30.1B67.4 GB–A100$1.21/hr

40B to 150B parameters

ModelParametersVRAM neededLicenseCheapest live fitEst. $/hr
Cosmos3-Super-Image2Video64.6B144 GB–RTX A5000 × 7$1.23/hr

VRAM is for the precision each model is published in, with the same overhead and an 8,192-token context assumed on every page; see the methodology. The license column shows the license where our catalog records one. The fit is the lowest-priced single GPU type that holds the model at that precision, or the lowest-priced multi-GPU set (up to 8) when none does.

Sources

Updated 2026-10-07.

More ways to choose a model

By GPU memory:

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.