Wan: Alibaba's open video generation models
Wan from Alibaba: Wan2.1 (1.3B and 14B) and Wan2.2 (MoE A14B, 5B, S2V), their tasks, licenses, and which generation to run for which job.
What Wan is
Wan is a family of open video generation models published by Alibaba under the Wan-AI organization on Hugging Face. Both generations are released under the Apache 2.0 license according to their GitHub READMEs, which makes the line one of the more permissive options for video. As of the Wan-AI organization page we read on 2026-10-07, no model newer than Wan2.2 is listed there, so Wan2.2 is the current generation.
Generations in order
Wan2.1 (February 2025)
Wan2.1 was released on February 25, 2025, with inference code and weights. It was added to Diffusers on March 3, 2025. The model sizes and tasks in the README are:
- Text-to-video (T2V): 1.3B and 14B, at 480P and 720P.
- Image-to-video (I2V): 14B, at 480P and 720P.
- First-last-frame-to-video (FLF2V): 14B at 720P, introduced April 17, 2025.
- VACE (video creation and editing): 1.3B and 14B, released May 14, 2025.
The README highlights that the 1.3B text-to-video model needs 8.19 GB of VRAM at 480P, and that the models can generate both Chinese and English text inside videos. The 1.3B size is the reason Wan2.1 stays relevant: it is the smallest checkpoint in the line.
Wan2.2 (July 2025)
Wan2.2 was released on July 28, 2025. Its main change is a Mixture-of-Experts design in the two flagship checkpoints, T2V-A14B and I2V-A14B. These have 27B parameters in total with 14B active per step. The README describes two experts: a high-noise expert that handles the early denoising stages and overall layout, and a low-noise expert that refines detail later. Both support 480P and 720P.
Wan2.2 also added:
- TI2V-5B: a 5B model that takes text or an image and generates 720P at 24 fps. The README says it runs on consumer GPUs such as an RTX 4090.
- S2V-14B (August 26, 2025): speech-to-video, audio-driven generation at 480P and 720P.
- Animate-14B (September 19, 2025): character animation and replacement.
The Wan-AI organization page also lists later Animate variants and a Wan-Dancer-14B checkpoint, all within the Wan2.2 naming. We did not open those model cards, so we say nothing further about them here.
Which one to use today
- Smallest footprint, text-to-video: Wan2.1 T2V 1.3B. It is the lowest-parameter option in the line and Apache 2.0.
- One checkpoint for text and image input on a single consumer card: Wan2.2 TI2V-5B, which the README positions for an RTX 4090 class GPU.
- Best fidelity from the open Wan weights: Wan2.2 T2V-A14B or I2V-A14B. Compute per step is that of a 14B model, but all 27B parameters need to be stored, so plan memory for the total, which the model pages compute for your hardware.
- Animating a still image: Wan2.2 I2V-A14B, or Wan2.1 I2V 14B if you need to stay with the first-generation pipeline.
- Video editing or controlled generation: Wan2.1 VACE (1.3B or 14B); the README lists it as the video creation and editing model.
- Audio-driven video: Wan2.2 S2V-14B.
- First and last frame control: Wan2.1 FLF2V, a 14B model at 720P.
Pick Wan2.2 when you start a new project; pick Wan2.1 when you need VACE or FLF2V, or the 1.3B text-to-video model, since those task models are listed under the first generation in the READMEs we read.
Side lines
Wan has no separate side line in this series; the task-specific models above (VACE, FLF2V, S2V, Animate) sit inside the two generations. Community distillations and LoRA merges of Wan are indexed under the generation hubs, Wan2.1 and Wan2.2.
Running Wan
Wan2.1 was merged into Diffusers in March 2025, and the Wan-AI organization publishes Diffusers-format checkpoints for Wan2.2. You can run these models on Aquanode GPUs.
Wan generations
Every Wan generation we track, oldest first. Each hub lists all of its models with VRAM at native, FP8 and INT4 precision.
| Hub | Models | Sizes | Smallest native VRAM | Cheapest live fit for it | Est. $/hr |
|---|---|---|---|---|---|
| Wan2.1 | 11 | 1.4B to 17.3B | 6.3 GB | V100 | $0.088/hr |
| Wan2.2 | 7 | 5.0B to 16.3B | 22.4 GB | RTX A5000 | $0.176/hr |
Smallest native VRAM is the lowest requirement among the publisher's own checkpoints in the hub, at the precision they are published in; see the methodology.
Best models by task
Wan models appear on these ranked task pages.
Sources
Facts in the text above were read from these pages and are the publisher's own statements, not benchmarks run by Aquanode. Last reviewed 2026-10-07.