Wan 2.1 guide: checkpoints, setup & current alternatives


Quick answer
Wan 2.1 is a downloadable Apache-2.0 video-model family with separate text-to-video and image-to-video checkpoints. Use the official Wan 2.1 repository for commands, files and hardware notes. It is still useful for reproducible older workflows, but Wan 2.2 is the current Wan release and LTX 2.5 or MiniMax H3 may be stronger starting points for a new open-weight deployment.
Status checked September 13, 2026. This guide does not declare Wan 2.1 the best model or promise a universal VRAM floor, render time or quality result. Those depend on the checkpoint, resolution, frame count, dtype, offloading, attention implementation and hardware.
Choose the correct Wan 2.1 checkpoint
T2V-1.3B: the smaller text-to-video route. The repository says it supports 480p and can run on consumer GPUs with 8.19 GB VRAM; treat that as the project’s documented configuration, not a guarantee for every environment.
T2V-14B: the larger text-to-video checkpoint for 480p or 720p generation.
I2V-14B-480P or I2V-14B-720P: image-to-video checkpoints. Match the checkpoint to the intended output resolution.
FLF2V-14B-720P: first-and-last-frame guidance at 720p.
VACE-1.3B or VACE-14B: video creation and editing checkpoints. Follow their task-specific examples rather than assuming the T2V command applies.
Install and run the official repository
The project currently requires Python 3.10–3.12 and PyTorch 2.4 or newer. Follow the repository’s CUDA and FlashAttention notes for your machine. The minimal repository setup is:
git clone https://github.com/Wan-Video/Wan2.1.git
cd Wan2.1
pip install -r requirements.txtDownload only the checkpoint you intend to run from the official Hugging Face or ModelScope links in the repository. Then copy the matching command from the repository. Do not mix a 1.3B command, a 14B checkpoint and unrelated community settings.
A reproducible first test
Start with the repository’s documented size, frame count and sample command for the selected checkpoint.
Save the checkpoint name and revision, repository commit, prompt, negative prompt, seed, size, frame count, solver, sampling steps and guidance scale.
For image-to-video, preserve the exact source image and crop. A changed crop is a changed test.
Generate at least three attempts before changing one variable. Record runtime, peak memory, failures and accepted outputs.
Review subject identity, anatomy, object permanence, camera motion, background stability, text and looping artifacts at full size.
Prompt Wan 2.1 without invented guarantees
Describe one shot: subject, action, environment, camera and lighting. For image-to-video, focus on desired motion and what must remain stable. Begin with one subject and one main action. Add complexity only after the baseline works.
A red ceramic mug sits on a dark wooden table beside a window. Steam rises slowly while the camera makes a gentle five-second push-in. Overcast daylight, fixed mug shape and logo, no hands entering frame.
A prompt cannot force a checkpoint to preserve every detail. If identity, a product label or composition must be exact, use a source image and inspect every frame. Our video prompt guide explains how to state motion and camera direction across current models.
When to use Wan 2.1 now
Keep it when an existing workflow, LoRA, benchmark or dependency is pinned to Wan 2.1.
Use it when the 1.3B route fits constrained local experimentation and you can validate the project’s documented memory configuration on your hardware.
Choose a current model when starting a new deployment and you do not need 2.1 compatibility.
Choose a managed API or web product when GPU setup, model downloads, queues and monitoring cost more than the control of self-hosting.
Current alternatives
Wan 2.2 is the current Wan family repository. It includes T2V and I2V 14B models plus a 5B text-and-image-to-video route documented for 720p at 24 fps. Move when you want the maintained Wan generation rather than 2.1 compatibility.
LTX 2.5 is a current open audio-video family with official inference tooling. MiniMax H3 publishes downloadable base checkpoints for text/keyframe and multimodal reference audio-video workflows. Compare licenses, hardware, native audio, supported inputs and the complete local-plus-hosted pipeline before selecting either.
The current open-source-friendly video guide compares those routes with managed APIs. Magic Hour’s model catalog lists available hosted models; check the live picker and credit estimate before assuming a particular version is exposed.
Try current video models in one workflow
Use a managed video workflow when you want to compare current models without downloading checkpoints or configuring a local GPU stack.
Explore AI ModelsBrowser alternatives
Use Text to Video when you want to start from a prompt, or Image to Video when a source image defines the subject or composition. Retain the same evaluation record you would use locally: selected model, prompt, source, duration, resolution, attempts and accepted result.
Frequently asked questions
The official code and checkpoints are published under Apache-2.0. Compute, storage, bandwidth and any hosted service still cost money. Review the repository license and the rights to your inputs before commercial use.
The repository reports 8.19 GB VRAM for its T2V-1.3B consumer-GPU configuration. That is specific to the documented route. Larger checkpoints and different settings require more memory; measure your own configuration.
Yes. The project publishes separate 14B image-to-video checkpoints for 480p and 720p plus a first/last-frame route. Select the task-specific command and checkpoint.
Use 2.1 for compatibility with an existing workflow or benchmark. Start a new Wan evaluation with 2.2, then keep 2.1 only if it wins on your hardware, output requirements or dependencies.
There is no universal winner. Wan 2.1 is an older family. Compare current Wan 2.2, LTX 2.5 and MiniMax H3 with the same inputs, retained outputs and accepted-output cost.





