Wan2.2-TI2V-5B — Video Generation AI Model

Wan2.2 TI2V 5B handles both jobs from one checkpoint — a prompt alone yields a five-second clip, a prompt plus an image makes that image frame one. Native 720P at 24fps on a single 4090-class card.

Details

  • CategoryVideo Generation
  • Year2026
  • LicenseApache-2.0

Compliance & Provenance

  • ProviderAlibaba (open weights) · Specialized
  • EU AI Act RiskLimited Risk
  • Art. 50 TransparencyRequired — AI-generated outputs are marked

Inputs & Outputs

  • TextInput · string

    Text prompt for video generation

  • ImageInput · image · optional

    Optional first-frame image. When provided, the image becomes the first frame of the generated video and the model animates subsequent motion guided by the text prompt (image-to-video mode). When omitted, the model generates the entire video from the prompt alone (text-to-video mode). The first frame is locked to this image throughout denoising — middle/last-frame conditioning is not supported.

  • VideoOutput · video

    Generated video

Tags

  • video-generation
  • text-to-video
  • diffusion
  • wan2.2
  • ti2v

Alternatives in Video Generation

  • Cosmos3 Nano

    NVIDIA world model for text- and image-to-video, with optional synced audio.

  • Helios Base

    On-device text-to-video, up to 60s @ 24fps. 3-10 min.

  • Helios Distilled

    Fast distilled text-to-video, fixed 640x384. 1-5 min.

  • MiniMax-H3 Ref2VA

    MiniMax-H3 Ref2VA (omni-reference) generates a 5-15s 24fps video with native stereo audio from a prompt plus a reference image and a reference clip, that clip's own soundtrack included. Video and its soundtrack come out of one denoising loop, so lip movement and sound land in sync. ~144GB bf16 weights — offloaded component-by-component, minutes-scale per clip.

Resources