Cosmos3-Nano — Video Generation AI Model

Cosmos3 Nano is NVIDIA's 16B world model, built to make video that obeys physics rather than merely looking plausible. Text in, or an image to anchor frame zero. Optional synced audio — describe the sound in the prompt.

Details

  • CategoryVideo Generation
  • Year2026
  • LicenseOpenMDW-1.1

Compliance & Provenance

  • ProviderNVIDIA (open) · Specialized
  • EU AI Act RiskLimited Risk
  • Art. 50 TransparencyRequired — AI-generated outputs are marked

Inputs & Outputs

  • TextInput · string

    Describe the scene and motion in one plain paragraph. With an input image, frame 0 carries the look and this text drives what happens over time; without one, the text alone sets both look and motion (text-to-video). When sound is on, mention the audio you want too.

  • ImageInput · image · optional

    Optional first frame. When provided, frame 0 is anchored to it (image-to-video); when omitted, the whole clip is generated from the text prompt alone (text-to-video).

  • VideoOutput · video

    Generated video

Tags

  • video-generation
  • text-to-video
  • image-to-video
  • diffusion
  • world-model
  • physical-ai
  • cosmos3
  • nvidia

Alternatives in Video Generation

  • Helios Base

    On-device text-to-video, up to 60s @ 24fps. 3-10 min.

  • Helios Distilled

    Fast distilled text-to-video, fixed 640x384. 1-5 min.

  • MiniMax-H3 Ref2VA

    MiniMax-H3 Ref2VA (omni-reference) generates a 5-15s 24fps video with native stereo audio from a prompt plus a reference image and a reference clip, that clip's own soundtrack included. Video and its soundtrack come out of one denoising loop, so lip movement and sound land in sync. ~144GB bf16 weights — offloaded component-by-component, minutes-scale per clip.

  • Wan2.2 TI2V 5B

    Wan2.2 dense 5B unified text+image-to-video. Text only → 5s video from prompt; text + image → image as the first frame of the video.

Resources