Cosmos3-Nano — Video Generation AI Model
Cosmos3 Nano is NVIDIA's 16B world model, built to make video that obeys physics rather than merely looking plausible. Text in, or an image to anchor frame zero. Optional synced audio — describe the sound in the prompt.
Details
- CategoryVideo Generation
- Year2026
- LicenseOpenMDW-1.1
Compliance & Provenance
- ProviderNVIDIA (open) · Specialized
- EU AI Act RiskLimited Risk
- Art. 50 TransparencyRequired — AI-generated outputs are marked
Inputs & Outputs
- TextInput · string
Describe the scene and motion in one plain paragraph. With an input image, frame 0 carries the look and this text drives what happens over time; without one, the text alone sets both look and motion (text-to-video). When sound is on, mention the audio you want too.
- ImageInput · image · optional
Optional first frame. When provided, frame 0 is anchored to it (image-to-video); when omitted, the whole clip is generated from the text prompt alone (text-to-video).
- VideoOutput · video
Generated video
Tags
Alternatives in Video Generation
- Helios Base
On-device text-to-video, up to 60s @ 24fps. 3-10 min.
- Helios Distilled
Fast distilled text-to-video, fixed 640x384. 1-5 min.
- MiniMax-H3 Ref2VA
MiniMax-H3 Ref2VA (omni-reference) generates a 5-15s 24fps video with native stereo audio from a prompt plus a reference image and a reference clip, that clip's own soundtrack included. Video and its soundtrack come out of one denoising loop, so lip movement and sound land in sync. ~144GB bf16 weights — offloaded component-by-component, minutes-scale per clip.
- Wan2.2 TI2V 5B
Wan2.2 dense 5B unified text+image-to-video. Text only → 5s video from prompt; text + image → image as the first frame of the video.