veo-3.1-generate-preview — Video Generation AI Model

VEO 3.1 makes 4, 6 or 8 second clips up to 4K with audio generated in sync — from text, from an image, between a first and last frame, or guided by a reference. Runs on Google's side, against your API key.

Details

  • CategoryVideo Generation
  • Year2026
  • LicenseGoogle Cloud Terms

Compliance & Provenance

  • ProviderGoogle · GPAI
  • EU AI Act RiskLimited Risk
  • Art. 50 TransparencyRequired — AI-generated outputs are marked

Inputs & Outputs

  • TextInput · string

    Text prompt for video generation

  • Image 1Input · image · optional

    Source image, first frame, or reference image 1

  • Image 2Input · image · optional

    Last frame or reference image 2

  • Image 3Input · image · optional

    Reference image 3 (Reference Image mode only)

  • VideoOutput · video

    Generated video with native audio

Tags

  • video-generation
  • text-to-video
  • image-to-video
  • frame-interpolation
  • multimodal
  • audio-synthesis

Alternatives in Image/Video Models

  • GPT Image

    OpenAI GPT-Image text-to-image / edit. Needs API key.

  • Nano Banana

    Google Gemini-based image gen/edit with multimodal reasoning. Needs API key.

  • Omni

    Google Gemini Omni: text/image/video in, video out, with conversational editing. Which input ports apply depends on Task — Text to Video: Text only. Image to Video: + Image 1. Reference to Video: + Image 1–3. Edit: + Video (aspect ratio is ignored). Unused ports are dropped even if wired. Needs API key.

  • Sora 2

    OpenAI Sora 2 text-to-video / img2video. Needs API key.

  • Video Analysis

    Gemini video analysis. Feed a video and get a transcript, timecodes, or highlight analysis as text. Needs API key.

Resources