gemini-3.5-flash — LLM AI Model

Hand Gemini a video and get text back: a transcript, MM:SS timecodes, chapter summaries, or highlight candidates, depending what you ask. It watches and listens rather than sampling frames. Google AI key required.

Details

  • CategoryLLM
  • Year2026
  • LicenseGoogle Cloud Terms

Compliance & Provenance

  • ProviderGoogle · GPAI
  • EU AI Act RiskMinimal Risk
  • Art. 50 TransparencyNot applicable

Inputs & Outputs

  • VideoInput · video

    Video to analyze

  • TextInput · string · optional

    Optional instruction (e.g. 'transcribe with timecodes' or 'find the highlights')

  • AnalysisOutput · segments

    Timecoded highlight segments (feeds the Video Trim node)

Tags

  • llm
  • multimodal
  • video-understanding
  • transcription
  • reasoning

Alternatives in Image/Video Models

  • GPT Image

    OpenAI GPT-Image text-to-image / edit. Needs API key.

  • Nano Banana

    Google Gemini-based image gen/edit with multimodal reasoning. Needs API key.

  • Omni

    Google Gemini Omni: text/image/video in, video out, with conversational editing. Which input ports apply depends on Task — Text to Video: Text only. Image to Video: + Image 1. Reference to Video: + Image 1–3. Edit: + Video (aspect ratio is ignored). Unused ports are dropped even if wired. Needs API key.

  • Sora 2

    OpenAI Sora 2 text-to-video / img2video. Needs API key.

  • VEO

    Google VEO 3.1 video gen (up to 4K, with native audio). Needs API key.

Resources