Whisper large-v3 — Speech Recognition AI Model

Feed a list of clips and Whisper large-v3 returns one timed transcript each, with clip-local word timestamps matched back by index. Around 99 languages, and the audio is pulled from every video for you.

Details

  • CategorySpeech Recognition
  • Year2026
  • LicenseApache-2.0

Compliance & Provenance

  • ProviderOpenAI (open-source) · Specialized
  • EU AI Act RiskMinimal Risk
  • Art. 50 TransparencyNot applicable

Inputs & Outputs

  • ClipsInput · list:video

    Video clips whose speech is transcribed (audio is extracted per clip).

  • TranscriptsOutput · list:transcript

    One timed-transcript JSON per clip: language, text, and clip-local per-word start/end seconds.

Tags

  • asr
  • speech-to-text
  • transcription
  • timestamps
  • subtitles
  • whisper
  • list
  • shorts

Alternatives in Audio Understanding

Resources