Whisper large-v3 — Speech Recognition AI Model

Whisper large-v3 transcribes the speech in a video and returns word-level timestamps rather than one block of text — which is what karaoke-style captions need. Around 99 languages, and it extracts the audio for you.

Details

  • CategorySpeech Recognition
  • Year2026
  • LicenseApache-2.0

Compliance & Provenance

  • ProviderOpenAI (open-source) · Specialized
  • EU AI Act RiskMinimal Risk
  • Art. 50 TransparencyNot applicable

Inputs & Outputs

  • VideoInput · video

    Video clip whose speech is transcribed (audio is extracted internally).

  • TranscriptOutput · transcript

    Timed-transcript JSON: language, full text, and per-word start/end seconds.

Tags

  • asr
  • speech-to-text
  • transcription
  • timestamps
  • subtitles
  • whisper

Alternatives in Audio Understanding

Resources