Whisper large-v3 — Speech Recognition AI Model
Whisper large-v3 transcribes the speech in a video and returns word-level timestamps rather than one block of text — which is what karaoke-style captions need. Around 99 languages, and it extracts the audio for you.
Details
- CategorySpeech Recognition
- Year2026
- LicenseApache-2.0
Compliance & Provenance
- ProviderOpenAI (open-source) · Specialized
- EU AI Act RiskMinimal Risk
- Art. 50 TransparencyNot applicable
Inputs & Outputs
- VideoInput · video
Video clip whose speech is transcribed (audio is extracted internally).
- TranscriptOutput · transcript
Timed-transcript JSON: language, full text, and per-word start/end seconds.
Tags
Alternatives in Audio Understanding
- Speech Recognition (Clips)
Recognize speech in each clip of a video list and convert it to timed transcripts.