Whisper large-v3 — Speech Recognition AI Model
Feed a list of clips and Whisper large-v3 returns one timed transcript each, with clip-local word timestamps matched back by index. Around 99 languages, and the audio is pulled from every video for you.
Details
- CategorySpeech Recognition
- Year2026
- LicenseApache-2.0
Compliance & Provenance
- ProviderOpenAI (open-source) · Specialized
- EU AI Act RiskMinimal Risk
- Art. 50 TransparencyNot applicable
Inputs & Outputs
- ClipsInput · list:video
Video clips whose speech is transcribed (audio is extracted per clip).
- TranscriptsOutput · list:transcript
One timed-transcript JSON per clip: language, text, and clip-local per-word start/end seconds.
Tags
Alternatives in Audio Understanding
- Speech Recognition (Video)
Recognize speech in a video and convert it to text.