What it is
It transcribes well, and it writes text when there is nothing to transcribe
The interface is as simple as it sounds. Audio goes in, text comes out, and with a flag you can ask for timestamps, or for English text from speech in another language.
Running a small Whisper checkpoint over three clips of public-domain speech, and then over ten seconds of nothing, shows both sides of it.

Two things in that run are worth carrying into any system you build.
The transcripts are formatted. Whisper writes punctuation, capitalisation and "Mr." for the spoken word mister, while the LibriSpeech reference is bare uppercase. A word error rate computed without normalising both sides will report failures that are only formatting.
The silence rows are the failure that bites in production. Given ten seconds of digital silence the model wrote "you", and given ten seconds of quiet noise it wrote a long run of repeated Georgian characters. Neither output is marked as uncertain.
Where it shows up
Where Whisper shows up
Meeting and call transcription
A common use. Audio arrives as a file, Whisper returns text with timestamps, and everything downstream treats it as a document.
Subtitles and captions
The timestamp output lines up text with the video, and the multilingual training means one model covers a catalogue in many languages.
Voice interfaces
Whisper turns the user's speech into text, an LLM decides what to do with it, and a text-to-speech model answers. Whisper wasn't built for live audio, so a voice interface either waits until the user stops talking and sends the whole utterance, or runs a streaming wrapper, which the section on 30-second windows covers. The newer speech-to-speech models collapse those three steps into one, which the speech-to-speech entry on this index covers.
Search over audio archives
Transcribe once, then index the text with embeddings. The audio becomes searchable without anyone listening to it.
Translating speech into English
Whisper has a translate mode that writes English text from speech in another language, in the same pass. It only translates into English, and it needs a multilingual checkpoint like large-v3 or medium. OpenAI's repository says the turbo model "is not trained for translation tasks" and returns the original language even when translation is requested, so a pipeline that switched to turbo for speed can stop translating without any error.
How it works
An audio encoder and a text decoder, in 30-second windows
The audio becomes a spectrogram, and the encoder reads it
The waveform arrives as a long list of numbers, 16,000 of them for every second of sound, and not one of them is a word. Whisper's first move is to turn that list into a log-Mel spectrogram, which is a picture of the sound with time running left to right and frequency running bottom to top, where each point says how much energy there was at that frequency at that moment. A voice shows up in it as a stack of horizontal bands, and a gap between words shows up as a pale column.
The encoder reads that picture. Two convolutions run across it and the second one steps two columns at a time, so the 3,000 columns of a 30-second window come out as 1,500 positions, one for every 20 milliseconds. Each position is a list of 768 numbers, and those numbers are the only thing the decoder ever gets of the audio.

The window is a fixed 30 seconds whatever you feed it, so a five-second clip is padded out with silence to fill the rest. That padding is why a clip with speech in it and a clip with nothing in it are the same shape to the encoder. Both arrive as 3,000 columns, both come out as 1,500 vectors, and the decoder is asked to write text from either one.
The decoder writes text conditioned on the whole clip
The decoder is an ordinary autoregressive text decoder with cross-attention into the encoder's output. Special tokens at the start tell it which language to use and whether to transcribe or translate, so one model covers all of it.
It works in 30-second windows
Whisper's input is a fixed 30-second window. Longer audio is split into chunks, each transcribed in turn, and the results stitched together. Chunk boundaries are where repeated or dropped words show up, and they are the reason different Whisper libraries give different results on the same long file.
The window also means Whisper has no streaming mode of its own. It writes text for a whole window after the window is full, and live captions need words while the person is still talking. Streaming wrappers work around this by running Whisper again and again on a growing buffer of recent audio and only showing words once two runs in a row agree on them. Macháček et al. (2023) built one this way, called Whisper-Streaming, and report 3.3 seconds of latency on long-form speech. That delay is fine for live subtitles and slow for a voice agent that has to answer quickly, where a model built for streaming is the better start.
The training bought robustness with scale
Radford et al. (2022) trained on 680,000 hours of weakly supervised audio collected from the web, meaning audio paired with transcripts that already existed online, and that scale is where the robustness to accents and noise comes from. Earlier systems trained on small curated corpora, with transcripts made for the purpose.
That training choice also explains the hallucinations. The model learned to produce plausible text for audio, and silence is audio, so it produces plausible text for silence.
The library decides the speed more than the checkpoint does
Much production use runs through a reimplementation of Whisper, in place of OpenAI's original code. faster-whisper uses CTranslate2, whisper.cpp targets CPUs, and WhisperKit targets Apple silicon. The downloads on Hugging Face show it plainly, with argmaxinc/whisperkit-coreml at 11.0 million a month and Systran/faster-whisper-small at 2.9 million.
Versions
The versions you'll see, as of September 2026
| Model | Released | License | Downloads a month |
|---|---|---|---|
argmaxinc/whisperkit-coreml | February 2024 | Check the card | 11.0M |
openai/whisper-large-v3-turbo | October 2024 | MIT | 6.6M |
openai/whisper-large-v3 | November 2023 | MIT | 4.6M |
Systran/faster-whisper-small | November 2023 | MIT | 2.9M |
openai/whisper-small | September 2022 | MIT | 2.9M |
nvidia/parakeet-tdt-0.6b-v2 | April 2025 | CC-BY-4.0 | 0.19M |
nvidia/canary-qwen-2.5b | June 2025 | CC-BY-4.0 | 0.03M |
whisper-large-v3-turbo has overtaken large-v3, and it is the checkpoint to start from for transcription. It is a large-v3 with a much smaller decoder, so it transcribes far faster at a small cost in accuracy. For translation into English, use large-v3 or medium, because turbo wasn't trained for it.
Whisper is no longer the most accurate open model on English. NVIDIA's Canary and Parakeet families sit above it on the Open ASR Leaderboard, and the leaderboard's paper compares 86 systems across 12 datasets, standardising word error rate against inverse real-time factor. Read the leaderboard itself for current figures, because they move and because blog posts about them go stale within weeks.
The download column shows what people run anyway. The NVIDIA models are ahead on accuracy and two orders of magnitude behind on usage, and their CC-BY-4.0 licence is a different commitment from Whisper's MIT.
Choosing
Choosing, and when to reach past Whisper
| The job | Reach for | Why |
|---|---|---|
| General transcription in many languages | whisper-large-v3-turbo | Fast, MIT-licensed, and it handles accents and noise well |
| Speech in another language, written out in English | whisper-large-v3 or medium with the translate task | Turbo wasn't trained for translation and returns the original language |
| Live captions while someone is talking | A streaming wrapper like Whisper-Streaming, or a model built for streaming | Whisper itself writes text a window at a time |
| The best English accuracy you can self-host | An NVIDIA Canary model | It leads the Open ASR Leaderboard, under CC-BY-4.0 |
| Raw throughput on a batch of files | A Parakeet model, or faster-whisper | Built for speed per hour of audio |
| Transcription on a phone or a laptop | WhisperKit or whisper.cpp | The same weights, compiled for the device |
| Knowing who spoke when | A speaker diarization model alongside it | Whisper writes words and never says who said them |
| A voice agent that answers out loud | A speech-to-speech model | One model in place of three, with far lower latency |
| Audio that is music or sound, with no speech in it | An audio understanding model | Whisper is trained to transcribe speech and nothing else |
Start with turbo unless you have a reason not to. Move to Canary when English accuracy is the binding constraint and the licence suits you, and move to a device-targeted build when the audio cannot leave the machine.
Try it
How to try it
Transcription is one call once the audio is in the shape the model expects, which is one channel at 16,000 samples a second. Audio files are often stereo at 44,100 or 48,000, so the loader converts both.
import librosa
import torch
from transformers import WhisperForConditionalGeneration, WhisperProcessor
model_id = "openai/whisper-large-v3-turbo" # MIT, the one to start from
proc = WhisperProcessor.from_pretrained(model_id)
model = WhisperForConditionalGeneration.from_pretrained(model_id).eval()
audio, sr = librosa.load("clip.wav", sr=16000, mono=True) # resample, downmix
feats = proc(audio, sampling_rate=sr, return_tensors="pt").input_features
# input_features covers the first 30 seconds only, the rest is cut off
with torch.no_grad():
ids = model.generate(feats, language="en", max_new_tokens=200)
print(proc.batch_decode(ids, skip_special_tokens=True)[0])Each step in that loader guards against a specific failure in the transformers feature extractor. A sampling rate other than 16,000 raises an error, so at least that one is loud. A stereo file loaded as a two-column array gets read as a batch, one tiny clip per row, and anything past 30 seconds is cut off without a warning, so a 10-minute file comes back as its first half minute of text.
For anything longer than 30 seconds, use faster-whisper, which handles the chunking and the timestamps. Its voice activity filter, the step that skips audio with no speech in it, is off by default in WhisperModel.transcribe and needs vad_filter=True. The batched pipeline, BatchedInferencePipeline, turns it on by default.
Measure Whisper on my own audio before I trust it. Take a folder of audio files with a CSV of reference transcripts. Load every file as mono 16 kHz audio and print any file that needed resampling or downmixing. Transcribe each file with the turbo model through faster-whisper, and compute word error rate after normalising both sides, which means lowercasing, stripping punctuation and expanding common abbreviations, then report the WER with and without that normalisation so I can see how much of the error is formatting. Report the real-time factor per file. Then run the same files with the voice activity filter turned off and on, and list any segment where the two runs differ, since those are where the model is writing over silence.
Limits
What it can't do
- It writes text over audio that has no speech in it. The run above produced the word "you" from ten seconds of silence and a run of repeated characters from quiet noise. Use a voice activity detector in front of it and drop segments it marks as silent.
- It doesn't tell you who spoke. Diarization is a separate model, and running the two together is the normal setup.
- It has no streaming mode. Live transcription needs a wrapper that re-runs it on a growing buffer, which adds a few seconds of delay.
- Turbo doesn't translate. Translation into English needs
large-v3ormedium, and turbo returns the original language without an error. - Long audio is chunked, and the seams show. The 30-second window means repeated or dropped words can appear at boundaries, and different libraries stitch them differently.
- Timestamps are approximate. Word-level timestamps come from extra machinery on top, and they drift, which is worth checking before you build subtitle alignment on them.
- It is no longer the accuracy leader. Current NVIDIA models score better on English on the Open ASR Leaderboard, which is worth knowing when transcription quality is the product.
- Quality varies a lot by language. The multilingual training is uneven, so a language with little training data gets noticeably worse results, and the only way to know is to test on your own recordings.
Go deeper
Radford et al. (2022): Robust Speech Recognition via Large-Scale Weak Supervision · Srivastav et al. (2025): Open ASR Leaderboard, reproducible multilingual and long-form evaluation · The Open ASR Leaderboard itself, which moves week to week · OpenAI: the whisper-large-v3-turbo model card · Vaswani et al. (2017): Attention Is All You Need, the encoder-decoder Whisper is built from · Macháček et al. (2023): Turning Whisper into Real-Time Transcription System · OpenAI's Whisper repository, with the note on turbo and translation
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
