MLGuerrillaStart with M1 →
Speech and audio·14 min read·Updated 24 September 2026

Whisper

Whisper is OpenAI's open speech-to-text model, and it is the default way to turn audio into text. It transcribes many languages well, and it writes fluent sentences over audio that contains no speech at all.

Whisper takes audio and writes down what was said. OpenAI released it with open weights in 2022, and it became the default answer for transcription because it handles accents, background noise and many languages without any setup.

It is an encoder-decoder. The encoder reads a spectrogram, which is a picture of how the sound's frequencies change over time, and the decoder writes text one token at a time while attending to what the encoder produced. That is the same shape as a translation model, with audio on the input side.

What it is

It transcribes well, and it writes text when there is nothing to transcribe

The interface is as simple as it sounds. Audio goes in, text comes out, and with a flag you can ask for timestamps, or for English text from speech in another language.

Running a small Whisper checkpoint over three clips of public-domain speech, and then over ten seconds of nothing, shows both sides of it.

A figure titled "It transcribes speech well, and it writes words over silence too", from openai/whisper-small at 242M parameters on a Mac CPU, using public-domain LibriSpeech audio. The upper card, three clips of real speech, notes that the reference transcript has no punctuation or casing and that Whisper adds both. The first clip is 5.9 seconds of audio transcribed in 2.6 seconds, giving "Mr. Quilter is the apostle of the middle classes, and we are glad to welcome his gospel." The second is 4.8 seconds transcribed in 1.4 seconds, giving "Nor is Mr. Quilter's manner less interesting than his matter." The third is 12.5 seconds transcribed in 1.7 seconds, giving "He tells us that at this festive season of the year, with Christmas and roast beef looming before us, symbolies drawn from eating and its results occur most readily to the mind." The lower card, outlined in terracotta and titled ten seconds with no speech in them, notes that the same model was given audio containing nothing to transcribe. Given silence it wrote the word "you". Given quiet noise it wrote a long run of repeated Georgian letters. The caption reads: nothing in the output marks the last two as different from the first three.
The reference for the third clip says SIMILES where Whisper wrote symbolies, which is the kind of error a word error rate counts. The last two rows are the kind it does not.

Two things in that run are worth carrying into any system you build.

The transcripts are formatted. Whisper writes punctuation, capitalisation and "Mr." for the spoken word mister, while the LibriSpeech reference is bare uppercase. A word error rate computed without normalising both sides will report failures that are only formatting.

The silence rows are the failure that bites in production. Given ten seconds of digital silence the model wrote "you", and given ten seconds of quiet noise it wrote a long run of repeated Georgian characters. Neither output is marked as uncertain.

Where it shows up

Where Whisper shows up

Meeting and call transcription

A common use. Audio arrives as a file, Whisper returns text with timestamps, and everything downstream treats it as a document.

Subtitles and captions

The timestamp output lines up text with the video, and the multilingual training means one model covers a catalogue in many languages.

Voice interfaces

Whisper turns the user's speech into text, an LLM decides what to do with it, and a text-to-speech model answers. Whisper wasn't built for live audio, so a voice interface either waits until the user stops talking and sends the whole utterance, or runs a streaming wrapper, which the section on 30-second windows covers. The newer speech-to-speech models collapse those three steps into one, which the speech-to-speech entry on this index covers.

Search over audio archives

Transcribe once, then index the text with embeddings. The audio becomes searchable without anyone listening to it.

Translating speech into English

Whisper has a translate mode that writes English text from speech in another language, in the same pass. It only translates into English, and it needs a multilingual checkpoint like large-v3 or medium. OpenAI's repository says the turbo model "is not trained for translation tasks" and returns the original language even when translation is requested, so a pipeline that switched to turbo for speed can stop translating without any error.

How it works

An audio encoder and a text decoder, in 30-second windows

The audio becomes a spectrogram, and the encoder reads it

The waveform arrives as a long list of numbers, 16,000 of them for every second of sound, and not one of them is a word. Whisper's first move is to turn that list into a log-Mel spectrogram, which is a picture of the sound with time running left to right and frequency running bottom to top, where each point says how much energy there was at that frequency at that moment. A voice shows up in it as a stack of horizontal bands, and a gap between words shows up as a pale column.

The encoder reads that picture. Two convolutions run across it and the second one steps two columns at a time, so the 3,000 columns of a 30-second window come out as 1,500 positions, one for every 20 milliseconds. Each position is a list of 768 numbers, and those numbers are the only thing the decoder ever gets of the audio.

A five-step figure titled "The sound becomes a picture, and the encoder reads the picture", from one public-domain LibriSpeech clip run through openai/whisper-small. Step one, the audio as it arrives, plots the real waveform and notes 5.86 seconds, 16,000 samples a second and 93,680 numbers, with the line: a wave of air pressure, nothing in it is split into words yet. Step two, the log-Mel spectrogram, shows the real spectrogram of that clip, 80 frequency bands from 0 Hz to 8 kHz up the side and time across the bottom, with stacked harmonic bands where the voice is and pale gaps between words. Step three, the window is always 30 seconds, draws a bar in which 585 real columns of speech take up the first fifth and the remaining 2,415 columns are padded silence. Step four, the encoder turns the picture into vectors, explains that two convolutions run over the columns and the second steps two at a time, so 3,000 columns come out as 1,500 positions, one every 20 milliseconds and each 768 numbers wide, and lists three real outputs: position 20 at 0.4 seconds beginning -0.17, +0.21, +2.17, +0.07, -3.17; position 200 at 4.0 seconds beginning -0.58, +1.35, +3.06, +2.03, -0.86; and position 1400 at 28.0 seconds beginning -1.13, -1.17, +0.19, -0.57, +0.93. It notes in terracotta that position 1400 is inside the padding and still produces a vector, because silence is a picture too. Step five, the decoder writes text while looking at all 1,500, shows the transcript it produced, which is Mr. Quilter is the apostle of the middle classes, and we are glad to welcome his gospel. The caption reads: the encoder never hears audio, it reads a 30-second picture, and it always has one to read.
The vectors in step four are the real output of whisper-small's encoder on this clip. Position 1400 sits in the padded region, which is worth holding on to when you get to the part about silence.

The window is a fixed 30 seconds whatever you feed it, so a five-second clip is padded out with silence to fill the rest. That padding is why a clip with speech in it and a clip with nothing in it are the same shape to the encoder. Both arrive as 3,000 columns, both come out as 1,500 vectors, and the decoder is asked to write text from either one.

The decoder writes text conditioned on the whole clip

The decoder is an ordinary autoregressive text decoder with cross-attention into the encoder's output. Special tokens at the start tell it which language to use and whether to transcribe or translate, so one model covers all of it.

It works in 30-second windows

Whisper's input is a fixed 30-second window. Longer audio is split into chunks, each transcribed in turn, and the results stitched together. Chunk boundaries are where repeated or dropped words show up, and they are the reason different Whisper libraries give different results on the same long file.

The window also means Whisper has no streaming mode of its own. It writes text for a whole window after the window is full, and live captions need words while the person is still talking. Streaming wrappers work around this by running Whisper again and again on a growing buffer of recent audio and only showing words once two runs in a row agree on them. Macháček et al. (2023) built one this way, called Whisper-Streaming, and report 3.3 seconds of latency on long-form speech. That delay is fine for live subtitles and slow for a voice agent that has to answer quickly, where a model built for streaming is the better start.

The training bought robustness with scale

Radford et al. (2022) trained on 680,000 hours of weakly supervised audio collected from the web, meaning audio paired with transcripts that already existed online, and that scale is where the robustness to accents and noise comes from. Earlier systems trained on small curated corpora, with transcripts made for the purpose.

That training choice also explains the hallucinations. The model learned to produce plausible text for audio, and silence is audio, so it produces plausible text for silence.

The library decides the speed more than the checkpoint does

Much production use runs through a reimplementation of Whisper, in place of OpenAI's original code. faster-whisper uses CTranslate2, whisper.cpp targets CPUs, and WhisperKit targets Apple silicon. The downloads on Hugging Face show it plainly, with argmaxinc/whisperkit-coreml at 11.0 million a month and Systran/faster-whisper-small at 2.9 million.

Versions

The versions you'll see, as of September 2026

Whisper and its alternatives, with monthly downloads read 2026-09-21
ModelReleasedLicenseDownloads a month
argmaxinc/whisperkit-coremlFebruary 2024Check the card11.0M
openai/whisper-large-v3-turboOctober 2024MIT6.6M
openai/whisper-large-v3November 2023MIT4.6M
Systran/faster-whisper-smallNovember 2023MIT2.9M
openai/whisper-smallSeptember 2022MIT2.9M
nvidia/parakeet-tdt-0.6b-v2April 2025CC-BY-4.00.19M
nvidia/canary-qwen-2.5bJune 2025CC-BY-4.00.03M

whisper-large-v3-turbo has overtaken large-v3, and it is the checkpoint to start from for transcription. It is a large-v3 with a much smaller decoder, so it transcribes far faster at a small cost in accuracy. For translation into English, use large-v3 or medium, because turbo wasn't trained for it.

Whisper is no longer the most accurate open model on English. NVIDIA's Canary and Parakeet families sit above it on the Open ASR Leaderboard, and the leaderboard's paper compares 86 systems across 12 datasets, standardising word error rate against inverse real-time factor. Read the leaderboard itself for current figures, because they move and because blog posts about them go stale within weeks.

The download column shows what people run anyway. The NVIDIA models are ahead on accuracy and two orders of magnitude behind on usage, and their CC-BY-4.0 licence is a different commitment from Whisper's MIT.

Choosing

Choosing, and when to reach past Whisper

Which model to reach for
The jobReach forWhy
General transcription in many languageswhisper-large-v3-turboFast, MIT-licensed, and it handles accents and noise well
Speech in another language, written out in Englishwhisper-large-v3 or medium with the translate taskTurbo wasn't trained for translation and returns the original language
Live captions while someone is talkingA streaming wrapper like Whisper-Streaming, or a model built for streamingWhisper itself writes text a window at a time
The best English accuracy you can self-hostAn NVIDIA Canary modelIt leads the Open ASR Leaderboard, under CC-BY-4.0
Raw throughput on a batch of filesA Parakeet model, or faster-whisperBuilt for speed per hour of audio
Transcription on a phone or a laptopWhisperKit or whisper.cppThe same weights, compiled for the device
Knowing who spoke whenA speaker diarization model alongside itWhisper writes words and never says who said them
A voice agent that answers out loudA speech-to-speech modelOne model in place of three, with far lower latency
Audio that is music or sound, with no speech in itAn audio understanding modelWhisper is trained to transcribe speech and nothing else

Start with turbo unless you have a reason not to. Move to Canary when English accuracy is the binding constraint and the licence suits you, and move to a device-targeted build when the audio cannot leave the machine.

Try it

How to try it

Transcription is one call once the audio is in the shape the model expects, which is one channel at 16,000 samples a second. Audio files are often stereo at 44,100 or 48,000, so the loader converts both.

transcribe.py — audio in, text out
import librosa
import torch
from transformers import WhisperForConditionalGeneration, WhisperProcessor

model_id = "openai/whisper-large-v3-turbo"     # MIT, the one to start from
proc = WhisperProcessor.from_pretrained(model_id)
model = WhisperForConditionalGeneration.from_pretrained(model_id).eval()

audio, sr = librosa.load("clip.wav", sr=16000, mono=True)  # resample, downmix
feats = proc(audio, sampling_rate=sr, return_tensors="pt").input_features
# input_features covers the first 30 seconds only, the rest is cut off

with torch.no_grad():
    ids = model.generate(feats, language="en", max_new_tokens=200)
print(proc.batch_decode(ids, skip_special_tokens=True)[0])

Each step in that loader guards against a specific failure in the transformers feature extractor. A sampling rate other than 16,000 raises an error, so at least that one is loud. A stereo file loaded as a two-column array gets read as a batch, one tiny clip per row, and anything past 30 seconds is cut off without a warning, so a 10-minute file comes back as its first half minute of text.

For anything longer than 30 seconds, use faster-whisper, which handles the chunking and the timestamps. Its voice activity filter, the step that skips audio with no speech in it, is off by default in WhisperModel.transcribe and needs vad_filter=True. The batched pipeline, BatchedInferencePipeline, turns it on by default.

Ask your AI coding tool

Measure Whisper on my own audio before I trust it. Take a folder of audio files with a CSV of reference transcripts. Load every file as mono 16 kHz audio and print any file that needed resampling or downmixing. Transcribe each file with the turbo model through faster-whisper, and compute word error rate after normalising both sides, which means lowercasing, stripping punctuation and expanding common abbreviations, then report the WER with and without that normalisation so I can see how much of the error is formatting. Report the real-time factor per file. Then run the same files with the voice activity filter turned off and on, and list any segment where the two runs differ, since those are where the model is writing over silence.

Limits

What it can't do

  • It writes text over audio that has no speech in it. The run above produced the word "you" from ten seconds of silence and a run of repeated characters from quiet noise. Use a voice activity detector in front of it and drop segments it marks as silent.
  • It doesn't tell you who spoke. Diarization is a separate model, and running the two together is the normal setup.
  • It has no streaming mode. Live transcription needs a wrapper that re-runs it on a growing buffer, which adds a few seconds of delay.
  • Turbo doesn't translate. Translation into English needs large-v3 or medium, and turbo returns the original language without an error.
  • Long audio is chunked, and the seams show. The 30-second window means repeated or dropped words can appear at boundaries, and different libraries stitch them differently.
  • Timestamps are approximate. Word-level timestamps come from extra machinery on top, and they drift, which is worth checking before you build subtitle alignment on them.
  • It is no longer the accuracy leader. Current NVIDIA models score better on English on the Open ASR Leaderboard, which is worth knowing when transcription quality is the product.
  • Quality varies a lot by language. The multilingual training is uneven, so a language with little training data gets noticeably worse results, and the only way to know is to test on your own recordings.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.