Parakeet vs Whisper on Apple silicon: our benchmark
We measured accuracy, speed and load time for four local speech models on an M4 Max. Here are the numbers, the method, and what they don't tell you.
Sayburst runs speech recognition on your Mac, so the model choice decides both how accurate it is and how quickly text appears. We benchmarked the four models we ship, using exactly the code that runs in the app.
The results
Measured on a Mac Studio (M4 Max, 36 GB), macOS 15.6, October 1, 2026.
| Model | Word error rate | Speed | Median time to text | First load |
|---|---|---|---|---|
| Parakeet v3 | 0.46% | 96× real time | 82 ms | 18.6 s |
| Parakeet Ultra | 0.65% | 120× real time | 69 ms | 21.1 s |
| Whisper Large v3 Turbo (q5) | 0.83% | 29× real time | 320 ms | 16.8 s |
| Whisper Base (English) | 1.66% | 154× real time | 39 ms | 0.1 s |
"Time to text" is how long transcription takes after you let go of the key, for short dictation-length clips. Cleanup adds less than a millisecond.
How we measured
- Accuracy: 40 utterances from LibriSpeech test-clean — 1,082 words from four speakers reading audiobooks aloud. We lowercase and strip punctuation before comparing.
- Latency: eight dictation-style clips made with macOS text-to-speech, timed from the end of the audio to the finished transcript.
- First load: the one-time preparation after installing a model, when macOS compiles it for your chip.
What the numbers don't tell you
- Real dictation is messier than audiobooks: background noise, fast speech, half-finished thoughts. Expect higher error rates for every model.
- Older and smaller Macs are slower. A base M1 will take noticeably longer, though Parakeet runs on the Neural Engine and stays fast.
- Numbers aren't normalized, so "10" vs "ten" counts as an error.
What we chose
Parakeet v3 is the default: lowest error rate in our test, fast, and it runs on the Neural Engine, leaving the GPU free. Whisper Large v3 Turbo is the pick for the 99 languages Parakeet doesn't cover. Parakeet Ultra is there if you want the newest model; in our test it was faster but not more accurate.
The benchmark tool and sample list are in the Sayburst repository, and we'll publish results from more Macs as we test them.
Credits
Parakeet TDT 0.6B v3 and v2 by NVIDIA, CC BY 4.0, converted to Core ML by FluidInference. Parakeet Ultra by moondream, CC BY 4.0. Whisper by OpenAI, MIT; ggml conversions by the whisper.cpp project.
Questions people ask
Is Parakeet more accurate than Whisper?
On our small English test set, Parakeet v3 made fewer errors than Whisper Large v3 Turbo (0.46% vs 0.83% word error rate), but the sample is too small to call a definitive winner. Whisper supports far more languages.
Which model does Sayburst use by default?
Parakeet v3 for English and 24 European languages, and Whisper Large v3 Turbo for other languages.