Models & hardware
Every model runs on your Mac. Pick one in Settings › Models; you can install several and switch any time.
Requirements
- Apple silicon Mac (M1 or later), macOS 14 or later.
- Memory: each model needs roughly the amount listed below while loaded. We expect 8 GB Macs to work, but have only measured on a 36 GB Mac so far.
- Disk space for each model you install (sizes below).
Models
| Model | Languages | Download | Memory | Runs on |
|---|---|---|---|---|
| Parakeet v3 Fast and accurate for English and 24 other European languages. | 25 languages | 483 MB | ~800 MB | Neural Engine |
| Parakeet Ultra Most accurate Parakeet. Same languages and speed as v3, larger download. | 25 languages | 632 MB | ~900 MB | Neural Engine |
| Parakeet v2 (English) English only. Strong on accents and long dictation. | English | 464 MB | ~800 MB | Neural Engine |
| Whisper Large v3 Turbo 99 languages, including Chinese, Japanese, Korean, Hindi and Arabic. Slower than Parakeet. | 99 languages | 574 MB | ~1200 MB | GPU (Metal) |
| Whisper Small 99 languages in a smaller download. Less accurate than Turbo. | 99 languages | 488 MB | ~850 MB | GPU (Metal) |
| Whisper Base (English) Smallest download, English only. For older or low-memory Macs. | English | 148 MB | ~400 MB | GPU (Metal) |
Parakeet models detect which of their languages you're speaking. Whisper models detect any of 99 languages, or you can set one in Settings › Models for better accuracy.
Measured performance
Run on Mac Studio (M4 Max, 36 GB), Version 15.6 (Build 24G84), 2026-10-01, using the same code that ships in the app.
| Model | Word error rate | Speed | Median time to text | First load |
|---|---|---|---|---|
| Parakeet v3 | 0.46% | 96× real time | 82 ms | 18.6 s |
| Parakeet Ultra | 0.65% | 120× real time | 69 ms | 21.1 s |
| Whisper Large v3 Turbo | 0.83% | 29× real time | 320 ms | 16.8 s |
| Whisper Base (English) | 1.66% | 154× real time | 39 ms | 0.1 s |
How this was measured — and its limits
- Accuracy: 40 utterances (1,082 words, 4 speakers) from LibriSpeech test-clean — clearly read audiobook speech. Real dictation is messier, so expect higher error rates. With this few words, a difference of a few errors between models isn't meaningful.
- Normalization: lowercase, punctuation removed; numbers aren't converted, so “10” vs “ten” counts as an error.
- Time to text: 8 short dictation-style clips made with macOS text-to-speech, timed from the end of audio to the transcript. Cleanup adds under a millisecond.
- First load happens once after installing a model, while macOS compiles it for your chip. Later launches are faster.
- Only one Mac was measured. Base M1 Macs will be noticeably slower; we'll publish more results as we test more hardware.
The benchmark tool and sample list are in the Sayburst source repository (bench/).
Credits
- Parakeet TDT 0.6B v3 by NVIDIA, CC BY 4.0. Converted to Core ML by FluidInference. Source · CC BY 4.0
- Parakeet Ultra by moondream, post-trained from NVIDIA Parakeet TDT 0.6B v3, CC BY 4.0. Converted to Core ML by FluidInference. Source · CC BY 4.0
- Parakeet TDT 0.6B v2 by NVIDIA, CC BY 4.0. Converted to Core ML by FluidInference. Source · CC BY 4.0
- Whisper by OpenAI, MIT. ggml conversion by the whisper.cpp project, MIT. Source · MIT
- Whisper by OpenAI, MIT. ggml conversion by the whisper.cpp project, MIT. Source · MIT
- Whisper by OpenAI, MIT. ggml conversion by the whisper.cpp project, MIT. Source · MIT