Models & hardware

Every model runs on your Mac. Pick one in Settings › Models; you can install several and switch any time.

Requirements

  • Apple silicon Mac (M1 or later), macOS 14 or later.
  • Memory: each model needs roughly the amount listed below while loaded. We expect 8 GB Macs to work, but have only measured on a 36 GB Mac so far.
  • Disk space for each model you install (sizes below).

Models

ModelLanguagesDownloadMemoryRuns on
Parakeet v3
Fast and accurate for English and 24 other European languages.
25 languages483 MB~800 MBNeural Engine
Parakeet Ultra
Most accurate Parakeet. Same languages and speed as v3, larger download.
25 languages632 MB~900 MBNeural Engine
Parakeet v2 (English)
English only. Strong on accents and long dictation.
English464 MB~800 MBNeural Engine
Whisper Large v3 Turbo
99 languages, including Chinese, Japanese, Korean, Hindi and Arabic. Slower than Parakeet.
99 languages574 MB~1200 MBGPU (Metal)
Whisper Small
99 languages in a smaller download. Less accurate than Turbo.
99 languages488 MB~850 MBGPU (Metal)
Whisper Base (English)
Smallest download, English only. For older or low-memory Macs.
English148 MB~400 MBGPU (Metal)

Parakeet models detect which of their languages you're speaking. Whisper models detect any of 99 languages, or you can set one in Settings › Models for better accuracy.

Measured performance

Run on Mac Studio (M4 Max, 36 GB), Version 15.6 (Build 24G84), 2026-10-01, using the same code that ships in the app.

ModelWord error rateSpeedMedian time to textFirst load
Parakeet v30.46%96× real time82 ms18.6 s
Parakeet Ultra0.65%120× real time69 ms21.1 s
Whisper Large v3 Turbo0.83%29× real time320 ms16.8 s
Whisper Base (English)1.66%154× real time39 ms0.1 s

How this was measured — and its limits

  • Accuracy: 40 utterances (1,082 words, 4 speakers) from LibriSpeech test-clean — clearly read audiobook speech. Real dictation is messier, so expect higher error rates. With this few words, a difference of a few errors between models isn't meaningful.
  • Normalization: lowercase, punctuation removed; numbers aren't converted, so “10” vs “ten” counts as an error.
  • Time to text: 8 short dictation-style clips made with macOS text-to-speech, timed from the end of audio to the transcript. Cleanup adds under a millisecond.
  • First load happens once after installing a model, while macOS compiles it for your chip. Later launches are faster.
  • Only one Mac was measured. Base M1 Macs will be noticeably slower; we'll publish more results as we test more hardware.

The benchmark tool and sample list are in the Sayburst source repository (bench/).

Credits

  • Parakeet TDT 0.6B v3 by NVIDIA, CC BY 4.0. Converted to Core ML by FluidInference. Source · CC BY 4.0
  • Parakeet Ultra by moondream, post-trained from NVIDIA Parakeet TDT 0.6B v3, CC BY 4.0. Converted to Core ML by FluidInference. Source · CC BY 4.0
  • Parakeet TDT 0.6B v2 by NVIDIA, CC BY 4.0. Converted to Core ML by FluidInference. Source · CC BY 4.0
  • Whisper by OpenAI, MIT. ggml conversion by the whisper.cpp project, MIT. Source · MIT
  • Whisper by OpenAI, MIT. ggml conversion by the whisper.cpp project, MIT. Source · MIT
  • Whisper by OpenAI, MIT. ggml conversion by the whisper.cpp project, MIT. Source · MIT