best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is WER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Spanish · 13 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 whisper-large-v3most accurate 2.702.41–3.01 2.1× 0.6× 1.55 GB Apache-2.0
2 canary-1b-v2!overall pickfast pick 3.102.78–3.44 13.2× 6.7× 1.10 GB CC-BY-4.0
3 whisper-large-v3-turbo 3.122.80–3.48 3.4× 0.8× 0.83 GB Apache-2.0
4 Voxtral-Mini-4B-Realtime-2602 3.242.86–3.68 0.9× 0.6× 4.73 GB Apache-2.0
5 Qwen3-ASR-1.7B 3.312.98–3.65 3.8× 2.0× 2.04 GB Apache-2.0
6 parakeet-tdt-0.6b-v3 3.653.32–4.00 12.5× 7.5× 0.72 GB CC-BY-4.0
7 whisper-medium 3.803.42–4.20 4.3× 1.1× 0.77 GB Apache-2.0
8 parakeet-primeline 3.853.48–4.21 12.5× 7.5× 0.72 GB CC-BY-4.0
9 cohere-transcribe-03-2026! 3.973.56–4.39 8.0× 3.0× 2.41 GB Apache-2.0
10 Qwen3-ASR-0.6B 4.884.50–5.29 8.0× 4.3× 0.79 GB Apache-2.0
11 whisper-small 5.925.49–6.37 12.1× 3.4× 0.25 GB Apache-2.0
12 nemotron-3.5-asr-streaming-0.6b 6.305.75–6.89 14.5× 7.5× 0.70 GB OpenMDW-1.1
13 canary-180m-flash! 6.545.92–7.18 31.9× 21.3× 0.20 GB CC-BY-4.0