best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is WER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

German · 13 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 whisper-large-v3most accurate 4.133.74–4.51 2.1× 0.6× 1.55 GB Apache-2.0
2 Qwen3-ASR-1.7B 4.253.86–4.64 3.8× 2.0× 2.04 GB Apache-2.0
3 canary-1b-v2!overall pickfast pick 4.464.05–4.89 13.2× 6.7× 1.10 GB CC-BY-4.0
4 whisper-large-v3-turbo 4.544.14–4.96 3.4× 0.8× 0.83 GB Apache-2.0
5 cohere-transcribe-03-2026! 5.064.53–5.65 8.0× 3.0× 2.41 GB Apache-2.0
6 parakeet-tdt-0.6b-v3 5.244.83–5.66 12.5× 7.5× 0.72 GB CC-BY-4.0
7 Voxtral-Mini-4B-Realtime-2602 5.835.19–6.51 0.9× 0.6× 4.73 GB Apache-2.0
8 parakeet-primeline 5.985.53–6.44 12.5× 7.5× 0.72 GB CC-BY-4.0
9 whisper-medium 6.235.74–6.71 4.3× 1.1× 0.77 GB Apache-2.0
10 Qwen3-ASR-0.6B 6.806.33–7.30 8.0× 4.3× 0.79 GB Apache-2.0
11 canary-180m-flash! 7.336.67–8.00 31.9× 21.3× 0.20 GB CC-BY-4.0
12 whisper-small 9.869.23–10.49 12.1× 3.4× 0.25 GB Apache-2.0
13 nemotron-3.5-asr-streaming-0.6b 10.339.58–11.13 14.5× 7.5× 0.70 GB OpenMDW-1.1