best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is WER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Italian · 12 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 whisper-large-v3most accurate 2.542.16–2.98 2.1× 0.6× 1.55 GB Apache-2.0
2 Qwen3-ASR-1.7B 2.682.38–2.99 3.8× 2.0× 2.04 GB Apache-2.0
3 whisper-large-v3-turbo 2.772.45–3.08 3.4× 0.8× 0.83 GB Apache-2.0
4 parakeet-tdt-0.6b-v3 3.022.72–3.34 12.5× 7.5× 0.72 GB CC-BY-4.0
5 canary-1b-v2!overall pick 3.102.77–3.46 13.2× 6.7× 1.10 GB CC-BY-4.0
6 parakeet-primeline 3.172.88–3.48 12.5× 7.5× 0.72 GB CC-BY-4.0
7 cohere-transcribe-03-2026! 3.242.86–3.61 8.0× 3.0× 2.41 GB Apache-2.0
8 Voxtral-Mini-4B-Realtime-2602 3.453.00–3.94 0.9× 0.6× 4.73 GB Apache-2.0
9 whisper-medium 4.173.75–4.65 4.3× 1.1× 0.77 GB Apache-2.0
10 Qwen3-ASR-0.6B 5.194.78–5.64 8.0× 4.3× 0.79 GB Apache-2.0
11 nemotron-3.5-asr-streaming-0.6bfast pick 5.785.24–6.28 14.5× 7.5× 0.70 GB OpenMDW-1.1
12 whisper-small 7.977.40–8.52 12.1× 3.4× 0.25 GB Apache-2.0