best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is WER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Russian · 12 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 whisper-large-v3most accurate 4.964.51–5.41 2.1× 0.6× 1.55 GB Apache-2.0
2 gigaam-v3-e2e-rnntoverall pickfast pick 5.354.85–5.90 22.0× 8.0× 0.25 GB MIT
3 whisper-large-v3-turbo 5.934.94–7.56 3.4× 0.8× 0.83 GB Apache-2.0
4 Voxtral-Mini-4B-Realtime-2602 6.145.55–6.78 0.9× 0.6× 4.73 GB Apache-2.0
5 Qwen3-ASR-1.7B 6.255.74–6.79 3.8× 2.0× 2.04 GB Apache-2.0
6 parakeet-tdt-0.6b-v3 6.546.08–7.06 12.5× 7.5× 0.72 GB CC-BY-4.0
7 whisper-medium 7.306.71–7.85 4.3× 1.1× 0.77 GB Apache-2.0
8 parakeet-primeline 7.817.21–8.42 12.5× 7.5× 0.72 GB CC-BY-4.0
9 canary-1b-v2! 7.837.24–8.48 13.2× 6.7× 1.10 GB CC-BY-4.0
10 Qwen3-ASR-0.6B 10.309.59–11.00 8.0× 4.3× 0.79 GB Apache-2.0
11 whisper-small 11.9011.19–12.62 12.1× 3.4× 0.25 GB Apache-2.0
12 nemotron-3.5-asr-streaming-0.6b 12.6111.87–13.40 14.5× 7.5× 0.70 GB OpenMDW-1.1