best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is WER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Hindi · 9 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 Qwen3-ASR-1.7Bmost accurate 7.847.09–8.70 3.8× 2.0× 2.04 GB Apache-2.0
2 nemotron-3.5-asr-streaming-0.6boverall pickfast pick 8.617.67–9.70 14.5× 7.5× 0.70 GB OpenMDW-1.1
3 Qwen3-ASR-0.6B 12.6811.63–13.83 8.0× 4.3× 0.79 GB Apache-2.0
4 Voxtral-Mini-4B-Realtime-2602 17.0415.34–18.80 0.9× 0.6× 4.73 GB Apache-2.0
5 whisper-large-v3 17.0615.97–18.29 2.1× 0.6× 1.55 GB Apache-2.0
6 whisper-large-v3-turbo 18.8517.82–20.10 3.4× 0.8× 0.83 GB Apache-2.0
7 whisper-medium 26.0924.77–27.63 4.3× 1.1× 0.77 GB Apache-2.0
8 whisper-small 42.0539.97–44.32 12.1× 3.4× 0.25 GB Apache-2.0
9 Fun-ASR-MLT-Nano-2512! 43.9639.95–48.08 9.0× 4.5× 0.83 GB FunASR Model Open Source License Agreement v1.1