best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is CER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Japanese · 11 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 Fun-ASR-MLT-Nano-2512!overall pickmost accuratefast pick 2.322.01–2.65 9.0× 4.5× 0.83 GB FunASR Model Open Source License Agreement v1.1
2 whisper-large-v3 4.814.25–5.61 2.1× 0.6× 1.55 GB Apache-2.0
3 whisper-large-v3-turbo 4.824.40–5.30 3.4× 0.8× 0.83 GB Apache-2.0
4 cohere-transcribe-03-2026! 5.134.48–5.84 8.0× 3.0× 2.41 GB Apache-2.0
5 Qwen3-ASR-1.7B 5.294.81–5.80 3.8× 2.0× 2.04 GB Apache-2.0
6 whisper-medium 7.356.79–7.91 4.3× 1.1× 0.77 GB Apache-2.0
7 SenseVoiceSmall 7.637.10–8.22 32.5× 15.5× 0.24 GB model-license (FunASR MODEL_LICENSE)
8 Qwen3-ASR-0.6B 8.617.98–9.28 8.0× 4.3× 0.79 GB Apache-2.0
9 Voxtral-Mini-4B-Realtime-2602 9.458.20–11.05 0.9× 0.6× 4.73 GB Apache-2.0
10 whisper-small 12.8112.05–13.52 12.1× 3.4× 0.25 GB Apache-2.0
11 nemotron-3.5-asr-streaming-0.6b 13.5212.78–14.27 14.5× 7.5× 0.70 GB OpenMDW-1.1