best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is CER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Korean · 11 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 Qwen3-ASR-1.7Bmost accurate 4.603.62–5.65 3.8× 2.0× 2.04 GB Apache-2.0
2 whisper-large-v3 4.893.93–5.90 2.1× 0.6× 1.55 GB Apache-2.0
3 Fun-ASR-MLT-Nano-2512!overall pick 5.204.21–6.29 9.0× 4.5× 0.83 GB FunASR Model Open Source License Agreement v1.1
4 whisper-large-v3-turbo 5.244.30–6.25 3.4× 0.8× 0.83 GB Apache-2.0
5 whisper-medium 5.464.53–6.45 4.3× 1.1× 0.77 GB Apache-2.0
6 Qwen3-ASR-0.6B 5.824.86–6.83 8.0× 4.3× 0.79 GB Apache-2.0
7 Voxtral-Mini-4B-Realtime-2602 5.994.92–7.17 0.9× 0.6× 4.73 GB Apache-2.0
8 cohere-transcribe-03-2026! 6.575.50–7.65 8.0× 3.0× 2.41 GB Apache-2.0
9 whisper-small 7.706.65–8.76 12.1× 3.4× 0.25 GB Apache-2.0
10 SenseVoiceSmallfast pick 8.277.13–9.45 32.5× 15.5× 0.24 GB model-license (FunASR MODEL_LICENSE)
11 nemotron-3.5-asr-streaming-0.6b 8.897.78–10.06 14.5× 7.5× 0.70 GB OpenMDW-1.1