best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is CER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Cantonese · 6 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 Qwen3-ASR-1.7Bmost accurate 6.135.55–6.68 3.8× 2.0× 2.04 GB Apache-2.0
2 Qwen3-ASR-0.6Boverall pickfast pick 7.917.28–8.52 8.0× 4.3× 0.79 GB Apache-2.0
3 Fun-ASR-MLT-Nano-2512! 12.7211.82–13.59 9.0× 4.5× 0.83 GB FunASR Model Open Source License Agreement v1.1
4 whisper-large-v3 22.0620.16–24.13 2.1× 0.6× 1.55 GB Apache-2.0
5 whisper-large-v3-turbo 34.6233.41–36.13 3.4× 0.8× 0.83 GB Apache-2.0
6 SenseVoiceSmall 37.4436.69–38.21 32.5× 15.5× 0.24 GB model-license (FunASR MODEL_LICENSE)