best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is CER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Mandarin Chinese · 12 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 Qwen3-ASR-1.7Bmost accurate 7.146.26–8.12 3.8× 2.0× 2.04 GB Apache-2.0
2 Qwen3-ASR-0.6Boverall pick 7.576.70–8.42 8.0× 4.3× 0.79 GB Apache-2.0
3 whisper-large-v3 7.987.12–8.82 2.1× 0.6× 1.55 GB Apache-2.0
4 Breeze-ASR-25 8.107.28–8.94 * 2.1× 0.6× 1.67 GB Apache-2.0
5 whisper-large-v3-turbo 8.507.65–9.40 3.4× 0.8× 0.83 GB Apache-2.0
6 Fun-ASR-MLT-Nano-2512! 8.647.70–9.55 9.0× 4.5× 0.83 GB FunASR Model Open Source License Agreement v1.1
7 SenseVoiceSmallfast pick 10.129.16–11.08 32.5× 15.5× 0.24 GB model-license (FunASR MODEL_LICENSE)
8 Voxtral-Mini-4B-Realtime-2602 10.419.32–11.52 0.9× 0.6× 4.73 GB Apache-2.0
9 cohere-transcribe-03-2026! 11.1810.19–12.21 8.0× 3.0× 2.41 GB Apache-2.0
10 whisper-medium 13.1311.97–14.25 4.3× 1.1× 0.77 GB Apache-2.0
11 nemotron-3.5-asr-streaming-0.6b 18.8717.82–19.91 14.5× 7.5× 0.70 GB OpenMDW-1.1
12 whisper-small 23.0621.79–24.35 12.1× 3.4× 0.25 GB Apache-2.0