best model by language
Benchmarked on the FLEURS test split at Q8_0. Error is CER — lower is better. Speed is × realtime on Handy target hardware.
accuracy vs speed
Mandarin Chinese · 12 models · left is slower, down is better · log speed axis
AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
| | | | | | |
| 1 | Qwen3-ASR-1.7Bmost accurate | 7.146.26–8.12 | 3.8× | 2.0× | 2.04 GB | Apache-2.0 |
| 2 | Qwen3-ASR-0.6Boverall pick | 7.576.70–8.42 | 8.0× | 4.3× | 0.79 GB | Apache-2.0 |
| 3 | whisper-large-v3 | 7.987.12–8.82 | 2.1× | 0.6× | 1.55 GB | Apache-2.0 |
| 4 | Breeze-ASR-25 | 8.107.28–8.94 * | 2.1× | 0.6× | 1.67 GB | Apache-2.0 |
| 5 | whisper-large-v3-turbo | 8.507.65–9.40 | 3.4× | 0.8× | 0.83 GB | Apache-2.0 |
| 6 | Fun-ASR-MLT-Nano-2512! | 8.647.70–9.55 | 9.0× | 4.5× | 0.83 GB | FunASR Model Open Source License Agreement v1.1 |
| 7 | SenseVoiceSmallfast pick | 10.129.16–11.08 | 32.5× | 15.5× | 0.24 GB | model-license (FunASR MODEL_LICENSE) |
| 8 | Voxtral-Mini-4B-Realtime-2602 | 10.419.32–11.52 | 0.9× | 0.6× | 4.73 GB | Apache-2.0 |
| 9 | cohere-transcribe-03-2026! | 11.1810.19–12.21 | 8.0× | 3.0× | 2.41 GB | Apache-2.0 |
| 10 | whisper-medium | 13.1311.97–14.25 | 4.3× | 1.1× | 0.77 GB | Apache-2.0 |
| 11 | nemotron-3.5-asr-streaming-0.6b | 18.8717.82–19.91 | 14.5× | 7.5× | 0.70 GB | OpenMDW-1.1 |
| 12 | whisper-small | 23.0621.79–24.35 | 12.1× | 3.4× | 0.25 GB | Apache-2.0 |