best model by language
Benchmarked on the FLEURS test split at Q8_0. Error is WER — lower is better. Speed is × realtime on Handy target hardware.
accuracy vs speed
English · 17 models · left is slower, down is better · log speed axis
AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
| 1 | Qwen3-ASR-1.7Bmost accurate | 3.232.85–3.65 | 3.8× | 2.0× | 2.04 GB | Apache-2.0 |
| 2 | parakeet-unified-en-0.6boverall pickfast pick | 3.993.60–4.42 | 12.5× | 7.5× | 0.71 GB | CC-BY-4.0 |
| 3 | whisper-large-v3 | 4.033.59–4.46 | 2.1× | 0.6× | 1.55 GB | Apache-2.0 |
| 4 | Breeze-ASR-25 | 4.143.70–4.61 | 2.1× | 0.6× | 1.67 GB | Apache-2.0 |
| 5 | Qwen3-ASR-0.6B | 4.233.76–4.69 | 8.0× | 4.3× | 0.79 GB | Apache-2.0 |
| 6 | whisper-large-v3-turbo | 4.383.95–4.84 | 3.4× | 0.8× | 0.83 GB | Apache-2.0 |
| 7 | canary-1b-v2 | 4.474.03–4.95 | 13.2× | 6.7× | 1.10 GB | CC-BY-4.0 |
| 8 | whisper-medium | 4.644.20–5.15 | 4.3× | 1.1× | 0.77 GB | Apache-2.0 |
| 9 | parakeet-primeline | 4.824.39–5.34 | 12.5× | 7.5× | 0.72 GB | CC-BY-4.0 |
| 10 | parakeet-tdt-0.6b-v3 | 4.834.39–5.30 | 12.5× | 7.5× | 0.72 GB | CC-BY-4.0 |
| 11 | Fun-ASR-MLT-Nano-2512 | 4.904.40–5.48 | 9.0× | 4.5× | 0.83 GB | FunASR Model Open Source License Agreement v1.1 |
| 12 | cohere-transcribe-03-2026 | 5.084.56–5.59 | 8.0× | 3.0× | 2.41 GB | Apache-2.0 |
| 13 | canary-180m-flash | 5.985.26–6.78 | 31.9× | 21.3× | 0.20 GB | CC-BY-4.0 |
| 14 | whisper-small | 6.515.88–7.22 | 12.1× | 3.4× | 0.25 GB | Apache-2.0 |
| 15 | SenseVoiceSmall | 7.146.54–7.77 | 32.5× | 15.5× | 0.24 GB | model-license (FunASR MODEL_LICENSE) |
| 16 | nemotron-3.5-asr-streaming-0.6b | 7.907.23–8.54 | 14.5× | 7.5× | 0.70 GB | OpenMDW-1.1 |
| 17 | Voxtral-Mini-4B-Realtime-2602 | 11.7710.02–13.71 | 0.9× | 0.6× | 4.73 GB | Apache-2.0 |