best model by language

Benchmarked on the FLEURS test split at Q8_0. Error is WER — lower is better. Speed is × realtime on Handy target hardware.

accuracy vs speed

Arabic · 11 models · left is slower, down is better · log speed axis

AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
1 cohere-transcribe-arabic-07-2026!overall pickmost accurate 11.069.62–12.60 8.0× 3.0× ?
2 cohere-transcribe-03-2026! 13.6012.12–15.16 8.0× 3.0× 2.41 GB Apache-2.0
3 Qwen3-ASR-1.7B 14.9113.53–16.42 3.8× 2.0× 2.04 GB Apache-2.0
4 whisper-large-v3 14.9213.55–16.39 2.1× 0.6× 1.55 GB Apache-2.0
5 whisper-large-v3-turbo 15.4814.10–16.99 3.4× 0.8× 0.83 GB Apache-2.0
6 nemotron-3.5-asr-streaming-0.6bfast pick 15.9314.47–17.54 14.5× 7.5× 0.70 GB OpenMDW-1.1
7 Voxtral-Mini-4B-Realtime-2602 16.0514.29–18.09 0.9× 0.6× 4.73 GB Apache-2.0
8 whisper-medium 21.9020.40–23.52 4.3× 1.1× 0.77 GB Apache-2.0
9 Qwen3-ASR-0.6B 24.5122.15–28.02 8.0× 4.3× 0.79 GB Apache-2.0
10 Fun-ASR-MLT-Nano-2512! 25.7924.27–27.42 9.0× 4.5× 0.83 GB FunASR Model Open Source License Agreement v1.1
11 whisper-small 32.1530.52–33.87 12.1× 3.4× 0.25 GB Apache-2.0

Benchmarked but not listed on the model card: cohere-transcribe-arabic-07-2026.