best model by language
Benchmarked on the FLEURS test split at Q8_0. Error is CER — lower is better. Speed is × realtime on Handy target hardware.
accuracy vs speed
Cantonese · 6 models · left is slower, down is better · log speed axis
AMD Ryzen 4750U GPU (Vulkan)
pareto frontier (no model is both faster and more accurate) other models
| | | | | | |
| 1 | Qwen3-ASR-1.7Bmost accurate | 6.135.55–6.68 | 3.8× | 2.0× | 2.04 GB | Apache-2.0 |
| 2 | Qwen3-ASR-0.6Boverall pickfast pick | 7.917.28–8.52 | 8.0× | 4.3× | 0.79 GB | Apache-2.0 |
| 3 | Fun-ASR-MLT-Nano-2512! | 12.7211.82–13.59 | 9.0× | 4.5× | 0.83 GB | FunASR Model Open Source License Agreement v1.1 |
| 4 | whisper-large-v3 | 22.0620.16–24.13 | 2.1× | 0.6× | 1.55 GB | Apache-2.0 |
| 5 | whisper-large-v3-turbo | 34.6233.41–36.13 | 3.4× | 0.8× | 0.83 GB | Apache-2.0 |
| 6 | SenseVoiceSmall | 37.4436.69–38.21 | 32.5× | 15.5× | 0.24 GB | model-license (FunASR MODEL_LICENSE) |