--- library_name: transformers license: apache-2.0 tags: - pytorch - whisper - automatic-speech-recognition base_model: - biodatlab/whisper-th-medium-combined language: - th --- # Typhoon-isan-asr-whisper | [![Model architecture](https://img.shields.io/badge/Model_Arch-Whisper_(Encoder--Decoder)-lightgrey#model-badge)](#model-architecture) | [![Model size](https://img.shields.io/badge/Params-769M-lightgrey#model-badge)](#model-architecture) | [![Language](https://img.shields.io/badge/Language-th-lightgrey#model-badge)](#datasets) **Typhoon Isan ASR Whisper** is a specialized, fine-tuned version of the [biodatlab/whisper-th-medium-combined](https://huggingface.co/biodatlab/whisper-th-medium-combined) model, optimized specifically for the Isan dialect of the Thai language. Built for high-accuracy offline transcription, it delivers state-of-the-art performance on Isan speech. This enables users to host their own ASR service for Isan dialect recognition, reducing costs and avoiding the need to send sensitive data to third-party cloud services. The model is based on [OpenAI's Whisper architecture](https://huggingface.co/openai/whisper-medium), utilizing the robust Thai-enhanced foundation from biodatlab to ensure superior understanding of regional tones and vocabulary. By using this model, you agree to the OpenTyphoon Terms and Conditions and acknowledge the Privacy Notice: https://opentyphoon.ai/tac · https://opentyphoon.ai/privacy **Try our demo available on [Demo]()** **Code / Examples available on [Github]()** **Release Blog available on [OpenTyphoon Blog]()** *** ### Performance cer comparison **Note on Baseline:** The `scb10x/whisper-medium-slscu-nectec` included in the comparison is a model we fine-tuned specifically for this benchmark using existing dialect data from NECTEC and SLSCU. It serves as a representative baseline for performance based on public data, distinct from the [SLSCU_korat_model](https://huggingface.co/SLSCU/thai-dialect_korat_model) (the prominent previous work for Isan dialect ASR). This helps to determine the clear gap between capabilities derived from previously available resources and the new Typhoon Isan ASR. ### Key Findings * **Outperforming Proprietary State-of-the-Art:** The `typhoon-isan-asr-whisper` model achieves a Character Error Rate (CER) of **0.0885**, surpassing **Gemini-2.5-pro** (0.1020) by a clear margin. This result validates that a specialized, fine-tuned open model can deliver superior accuracy compared to massive, general-purpose proprietary systems for the Isan dialect. * **The Definitive Choice for Accuracy:** Among all architectures tested—including latency-optimized realtime models and historical baselines—the Whisper-based model stands out as the absolute leader. While our `typhoon-isan-asr-realtime` model is highly competitive (0.1065), the `typhoon-isan-asr-whisper` offers the highest possible fidelity, making it the preferred choice for offline transcription where precision is paramount. * ## **Follow us** **https://twitter.com/opentyphoon** ## **Support** **https://discord.gg/us5gAYmrxw**