--- pipeline_tag: automatic-speech-recognition language: mri license: apache-2.0 tags: - trimmed library_name: transformers base_model: openai/whisper-large-v3-turbo base_model_relation: quantized datasets: - lbourdois/fineweb-2-trimming --- # whisper-large-v3-turbo-mri-32768 This model is a **3.02% whisper-large-v3-turbo** version of [openai/whisper-large-v3-turbo](https://huggingface.co/openai/whisper-large-v3-turbo) optimized for **Maori** language via vocabulary size reduction using the [trimming](https://huggingface.co/blog/lbourdois/introduction-to-trimming) method. This trimmed model should perform similarly to the original model with only 32,768 tokens and a much smaller memory footprint. However, it may not perform well for other languages as tokens not commonly used in the selected languages were removed from the vocabulary. ## Model Statistics | Metric | Original | Trimmed | Reduction | |--------|----------|---------|-----------| | **Vocabulary size** | 51,865 tokens | 32,768 tokens | **36.82%** | | **Model size** | 808,878,080 params | 784,432,640 params | **3.02%** | ![image](https://raw.githubusercontent.com/lbourdois/blog/refs/heads/master/assets/images/Trimming/whisper-large-v3-turbo-32768.png) ## Mining Dataset Statistics - **Number of texts used for mining**: 158,804 texts - **Dataset**: [lbourdois/fineweb-2-trimming](https://huggingface.co/datasets/lbourdois/fineweb-2-trimming) ## Usage ```python from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline import librosa # Pipeline function processor = AutoProcessor.from_pretrained("alphaedge-ai/whisper-large-v3-turbo-mri-32768") pipe = pipeline( "automatic-speech-recognition", model="alphaedge-ai/whisper-large-v3-turbo-mri-32768", tokenizer=processor.tokenizer, feature_extractor=processor.feature_extractor, generate_kwargs={"language": "maori", "task": "transcribe"}, ) # Loading and resampling at 16 kHz (required by Whisper) audio_array, sampling_rate = librosa.load(audio_path, sr=16000) # Result result = pipe(audio_array) print("Transcription :", result["text"]) ``` ## Citations #### Whisper ``` @misc{radford2022whisper, doi = {10.48550/ARXIV.2212.04356}, url = {https://arxiv.org/abs/2212.04356}, author = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya}, title = {Robust Speech Recognition via Large-Scale Weak Supervision}, publisher = {arXiv}, year = {2022}, copyright = {arXiv.org perpetual, non-exclusive license} } ``` #### Trimming blog post ``` @misc{hf_blogpost_trimming, title={Introduction to Trimming}, author={Loïck BOURDOIS and Tom AARSEN and Bram VANROY and Christopher AKIKI and Woojun JUNG and Manuel ROMERO and Prithiv SAKTHI}, year={2026}, url={https://huggingface.co/blog/lbourdois/introduction-to-trimming}, } ```