| --- |
| license: apache-2.0 |
| tags: |
| - executorch |
| - xnnpack |
| - pte |
| - on-device |
| - automatic-speech-recognition |
| base_model: |
| - openai/whisper-tiny |
| --- |
| # Whisper-tiny β ExecuTorch (encoder + decoder) |
|
|
| Speech recognition in two `.pte` files: the encoder runs once per 30-second window, the |
| decoder once per generated token. Putting them in one graph would re-encode the audio on |
| every step. |
|
|
| | graph | build | file | size (MB) | corr vs fp32 eager | ms | eager ms | |
| |---|---|---|---|---|---|---| |
| | encoder | XNNPACK fp32 | `whisper_tiny_encoder_xnnpack_fp32.pte` | 32.9 | 1.000000 | 60.2 | 22.1 | |
| | encoder | XNNPACK fp16 | `whisper_tiny_encoder_xnnpack_fp16.pte` | 17.6 | 1.000000 | 110.5 | 22.8 | |
| | encoder | XNNPACK int8 | `whisper_tiny_encoder_xnnpack_int8.pte` | 11.7 | 0.999439 | 58.9 | 22.0 | |
| | encoder | Core ML | `whisper_tiny_encoder_coreml_all.pte` | 16.6 | 0.999992 | 13.5 | 22.3 | |
| | decoder | XNNPACK fp32 | `whisper_tiny_decoder_xnnpack_fp32.pte` | 198.0 | 1.000000 | 20.6 | 11.9 | |
| | decoder | XNNPACK fp16 | `whisper_tiny_decoder_xnnpack_fp16.pte` | 99.1 | 0.999991 | 48.7 | 12.1 | |
| | decoder | Core ML | `whisper_tiny_decoder_coreml_all.pte` | 59.3 | 0.999892 | 3.0 | 11.8 | |
|
|
| Every file takes and returns fp32 tensors (token ids stay int64), so any encoder pairs with |
| any decoder. The lightest working pair is 71.0 MB. |
|
|
| - **Source**: [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny) |
| - **License**: Apache-2.0 |
| - **Encoder input**: log-mel spectrogram `[1, 80, 3000]` β 30 s at 16 kHz, 80 mel bins, hop |
| 160, window 400, exactly what `WhisperFeatureExtractor` produces |
| - **Decoder input**: the encoder output plus `decoder_input_ids [1, 128]` int64, |
| left-aligned and padded. Start with `<|startoftranscript|>`, a language token, |
| `<|transcribe|>`, `<|notimestamps|>`. |
|
|
| ## Decoding |
|
|
| No KV cache: the decoder is a static graph over a fixed 128-token window, so a greedy step |
| is take `argmax` of row `len-1`, append it, run again. Stop at `<|endoftext|>` (50257). 128 |
| tokens covers a 30-second window of ordinary speech; past that, start a new window. |
|
|
| That costs a full 128-position forward pass per token, which is the price of a static graph |
| that runs unchanged across runtimes and precisions. |
|
|
| ## Verification (Mac arm64, executorch 1.4.0, torch 2.13.0) |
|
|
| The two wrappers compose back to `WhisperForConditionalGeneration` exactly β max_abs_diff |
| **0.000e+00** β and every graph matches torch fp32 eager at the correlations above. Timings |
| are medians over 5 runs in one process: a relative reference, not a device number. |
|
|
| ## Two things worth knowing about the sizes |
|
|
| **The decoder `.pte` is larger than the decoder's weights.** Whisper ties `proj_out.weight` |
| to `decoder.embed_tokens.weight`, but the two uses need different representations: an |
| embedding table the portable kernels index into, and the same values packed into the XNNPACK |
| delegate's blob for the output matmul. Tying them in PyTorch does not tie them here. |
| Referencing the weight through `F.linear` instead of the `proj_out` module does not either β |
| exported both ways, whisper-tiny's decoder comes out at 198.0 MB exactly. |
|
|
| **The decoder's int8 build is not shipped.** Dynamic int8 quantizes the linear weights and |
| leaves the token embedding table in fp32, and that table is 79.7 MB β 51,865 tokens at |
| 384 dimensions. On this size that table is most of the file, so int8 lands at 110.2 MB against fp16's 99.1 MB β larger, because fp16 halves the table too. It converts and holds, but nothing would pick it, so it is not shipped. |
|
|
| Until recently there was no decoder int8 build at all, and this card said PT2E was observing |
| the int64 `decoder_input_ids`. That was wrong on both halves. |
| `XNNPACKQuantizer.transform_for_annotation` rewrites every scalar argument of |
| `add.Tensor`/`mul.Tensor` as `torch.tensor(float(arg))` whatever the node's dtype β one line |
| in ExecuTorch's `backends/xnnpack/quantizer/xnnpack_quantizer_utils.py`, still present on |
| main. In this decoder the casualty is `position_ids = torch.arange(...) + past_key_values_length` |
| (`modeling_whisper.py:749`, `past_key_values_length` being a python `int`): it comes back |
| float32, and the failure lands on `self.weight[position_ids]` β the **position** embedding |
| lookup, not the token ids, and no observer involved. Measured by running `prepare_pt2e` |
| with an empty quantizer and printing the failing node. |
|
|
| ### Checked in the task's own units |
|
|
| Correlation is a first filter. These are the numbers that decide: |
|
|
| - **encoder int8** β measured end to end β word error rate against the fp32 encoder: mean WER 0.0% (worst clip 0.0%) over 5 spoken sentences, int8 encoder against the fp32 encoder with the same fp32 decoder and the same waveform; the fp32 arm transcribes all five correctly, so the comparison is against a working control rather than against noise. |
|
|
| The sensitivity of that test, measured by injecting random noise into whisper-tiny's encoder output: rel_l2 0.03 (what int8 actually costs) and 0.10 both give WER 0.000; 0.20 and 0.40 give 0.025. Five clean sentences leave headroom, so a pass means *does not break the transcript*, not *indistinguishable at any error level*. |
| |
| ## Conversion |
| |
| ```bash |
| python convert/export_whisper.py tiny |
| ``` |
| |
| The ExecuTorch tree ships a single-graph Whisper example under `examples/models/whisper`; |
| this is that model with the halves separated. |
| |
| (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models)) |
| |