Qwen3.5-0.8B Parakeet FullFT — MLX 5-bit

This is a 5-bit MLX conversion of rdsm/Qwen3.5-0.8B-parakeet-52k-FullFT, published by TalkMate for local English transcript cleanup on Apple silicon.

The source model is a full fine-tune of Qwen/Qwen3.5-0.8B trained on rdsm/parakeet-stt-redone. It cleans already-transcribed text; it is not a speech-recognition model.

Provenance

Source model rdsm/Qwen3.5-0.8B-parakeet-52k-FullFT
Source revision d1b8e0c7371779e7da322332ccea4a019b4c3910
Source format BF16 safetensors
Conversion tools mlx-lm 0.31.3, MLX 0.32.0
Quantization Uniform affine, 5 bits, group size 64
Effective quantization 5.507 bits per weight
Weight size 518,013,158 bytes
Weight SHA-256 ce91877f73def5cd98d1a4a91030607db88849c1adb60a577c14948138a642ea

The conversion was produced with:

hf download rdsm/Qwen3.5-0.8B-parakeet-52k-FullFT \
  --revision d1b8e0c7371779e7da322332ccea4a019b4c3910 \
  --local-dir source

mlx_lm.convert \
  --hf-path source \
  --mlx-path mlx-5bit \
  -q \
  --q-bits 5

TalkMate evaluation

TalkMate selected 5-bit after comparing the pinned BF16 source with local 8-bit, 5-bit, and 4-bit MLX conversions on a deliberately authored 21-case transcript-cleanup corpus.

The 5-bit artifact completed three runs through TalkMate's production path:

  • 21 / 21 important requirements passed in every run;
  • 6 / 8 nice-to-have requirements passed in every run;
  • all 21 final outputs were byte-for-byte identical across the three runs;
  • mean cleanup latency was 0.678 seconds and aggregate p95 was 0.936 seconds on a 16 GB M4 Mac with the transcription model unloaded.

Important requirements covered short fragments, punctuation, disfluency removal, Personal Vocabulary casing, numbers, negation, long-form meaning, instruction-like dictation, and preservation of both spoken course-correction clauses.

These are TalkMate application-path results, not an independent raw-model benchmark. The path parses the model's cleaned_text JSON field and applies TalkMate's vocabulary guard, faithfulness guard, and conservative deterministic list formatter. The target 8 GB M1 memory test with Parakeet Accurate resident remains outstanding.

Evaluation configuration

TalkMate used the source model's transcript-cleanup prompt contract with:

  • thinking disabled;
  • temperature 0.2;
  • a 160-token benchmark limit;
  • JSON output read from the cleaned_text field.

Applications integrating this model should independently validate output faithfulness and fall back to the original transcript if meaningful content is added, removed, or changed.

Limitations

  • The TalkMate corpus is intentionally small and product-specific.
  • Results do not establish quality for every accent, ASR engine, language, or long-form transcript.
  • This quantization is published by TalkMate and is not the author's separately referenced 8-bit conversion.
  • Quantized models can behave differently from their source checkpoint.
  • The model should not be used without application-level checks when exact transcript fidelity matters.

License and attribution

The source checkpoint and this conversion are distributed under the Apache License 2.0. Please retain attribution to the source model author, the Qwen/Qwen3.5-0.8B base model, and the original training dataset.

Downloads last month
48
Safetensors
Model size
0.1B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TalkMate/Qwen3.5-0.8B-parakeet-52k-FullFT-MLX-5bit

Quantized
(1)
this model