Qwen3.5-0.8B V15 LoRA Dictation Corrector (MLX mxfp4 g=32)

RAM-constrained V15 R-3 corrector for low-memory Apple Silicon (M1 8 GB and M2 16 GB-1 hosts). PTQ on V15's bf16-fused weights via mlx-lm --q-mode mxfp4 --q-group-size 32. No adapter retraining; sidesteps the noisy-Q4-gradient failure mode that doomed the train-on-Q4 V14/V16 lineage.

Eval results — three independent angles

Eval set V15 prod (Q8 g=64) This mirror V14 4-bit baseline
seed_v5 wild (50 rows, hard-neg discourse) 100% 100% 74.1%
ECHO15 long-form WER vs Gemini-3.1 ref 0.210 0.164 not tested
Dev-corpus long-form WER vs Gemini-3.5-Flash ref 0.141 0.140 not tested
Disk size 782 MB 400 MB 424 MB
Warm RSS (process) 1573 MB 1244 MB ~1100 MB
Bits/weight (avg) ~8.5 {4.258} ~4.5

This mirror is strictly Pareto-optimal vs V15 Q8 on every axis: smaller, lower RSS, equal-or-better quality. See companion mirror qwen3-5-0.8b-dictation-corrector-mlx-mixed46-g32 for the M2 16 GB+ (16 GiB+ unified) tier.

Tier-of-use

Best for: M1 8 GB (8-15 GiB unified).

The voice-scribe-macos installer (WP#1074) ships both this mirror and its sibling and picks one at install time via sysctl hw.memsize:

  • 8-15 GiB → mxfp4 g=32 (this mirror, 400 MB / 1244 MB RSS)
  • 16+ GiB → mixed46 g=32 (482 MB / 1326 MB RSS)

Quickstart

from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load("VoiceScribe/qwen3-5-0.8b-dictation-corrector-mlx-mxfp4-g32")

messages = [
    {"role": "system", "content": "Корректор русской диктовки. Убери слова-паразиты. Нормализуй IT-термины. Не меняй смысл."},
    {"role": "user", "content": "Эм, докер мониторит, ну, бэкенд через гитхаб экшнс"},
]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
sampler = make_sampler(temp=0.0, top_p=1.0)
out = generate(model, tokenizer, prompt=prompt, max_tokens=200, sampler=sampler)
print(out)  # → "Docker мониторит бэкенд через GitHub Actions"

Quantisation recipe — no retraining

Source: V15 R-3 LoRA-fused bf16 weights (the predecessor of the V15 Q8 production mirror). PTQ via mlx-lm 0.31.x:

python -m mlx_lm convert \
    --hf-path /path/to/v15-r3-fused-bf16 \
    --mlx-path ./qwen3-5-0.8b-dictation-corrector-mlx-mxfp4-g32 \
{    --quantize --q-mode mxfp4 --q-group-size 32 \
}    --dtype bfloat16

The bf16-fused weights are the canonical V15 source; the V15 Q8 production mirror was produced from the same source via uniform Q8 quantisation. This mirror simply uses a smarter quantiser on the same proven training signal.

Why PTQ (not retraining on Q4 base)

Predecessor approaches (V14, V16) attempted to train a LoRA adapter directly on a Q4-quantised base. Both produced regressions because Q4 forward gradients are too noisy — the adapter learns a sloppier signal than the bf16-base V15 path.

This mirror reverses the order: train on bf16 (proven V15 path), then quantise smarter. Result: the adapter signal is preserved faithfully; the only loss is from re-quantising weights with a more aggressive bit budget on layers where it doesn't hurt.

The {mxfp4 (E2M1 floating-point) numerics} keeps the high-magnitude channels (attention output projections — where the LoRA-trained signal concentrates) at high precision while compressing the MLP feedforward layers (which tolerate aggressive quant). Result: 38-49% smaller disk + 16-21% lower RSS at zero (or better) quality vs the uniform Q8 mirror.

Intended use

  • Yes: Russian dictation cleanup after ASR (GigaAM, Whisper, Parakeet). Removes fillers, normalises Cyrillic IT terms (докер→Docker, гитхаб→GitHub), preserves meaning verbatim.
  • No: General text editing, English text, summarisation, translation, creative writing. Trained for a strict conservative-edit policy; will not paraphrase.

Limitations

  • Numeric edge cases: rare numeral-word sequences may regenerate with substitution errors.
  • OOD brand normalisation: brands not in training data may stay in Cyrillic transliteration.
  • English-language inputs: not supported (Russian only).
  • Small eval sample: 50 + 30 + 2 (long-form) files. Real-world variance higher than reported confidence interval.

Ethical considerations

  • Privacy: runs entirely on-device. No telemetry, no cloud round-trip.
  • No user data in training: all training prompts are synthetic (authored by maintainer with AI assistance). Production usage does not contribute to future training.
  • Conservative policy: preserves exact meaning; never paraphrases.

Citation

@software{{voicescribe-v15-mxfp4-2026,
  title  = {{Voice Scribe Russian Dictation Corrector V15 R-3 (mxfp4 {Apple Silicon M1 8 GB tier})}},
  author = {{Sabynin, Andrey}},
  year   = {{2026}},
  note   = {{WP#1067 R&D + WP#1074 productisation}},
  url    = {{https://huggingface.co/VoiceScribe/qwen3-5-0.8b-dictation-corrector-mlx-mxfp4-g32}}
}}

Related repos

Downloads last month
28
Safetensors
Model size
0.1B params
Tensor type
U8
·
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VoiceScribe/qwen3-5-0.8b-dictation-corrector-mlx-mxfp4-g32

Adapter
(198)
this model