ToPo-ToPo commited on
Commit
77f84a7
·
verified ·
1 Parent(s): cc8a705

Add/standardize MTP usage with the model's matching drafter

Browse files
Files changed (1) hide show
  1. README.md +30 -0
README.md CHANGED
@@ -21,3 +21,33 @@ model, processor = load("ToPo-ToPo/gemma-4-E4B-it-qat-mlx-4bit")
21
 
22
  ## License
23
  Derivative of Google Gemma; governed by the Gemma Terms of Use (https://ai.google.dev/gemma/terms) and Prohibited Use Policy. Converted to MLX.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
21
 
22
  ## License
23
  Derivative of Google Gemma; governed by the Gemma Terms of Use (https://ai.google.dev/gemma/terms) and Prohibited Use Policy. Converted to MLX.
24
+
25
+ ## ⚡ Faster generation with MTP (speculative decoding, lossless)
26
+
27
+ **Recommended drafter: `google/gemma-4-E4B-it-assistant`** — Google's official MTP drafter for this
28
+ model. It loads **directly in mlx-vlm (no conversion needed)** and gives up to
29
+ ~3x faster generation (≈1.4–1.5x measured on short prompts); output is
30
+ **identical** to non-MTP decoding.
31
+
32
+ ```python
33
+ # requires: pip install "mlx-vlm>=0.6.3"
34
+ from mlx_vlm import load, generate
35
+ from mlx_vlm.prompt_utils import apply_chat_template
36
+ from mlx_vlm.utils import load_config
37
+
38
+ model, processor = load("ToPo-ToPo/gemma-4-E4B-it-qat-mlx-4bit")
39
+ draft_model, _ = load("google/gemma-4-E4B-it-assistant")
40
+ config = load_config("ToPo-ToPo/gemma-4-E4B-it-qat-mlx-4bit")
41
+
42
+ prompt = apply_chat_template(processor, config, "Hello!", num_images=0)
43
+ out = generate(model, processor, prompt,
44
+ draft_model=draft_model, draft_kind="mtp", max_tokens=256)
45
+ ```
46
+
47
+ CLI (draft_kind auto-detected):
48
+ `mlx_vlm.generate --model ToPo-ToPo/gemma-4-E4B-it-qat-mlx-4bit --draft-model google/gemma-4-E4B-it-assistant`
49
+
50
+ ### Notes
51
+ - `draft_kind="mtp"` is required in the Python API (the CLI auto-detects it).
52
+ - Use **this model's own** drafter above — drafters are size-specific and not interchangeable across Gemma 4 variants.
53
+ - Needs **mlx-vlm >= 0.6.3**. MTP is lossless — if output differs from non-MTP, your versions are mismatched.