AEON-7 commited on
Commit
1af382a
Β·
verified Β·
1 Parent(s): c83d4f7

Make MTP the default quickstart (drafter auto-pulled); no-MTP optional

Browse files
Files changed (1) hide show
  1. README.md +13 -6
README.md CHANGED
@@ -56,12 +56,17 @@ This is the **fidelity** member of the MLX quant grid (29.5 GB on disk, 8.634 bp
56
  ```bash
57
  curl -LsSf https://astral.sh/uv/install.sh | sh && source $HOME/.local/bin/env # one-time: install uv
58
 
59
- # serve β€” uv fetches Python 3.12 + mlx-vlm(main) on first run Β· MLX-8bit (max fidelity)
 
 
60
  uv run --python 3.12 --with "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm" -- \
61
  python -m mlx_vlm.server --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit \
 
62
  --port 8080 --trust-remote-code
63
  ```
64
 
 
 
65
  Call it like an OpenAI endpoint (`POST http://localhost:8080/v1/chat/completions`) with the request `"model"` set to the launched id. *(While this repo is private, run `hf auth login` first β€” or pass a local `--model` path.)*
66
 
67
  **Sampling β€” set `temperature: 1.0`.** The MLX server defaults to *greedy* decoding (`temperature 0`), which can repeat or loop on long prompts. This model is tuned for its native sampling β€” **`temperature 1.0`** (`top_p 0.95`, `top_k ~64`). Pass it in every request (clients that send no sampling params fall back to greedy):
@@ -84,18 +89,20 @@ python -m mlx_vlm.generate --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-M
84
  ```
85
  </details>
86
 
87
- ### ⚑⚑ Optional β€” +MTP self-speculation (lossless throughput boost)
88
 
89
- Qwen ships a properly-trained **MTP head**, packaged here as a native `qwen3_5_mtp` drafter β€” it *proposes* tokens this model then *verifies*. Every token is verified, so the **output is identical** β€” purely a throughput boost. Pull the [drafter repo](https://huggingface.co/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter) and add the three `--draft-*` flags. **Use `--draft-block-size 3`** β€” the benchmarked sweet spot.
90
 
91
  ```bash
92
  uv run --python 3.12 --with "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm" -- \
93
  python -m mlx_vlm.server --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit \
94
- --port 8080 --trust-remote-code \
95
- --draft-model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter --draft-kind mtp --draft-block-size 3
96
  ```
 
 
 
97
 
98
- Remove the three `--draft-*` flags to disable. *(The MTP sweep below was measured on the FP4 build β€” `bs=3` lands 1.78Γ— lossless there. The same drafter pairs with this 8-bit target; throughput gains apply on top of the 8-bit baseline.)*
99
 
100
  KV-cache quant for long context (optional): `--kv-bits 8 --kv-group-size 64 --quantized-kv-start 1024`. `--max-kv-size` is ignored under `--kv-bits`; `--prefill-step-size` is inert under MTP.
101
 
 
56
  ```bash
57
  curl -LsSf https://astral.sh/uv/install.sh | sh && source $HOME/.local/bin/env # one-time: install uv
58
 
59
+ # serve MLX-8bit + MTP self-speculation (recommended default β€” lossless throughput boost).
60
+ # uv fetches Python 3.12 + mlx-vlm(main) on first run. --model and --draft-model are HF repo ids,
61
+ # so mlx-vlm pulls BOTH the 29.5 GB model and the 821 MB MTP drafter automatically on first run.
62
  uv run --python 3.12 --with "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm" -- \
63
  python -m mlx_vlm.server --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit \
64
+ --draft-model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter --draft-kind mtp --draft-block-size 3 \
65
  --port 8080 --trust-remote-code
66
  ```
67
 
68
+ `--draft-block-size 3` is the benchmarked sweet spot. MTP is **lossless** β€” every drafted token is verified against the target, so the output is byte-identical to running without it, just faster. *(Prefer to pre-fetch the drafter explicitly? `hf download AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter`.)*
69
+
70
  Call it like an OpenAI endpoint (`POST http://localhost:8080/v1/chat/completions`) with the request `"model"` set to the launched id. *(While this repo is private, run `hf auth login` first β€” or pass a local `--model` path.)*
71
 
72
  **Sampling β€” set `temperature: 1.0`.** The MLX server defaults to *greedy* decoding (`temperature 0`), which can repeat or loop on long prompts. This model is tuned for its native sampling β€” **`temperature 1.0`** (`top_p 0.95`, `top_k ~64`). Pass it in every request (clients that send no sampling params fall back to greedy):
 
89
  ```
90
  </details>
91
 
92
+ <details><summary>Run <strong>without</strong> MTP β€” ~1.5 GB less unified memory, but slower</summary>
93
 
94
+ MTP is the recommended default (lossless + faster). If you're tight on unified memory, drop the three `--draft-*` flags to serve the target alone β€” that frees the 821 MB drafter plus its speculative buffers (**~1.5 GB less peak RAM**) at the cost of the speedup:
95
 
96
  ```bash
97
  uv run --python 3.12 --with "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm" -- \
98
  python -m mlx_vlm.server --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit \
99
+ --port 8080 --trust-remote-code
 
100
  ```
101
+ </details>
102
+
103
+ ### ⚑⚑ Why MTP is on by default
104
 
105
+ Qwen ships a properly-trained native `qwen3_5_mtp` **MTP head**, packaged as a separate 821 MB drafter that *proposes* tokens this model then *verifies* β€” so the **output is identical**, purely a throughput boost. `--draft-block-size 3` is the benchmarked sweet spot. *(The MTP sweep below was measured on the FP4 build β€” `bs=3` lands 1.78Γ— lossless there. The same drafter pairs with this 8-bit target; throughput gains apply on top of the 8-bit baseline.)*
106
 
107
  KV-cache quant for long context (optional): `--kv-bits 8 --kv-group-size 64 --quantized-kv-start 1024`. `--max-kv-size` is ignored under `--kv-bits`; `--prefill-step-size` is inert under MTP.
108