Image-Text-to-Text
MLX
Safetensors
English
qwen3_5
mlx-vlm
apple-silicon
metal
mixed-precision
quantized
qwen
qwen3
qwen3.6
multimodal
vision
mtp
speculative-decoding
gated-deltanet
mamba
ssm
linear-attention
uncensored
abliterated
refusal-removed
aeon
aeon-7
m4-pro
on-device
conversational
8-bit precision
8bit
int8
Instructions to use AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit") config = load_config("AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit
Run Hermes
hermes
- OpenClaw new
How to use AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Make MTP the default quickstart (drafter auto-pulled); no-MTP optional
Browse files
README.md
CHANGED
|
@@ -56,12 +56,17 @@ This is the **fidelity** member of the MLX quant grid (29.5 GB on disk, 8.634 bp
|
|
| 56 |
```bash
|
| 57 |
curl -LsSf https://astral.sh/uv/install.sh | sh && source $HOME/.local/bin/env # one-time: install uv
|
| 58 |
|
| 59 |
-
# serve
|
|
|
|
|
|
|
| 60 |
uv run --python 3.12 --with "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm" -- \
|
| 61 |
python -m mlx_vlm.server --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit \
|
|
|
|
| 62 |
--port 8080 --trust-remote-code
|
| 63 |
```
|
| 64 |
|
|
|
|
|
|
|
| 65 |
Call it like an OpenAI endpoint (`POST http://localhost:8080/v1/chat/completions`) with the request `"model"` set to the launched id. *(While this repo is private, run `hf auth login` first β or pass a local `--model` path.)*
|
| 66 |
|
| 67 |
**Sampling β set `temperature: 1.0`.** The MLX server defaults to *greedy* decoding (`temperature 0`), which can repeat or loop on long prompts. This model is tuned for its native sampling β **`temperature 1.0`** (`top_p 0.95`, `top_k ~64`). Pass it in every request (clients that send no sampling params fall back to greedy):
|
|
@@ -84,18 +89,20 @@ python -m mlx_vlm.generate --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-M
|
|
| 84 |
```
|
| 85 |
</details>
|
| 86 |
|
| 87 |
-
|
| 88 |
|
| 89 |
-
|
| 90 |
|
| 91 |
```bash
|
| 92 |
uv run --python 3.12 --with "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm" -- \
|
| 93 |
python -m mlx_vlm.server --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit \
|
| 94 |
-
--port 8080 --trust-remote-code
|
| 95 |
-
--draft-model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter --draft-kind mtp --draft-block-size 3
|
| 96 |
```
|
|
|
|
|
|
|
|
|
|
| 97 |
|
| 98 |
-
|
| 99 |
|
| 100 |
KV-cache quant for long context (optional): `--kv-bits 8 --kv-group-size 64 --quantized-kv-start 1024`. `--max-kv-size` is ignored under `--kv-bits`; `--prefill-step-size` is inert under MTP.
|
| 101 |
|
|
|
|
| 56 |
```bash
|
| 57 |
curl -LsSf https://astral.sh/uv/install.sh | sh && source $HOME/.local/bin/env # one-time: install uv
|
| 58 |
|
| 59 |
+
# serve MLX-8bit + MTP self-speculation (recommended default β lossless throughput boost).
|
| 60 |
+
# uv fetches Python 3.12 + mlx-vlm(main) on first run. --model and --draft-model are HF repo ids,
|
| 61 |
+
# so mlx-vlm pulls BOTH the 29.5 GB model and the 821 MB MTP drafter automatically on first run.
|
| 62 |
uv run --python 3.12 --with "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm" -- \
|
| 63 |
python -m mlx_vlm.server --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit \
|
| 64 |
+
--draft-model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter --draft-kind mtp --draft-block-size 3 \
|
| 65 |
--port 8080 --trust-remote-code
|
| 66 |
```
|
| 67 |
|
| 68 |
+
`--draft-block-size 3` is the benchmarked sweet spot. MTP is **lossless** β every drafted token is verified against the target, so the output is byte-identical to running without it, just faster. *(Prefer to pre-fetch the drafter explicitly? `hf download AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter`.)*
|
| 69 |
+
|
| 70 |
Call it like an OpenAI endpoint (`POST http://localhost:8080/v1/chat/completions`) with the request `"model"` set to the launched id. *(While this repo is private, run `hf auth login` first β or pass a local `--model` path.)*
|
| 71 |
|
| 72 |
**Sampling β set `temperature: 1.0`.** The MLX server defaults to *greedy* decoding (`temperature 0`), which can repeat or loop on long prompts. This model is tuned for its native sampling β **`temperature 1.0`** (`top_p 0.95`, `top_k ~64`). Pass it in every request (clients that send no sampling params fall back to greedy):
|
|
|
|
| 89 |
```
|
| 90 |
</details>
|
| 91 |
|
| 92 |
+
<details><summary>Run <strong>without</strong> MTP β ~1.5 GB less unified memory, but slower</summary>
|
| 93 |
|
| 94 |
+
MTP is the recommended default (lossless + faster). If you're tight on unified memory, drop the three `--draft-*` flags to serve the target alone β that frees the 821 MB drafter plus its speculative buffers (**~1.5 GB less peak RAM**) at the cost of the speedup:
|
| 95 |
|
| 96 |
```bash
|
| 97 |
uv run --python 3.12 --with "mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm" -- \
|
| 98 |
python -m mlx_vlm.server --model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-8bit \
|
| 99 |
+
--port 8080 --trust-remote-code
|
|
|
|
| 100 |
```
|
| 101 |
+
</details>
|
| 102 |
+
|
| 103 |
+
### β‘β‘ Why MTP is on by default
|
| 104 |
|
| 105 |
+
Qwen ships a properly-trained native `qwen3_5_mtp` **MTP head**, packaged as a separate 821 MB drafter that *proposes* tokens this model then *verifies* β so the **output is identical**, purely a throughput boost. `--draft-block-size 3` is the benchmarked sweet spot. *(The MTP sweep below was measured on the FP4 build β `bs=3` lands 1.78Γ lossless there. The same drafter pairs with this 8-bit target; throughput gains apply on top of the 8-bit baseline.)*
|
| 106 |
|
| 107 |
KV-cache quant for long context (optional): `--kv-bits 8 --kv-group-size 64 --quantized-kv-start 1024`. `--max-kv-size` is ignored under `--kv-bits`; `--prefill-step-size` is inert under MTP.
|
| 108 |
|