Instructions to use wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8") config = load_config("wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Hermes Agent
How to use wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8
Run Hermes
hermes
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
--auth-choice custom-api-key \
--custom-base-url http://127.0.0.1:8080/v1 \
--custom-model-id "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8" \
--custom-provider-id mlx-lm \
--custom-compatibility openai \
--custom-text-input \
--accept-risk \
--skip-healthRun OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8
Mixed-precision quantization for Apple Silicon, with vision-language (VLM) capabilities preserved. Highest quality version in the series — closest to BF16 baseline.
Quantized from lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled using oMLX's oQ8 algorithm (sensitivity-aware mixed-precision quantization).
📊 Specs
| Field | Value |
|---|---|
| Base model | Qwen/Qwen3.6-35B-A3B (35B params, 128 experts MoE, A3B activation) |
| Fine-tune | LoRA distilled from Claude 4.7 Opus reasoning outputs (lordx64 dataset) |
| Quantization | oMLX oQ8 (mixed-precision, ~8.7 bpw average) |
| Modality | Vision + Text (VLM) |
| Format | MLX safetensors |
| Model size | ~35 GB |
| Inference memory | ~39 GB (incl. KV cache and runtime overhead) |
| Recommended hardware | Apple Silicon M2 Ultra 64GB+ / M3 Max / M5 Max |
🚀 Quick Start
Install
pip install mlx-vlm
# Or with uv:
uv tool install mlx-vlm --with torch --with torchvision
Inference (image + text)
mlx_vlm.generate \
--model wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8 \
--image /path/to/image.jpg \
--prompt "Describe this image in detail." \
--max-tokens 256
Python API
from mlx_vlm import load, generate
model, processor = load(
"wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8"
)
output = generate(
model,
processor,
image="/path/to/image.jpg",
prompt="What's in this image?",
max_tokens=512,
)
print(output)
OpenAI-compatible Server
mlx_vlm.server \
--model wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8 \
--port 8080
Drop-in compatible with OpenAI clients including AstrBot, Open WebUI, LibreChat, and Continue.dev.
📈 Measured Performance
Benchmarked on MacBook Pro M5 Max 128GB:
| Metric | Value |
|---|---|
| Prompt processing | ~637 tokens/s |
| Generation speed | ~96 tokens/s |
| Peak memory | 38.7 GB |
| Model load time | ~14 sec |
When to Choose oQ8 over oQ6
oQ8 retains slightly more precision than oQ6, but the observable quality difference is small for most tasks. Use oQ8 when:
- You need the absolute highest fidelity quantization
- Running quality benchmarks against the BF16 reference
- Memory budget is generous (39+ GB free)
For most users, oQ6 is the better choice — comparable quality at 8 GB less footprint.
🧠 Model Behavior
Inherits the Claude reasoning distillation: the model uses <think>...</think> tags to structure its chain-of-thought before producing the final response.
Best for:
- Multimodal reasoning tasks (image analysis with complex thinking)
- Agentic workflows / tool use
- Quality-sensitive applications where size is not a constraint
- Quantization quality reference / baseline comparisons
🔬 Quantization Details
- Source model:
Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled(BF16 MLX-converted) - Sensitivity model: 8-bit quantization of the same distilled model (self-referenced sens for tight distribution alignment)
- Non-quant weight dtype: bfloat16 (M3+ optimal)
- Text-Only mode: OFF (vision tower preserved)
- Quantizer: oMLX with mlx-vlm conversion path
The vision processor configurations (preprocessor_config.json, video_preprocessor_config.json, processor_config.json) were sourced from the official Qwen base model to ensure proper image input handling — these were missing from the upstream distilled checkpoint.
📦 Other Versions in This Series
| Version | Size | Best for |
|---|---|---|
| VLM-MLX-oQ4 | 19.6 GB | Memory-constrained inference |
| VLM-MLX-oQ6 | 27 GB | Recommended: best quality/size ratio |
| VLM-MLX-oQ8 | 35 GB | Quality reference baseline |
| Text-MLX-oQ4 | 19 GB | Text-only, fastest |
| Text-MLX-oQ6 | 27 GB | Text-only, balanced |
| Text-MLX-oQ8 | 34 GB | Text-only, max quality |
Choosing a version:
- Text-only workflows (coding, agents, dialogue) →
Textvariants are faster and lighter - Image input needed (OCR, visual analysis, screenshot understanding) →
VLMvariants - oQ6 is the sweet spot for most use cases. oQ8 yields diminishing returns relative to its size.
⚠️ Disclaimer
This model derives from a chain of upstream work:
- Base model
Qwen/Qwen3.6-35B-A3Bby Alibaba's Qwen team (Apache-2.0) - Distilled variant
lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilledby lordx64 using Claude 4.7 Opus reasoning outputs (Apache-2.0) - This quantization by @wangkezun using oMLX on Apple Silicon
This model is not affiliated with or endorsed by Anthropic, PBC. "Claude" is a trademark of Anthropic, PBC. The use of "Claude" in this model name is purely descriptive (nominative fair use) to indicate the upstream training data lineage.
By using this model, you agree to comply with:
- The Apache-2.0 license inherited from the base model
- Any applicable license terms of the upstream distillation dataset
- Local laws and regulations governing AI model usage in your jurisdiction
🙏 Acknowledgments
- Alibaba Qwen Team — for the Qwen3.6-35B-A3B base model
- lordx64 — for the reasoning-focused LoRA distillation
- Jundot (oMLX team) — for the oQ mixed-precision quantization algorithm
- Apple MLX team — for the MLX framework and tooling
- mlx-vlm contributors — for the VLM conversion path
📜 License
Apache-2.0 (inherited from base model).
Generated: 2026-04-26
Quantizer: @wangkezun
- Downloads last month
- 128
8-bit
Model tree for wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8
Base model
Qwen/Qwen3.6-35B-A3B
Start the MLX server
# Install MLX LM: uv tool install mlx-lm# Start a local OpenAI-compatible server: mlx_lm.server --model "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8"