Instructions to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451") model = AutoModelForMultimodalLM.from_pretrained("nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - MLX
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451") config = load_config("nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- vLLM
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
- SGLang
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Unsloth Studio
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451", max_seq_length=2048, ) - Pi
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
Run Hermes
hermes
- OpenClaw new
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451 with Docker Model Runner:
docker model run hf.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
Qwen3.6 27B family evaluation
Standby, A large family wide evaluation landing soon...
I have hit a ... strange quirk, assessing.. The BF16 model is showing regression from the GGUF Q8_0 model that I unpacked into BF16 and tested, was this QAT/QLORA merged, might explain part of it?
Both are showing an anomalous behavior in a NIAH test I've not observed in 27B variants before.
Implementing a feature in vllm that will help me dissect it, overall there IS an improvement that is incontrovertible vs the base model, but I'm trying to understand what's going on with the anomaly before I drop results.
I am running overnight a test for the BF16 made from source to establish a baseline. This will take about 10-12 hours.
Where it stands now can be viewed here, I am working to resolve the mentioned anomalies above, when done I will rerun the whole family to give each an honest shake down. In short, I'm implementing DRY in vllm to clamp thinking repetition across the whole family so they get a fair capability shakedown and not a "does this model loop" shakedown, because full transparency...every single 27B will thinking loop if stressed hard enough.
I looked at the metrics, they make complete sense.
I see higher numbers with qx86-hi and mxfp8 that are consistently higher than BF16 in previous models.
This is normal, and I rarely see a BF16 outperform a tuned quant. The GGUF XL mixed quants and David's NEO raise those metrics even further.
On MLX I see the same effect on qx86-hi and qx64-hi, here is for example the Qwen3.6-27B-Architect-DS9-Polaris2
arc arc/e boolq hswag obkqa piqa wino
bf16 0.692,0.863,0.911
mxfp8 0.699,0.871,0.910
q8-hi 0.694,0.865,0.910
qx86-hi 0.688,0.862,0.910
qx64-hi 0.700,0.862,0.907
mxfp4 0.694,0.872,0.909
Quant Perplexity Peak Memory Tokens/sec
bf16 3.898 ± 0.025 60.75 GB 226
q8-hi 3.895 ± 0.025 37.26 GB 215
mxfp8 3.921 ± 0.025 34.74 GB 218
qx86-hi 3.898 ± 0.025 32.36 GB 218
qx64-hi 3.918 ± 0.025 25.64 GB 217
mxfp4 3.999 ± 0.025 21.30 GB 225
If you compare now how the BF16 and F32 is quanting down, mxfp4 is the only one that did not change at all.
From F32, the lower quants are much better, while mxfp8 suffers, as expected. I did not benchmark the q8-hi from F32 to compare it to BF16, that would be an interesting test.
The F32 source is available.
Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
mxfp8 0.711,0.879,0.910,0.790,0.514,0.823,0.763
qx86-hi 0.696,0.876,0.912,0.791,0.518,0.824,0.760
qx64-hi 0.702,0.873,0.909,0.794,0.514,0.822,0.750
mxfp4 0.701,0.873,0.909,0.786,0.488,0.813,0.759
nvfp4 0.701,0.866,0.907
Quant Perplexity Peak Memory Tokens/sec
mxfp8 3.783 ± 0.023 34.74 GB 203
qx86-hi 3.735 ± 0.023 33.25 GB 183
qx64-hi 3.747 ± 0.023 27.03 GB 194
mxfp4 3.854 ± 0.024 21.30 GB 197
Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-F32
mxfp8 0.709,0.880,0.909
qx64-hi 0.706,0.873,0.908
nvfp4 0.704,0.868,0.908
mxfp4 0.701,0.873,0.909
Also, while you are at it, there is another one
https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-Tess
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.712,0.879,0.911,0.792,0.508,0.823,0.764
qx86-hi 0.701,0.877,0.911,0.794,0.518,0.823,0.758
qx64-hi 0.706,0.873,0.909,0.795,0.512,0.823,0.752
mxfp4 0.706,0.873,0.910,0.790,0.496,0.817,0.761
1M
mxfp8 0.703,0.877,0.909
qx86-hi 0.702,0.873,0.911,0.793,0.508,0.824,0.765
qx64-hi 0.706,0.873,0.909,0.795,0.512,0.823,0.752
mxfp4 0.701,0.874,0.912,0.789,0.500,0.817,0.759
Quant Perplexity Peak Memory Tokens/sec
mxfp8 3.797 ± 0.024 34.74 GB 183
qx86-hi 3.756 ± 0.023 33.25 GB 170
qx64-hi 3.765 ± 0.023 27.03 GB 181
mxfp4 3.876 ± 0.024 21.26 GB 176
1M
mxfp8 3.803 ± 0.024 34.70 GB 169
qx86-hi 3.758 ± 0.023 33.21 GB 175
qx64-hi 3.769 ± 0.023 26.99 GB 172
mxfp4 3.876 ± 0.024 21.26 GB 176
That validates some other model and quant testing I've done, it seems the distribution snapping impact of quantization seems to make certain behaviors have a higher probability in a way that is represented as better performance in testing.
I've been trying to figure out why my custom 4 bit always leads to better performance than 16 bit in certain tests so, while I might not fully understand the exact mechanism, its good to know others are also observing the effect. Thanks for the responses!