Text Generation
Transformers
Safetensors
qwen3_5_moe
qwen3.5
Mixture of Experts
fp8
quantized
abliterated
compressed-tensors
vllm
conversational
Instructions to use bjk110/Qwen3.5-122B-A10B-abliterated-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bjk110/Qwen3.5-122B-A10B-abliterated-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bjk110/Qwen3.5-122B-A10B-abliterated-FP8") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoProcessor, AutoModelForCausalLM processor = AutoProcessor.from_pretrained("bjk110/Qwen3.5-122B-A10B-abliterated-FP8") model = AutoModelForCausalLM.from_pretrained("bjk110/Qwen3.5-122B-A10B-abliterated-FP8", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bjk110/Qwen3.5-122B-A10B-abliterated-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bjk110/Qwen3.5-122B-A10B-abliterated-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bjk110/Qwen3.5-122B-A10B-abliterated-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bjk110/Qwen3.5-122B-A10B-abliterated-FP8
- SGLang
How to use bjk110/Qwen3.5-122B-A10B-abliterated-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bjk110/Qwen3.5-122B-A10B-abliterated-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bjk110/Qwen3.5-122B-A10B-abliterated-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bjk110/Qwen3.5-122B-A10B-abliterated-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bjk110/Qwen3.5-122B-A10B-abliterated-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use bjk110/Qwen3.5-122B-A10B-abliterated-FP8 with Docker Model Runner:
docker model run hf.co/bjk110/Qwen3.5-122B-A10B-abliterated-FP8
File size: 3,913 Bytes
ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 e9434b3 ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 1e92560 ee139d3 e9434b3 ee139d3 e9434b3 ee139d3 e9434b3 ee139d3 e9434b3 ee139d3 e9434b3 ee139d3 e9434b3 ee139d3 e9434b3 ee139d3 e9434b3 ee139d3 e9434b3 ee139d3 e9434b3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 | ---
license: apache-2.0
license_name: apache-2.0
license_link: https://hf.co/Qwen/Qwen3.5-122B-A10B/blob/main/LICENSE
base_model:
- wangzhang/Qwen3.5-122B-A10B-abliterated
tags:
- qwen3.5
- moe
- fp8
- quantized
- abliterated
- compressed-tensors
- vllm
language:
- en
- ko
- zh
- ja
library_name: transformers
pipeline_tag: text-generation
---
# Qwen3.5-122B-A10B-abliterated-FP8
FP8 quantized derivative of [wangzhang/Qwen3.5-122B-A10B-abliterated](https://huggingface.co/wangzhang/Qwen3.5-122B-A10B-abliterated), which itself is derived from [Qwen/Qwen3.5-122B-A10B](https://huggingface.co/Qwen/Qwen3.5-122B-A10B).
This repository provides a modified derivative checkpoint for local inference and serving. The primary changes in this repository are FP8 quantization, weight repacking / export formatting, and serving compatibility adjustments.
## Model Details
| Property | Value |
|----------|-------|
| Intermediate Base Model | [wangzhang/Qwen3.5-122B-A10B-abliterated](https://huggingface.co/wangzhang/Qwen3.5-122B-A10B-abliterated) |
| Original Base Model | [Qwen/Qwen3.5-122B-A10B](https://huggingface.co/Qwen/Qwen3.5-122B-A10B) |
| Architecture | Qwen3.5 MoE (256 routed experts, 10B active) |
| Quantization | FP8 |
| Original Size | 228 GB (BF16) |
| Quantized Size | **116 GB** |
| Format | safetensors |
## Quantization Method
This model was quantized from the abliterated BF16 checkpoint into FP8 format for more practical deployment while preserving compatibility with modern inference stacks.
### What is Quantized
| Component | Format | Notes |
|-----------|--------|-------|
| Expert weights | **FP8** | Quantized for reduced memory footprint |
| Attention projections | **FP8** | Quantized where supported |
| Selected sensitive components | **BF16** | Kept at higher precision where needed for stability |
| Embeddings / norms / control tensors | **BF16** | Preserved at full precision |
## Serving with vLLM
This model is intended for vLLM-based inference and may require tensor parallelism depending on available memory.
### Quick Start
```bash
# 1. Download the model
huggingface-cli download bjk110/Qwen3.5-122B-A10B-abliterated-FP8
# 2. Serve with vLLM
vllm serve /path/to/model \
--served-model-name Qwen3.5-122B-A10B-abliterated-FP8 \
--tensor-parallel-size 2 \
--max-model-len 131072 \
--max-num-seqs 4 \
--gpu-memory-utilization 0.90 \
--trust-remote-code \
--enable-prefix-caching \
--enable-chunked-prefill \
--reasoning-parser qwen3
```
### Docker Entrypoint Auto-Patch
Add to the beginning of your entrypoint.sh:
```bash
if [ -f /patches/patch_qwen35_moe_text.py ]; then
python3 /patches/patch_qwen35_moe_text.py || true
fi
```
Mount the patches volume in docker-compose.yml:
```yaml
volumes:
- ./vllm_patches:/patches:ro
```
## What the Patch Does
| Issue | Cause | Fix |
|-------|-------|-----|
| `Qwen3_5MoeForCausalLM` not recognized | Not in vLLM registry | Registers TextOnlyShim class |
| Hybrid cache page-size error | Bug in text-only CausalLM path | Reuses multimodal wrapper's cache-spec |
| Vision encoder init failure | Wrapper forces vision init | Skips vision encoder |
| TP2 block_k=128 error | vision_config.hidden_size=1152 | Injects dummy vision config |
The patch will become unnecessary once vLLM adds native support for `qwen3_5_moe_text`.
## Hardware Requirements
| Config | GPU Memory | Notes |
|--------|-----------|-------|
| TP=1 | ~115 GB | Requires GB200 or similar |
| **TP=2** | **~58 GB/GPU** | DGX Spark, H100×2, A100 80GB×2 |
| TP=4 | ~29 GB/GPU | A100 40GB×4 |
## Base Model
[wangzhang/Qwen3.5-122B-A10B-abliterated](https://huggingface.co/wangzhang/Qwen3.5-122B-A10B-abliterated) — uncensored via [Prometheus](https://github.com/wuwangzhang1216/prometheus) abliteration. Refusal rate 0.5% (1/200), KL divergence 0.0115.
## License
Follows the license of the base model.
|