--- language: - en - zh tags: - qwen - qwen3_5_moe - qwen3_6 - nvfp4 - fp4 - gguf - moe - vision - multimodal - fast license: apache-2.0 base_model: unsloth/Qwen3.6-35B-A3B-NVFP4-Fast pipeline_tag: image-text-to-text library_name: gguf --- # Qwen3.6-35B-A3B-Fast-NVFP4-GGUF GGUF NVFP4 quantization of [unsloth/Qwen3.6-35B-A3B-NVFP4-Fast](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast), a 35B parameter MoE model with 3B active parameters. ## What is the "Fast" Variant? Unsloth's **NVFP4 Fast** variant is a speed-optimized NVFP4 quantization that delivers **1.79x faster throughput** than other NVFP4 quants while maintaining competitive accuracy: | Variant | MMLU-Pro | GPQA | AIME 2025 | |---------|----------|------|-----------| | Unsloth NVFP4 Fast | 85.58 | 87.75 | 91.67 | | Unsloth NVFP4 | 85.85 | 86.74 | 92.29 | | NVIDIA NVFP4 | 85.60 | 87.12 | 91.88 | | BF16 | 85.75 | 86.36 | 92.50 | The Fast variant is calibrated on a mixture of Unsloth's dataset + UltraChat dataset, optimized for throughput on vLLM's native NVFP4 backend (cute-DSL/CUTLASS/flashinfer_trtllm). ## About the Model Qwen3.6-35B-A3B is a multimodal MoE model from Alibaba's Qwen team: - **35B total parameters**, 3B active per token (256 experts, 8 active) - **40-layer decoder** with Gated DeltaNet + full attention hybrid - **27-layer vision encoder** (SigLIP-based) for image/video understanding - **262K native context** (extensible to 1M+ via YaRN) - **Multi-Token Prediction (MTP)** for faster speculative decoding - **Agentic coding** with SWE-bench Verified 73.4, tool calling support ## Files | File | Size | Description | |------|------|-------------| | `qwen36-35b-a3b-fast-nvfp4.gguf` | ~18.8 GB | NVFP4 quantized text model | | `mmproj-qwen36-35b-a3b-f16.gguf` | ~0.84 GB | Vision encoder (F16) | ## Usage ### llama.cpp ```bash llama-server \ -m qwen36-35b-a3b-fast-nvfp4.gguf \ --mmproj mmproj-qwen36-35b-a3b-f16.gguf \ -ngl 99 \ --host 0.0.0.0 \ --port 8080 ``` ### vLLM (Recommended for Max Performance) ```bash pip install vllm flashinfer-python nvidia-cutlass-dsl vllm serve FreedomAISVR/Qwen3.6-35B-A3B-NVFP4-Fast-GGUF --dtype auto --max-model-len 4096 ``` ## Hardware Requirements - **Minimum**: 24 GB VRAM for partial offload - **Recommended**: 32+ GB VRAM for full GPU offload - **Optimal**: Blackwell B200 for max NVFP4 throughput ## Quantization Quantized from Qwen/Qwen3.6-35B-A3B BF16 weights using [llama.cpp](https://github.com/ggerganov/llama.cpp) (`llama-quantize.exe NVFP4`). ## License Apache 2.0 - same as the base model. ## Credits - Original model: [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) - NVFP4 Fast quantization: [unsloth/Qwen3.6-35B-A3B-NVFP4-Fast](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast) - GGUF conversion: FreedomAISVR