--- license: apache-2.0 base_model: - Qwen/Qwen3.6-35B-A3B pipeline_tag: text-generation library_name: transformers tags: - moe - mixture-of-experts - qwen3_5_moe - unsloth - sft - lora - bf16 - multimodal - vision - vlm --- ![grok-image-9f997638-c16b-47b7-8ea4-1675a600eafe](https://cdn-uploads.huggingface.co/production/uploads/69a27f2d114e4ac9de4dafc7/1jlDhhHzMbR4pyPR3XDAz.jpeg) # PINQWEN-3.6-35B-CLEAN-BF16 Full 16-bit merged weights of PINQWEN-3.6-35B-CLEAN, a supervised fine-tune of Qwen/Qwen3.6-35B-A3B by Blackfrost-AI. General-purpose assistant tuned for reasoning and agentic tool-use. Safety-aligned (CLEAN) variant. ## Model Description PINQWEN-3.6-35B-CLEAN is a supervised fine-tune (SFT) of Alibaba's **Qwen/Qwen3.6-35B-A3B** base model. This repository holds the **BF16** release: the full 16-bit merged weights. - **Developer:** Blackfrost-AI - **Base model:** [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (Alibaba Qwen team) - **Architecture:** `qwen3_5_moe` (`Qwen3_5MoeForConditionalGeneration`) — a Mixture-of-Experts model with 256 experts (8 routed + 1 shared per token), ~36.97B total parameters and ~3B active per token. Hybrid Gated-DeltaNet + gated attention. Unified vision-language model with 262K native context. Thinking-on by default. - **Variant:** CLEAN — the aligned variant. This is **not** an abliterated model; the base model's safety alignment is preserved. - **License:** Apache-2.0 - **Language/modality:** Text (see Limitations regarding vision). A 4-bit NVFP4 quantization of this same model is released separately at **Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-NVFP4**. ## PINQWEN & The Void **PINQWEN** is Blackfrost-AI's model name for this Qwen3.6-based series. The **CLEAN** suffix marks the safety-aligned build. The model was fine-tuned on **The Void (v4)**, Blackfrost's proprietary distillation corpus: roughly **5,032 high-quality multi-turn examples** that blend distilled reasoning / chain-of-thought and broad knowledge with agentic, ReAct-style tool-use trajectories drawn from multiple frontier teacher models. Refusal and denial data is scrubbed from the corpus to protect MoE routing quality. The corpus construction method is proprietary and not disclosed here. ## Training Procedure - **Method:** Supervised fine-tuning (SFT) with **bf16 LoRA** via [Unsloth](https://github.com/unslothai/unsloth). LoRA was run in bf16 rather than 4-bit — 4-bit QLoRA degrades this MoE. - **LoRA configuration:** rank 32, alpha 32, dropout 0. Adapters were applied to the attention projections (q/k/v/o) **and** the MoE expert projections (`gate_up_proj`, `down_proj`), giving **1.86B trainable parameters (5.04% of the model)**. - **Schedule:** 3 epochs, learning rate 2e-4, linear schedule, length-grouped batching. - **Objective:** Multi-turn SFT with loss masked to assistant turns only. - **Hardware:** 8× NVIDIA B200, DDP. - **Release:** The LoRA adapter was merged back to 16-bit for this release. ## Vision (multimodal) This is a **vision-language model**. It carries a full vision tower inherited from the Qwen3.6-35B vision-language base, so it accepts **images and video** alongside text. **Scope of Blackfrost's work:** training here was **text-only**; the vision tower is **inherited unchanged from the base** and was **not tuned or evaluated by Blackfrost**. Multimodal behavior tracks the base model — validate it for your use case. ## Intended Uses General-purpose text assistant for reasoning and agentic / tool-use (ReAct-style) tasks. ## Limitations - **Text-focused SFT.** The base model is vision-language, but vision was **not** specifically tuned or evaluated in this work. Treat vision behavior as untuned base-model behavior. - **No public benchmark numbers are claimed yet.** Internal evaluations are pending; no scores are reported here. - The model may inherit base-model limitations and can hallucinate. **Verify important facts.** ## How to Use ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "Blackfrost-AI/PINQWEN-3.6-35B-CLEAN-BF16" tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True, ) messages = [ {"role": "user", "content": "Explain what a Mixture-of-Experts model is in two sentences."}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_tensors="pt", ).to(model.device) outputs = model.generate(inputs, max_new_tokens=512) print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True)) ``` ## License & Attribution Released under **Apache-2.0**. Built on **Qwen/Qwen3.6-35B-A3B** by Alibaba's Qwen team (Apache-2.0). Fine-tuning and release by Blackfrost-AI. ## Responsible Use This model retains the base model's safety alignment. Do not use it to generate content that exploits minors or promotes self-harm. Standard responsible-use expectations apply.