--- license: apache-2.0 base_model: - Qwen/Qwen3.6-35B-A3B - nightmedia/Qwen3.6-35B-A3B-Holo3-Qwopus-AgentWorld-qx64-hi-mlx - Hcompany/Holo3-35B-A3B - samuelcardillo/Qwopus-MoE-35B-A3B - Qwen/Qwen-AgentWorld-35B-A3B - llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved tags: - gguf - llama.cpp - qwen3_5_moe - moe - merge - vision-language - security - image-text-to-text --- # RazorStrike-v1 GGUF GGUF build of `lancejames221b/razorstrike-v1`. **Same lineage as the current MLX repo**, not the older DARE-TIES build documented in `lancejames221b/razorstrike-v1-bf16` (that lineage was superseded 2026-07-21 and has no GGUF/MLX quant under this org). **Base**: a 4-bit-equivalent quantization of `nightmedia/Qwen3.6-35B-A3B-Holo3-Qwopus-AgentWorld-qx64-hi-mlx` (itself a merge of `Hcompany/Holo3-35B-A3B`, `samuelcardillo/Qwopus-MoE-35B-A3B`, and `llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved` with `Qwen/Qwen-AgentWorld-35B-A3B`) with the Qwen3.6-35B-A3B vision tower spliced back in — see `lancejames221b/razorstrike-v1`'s README for the full provenance and known-issues notes (including the inherited repetition/looping caveat). ## Files - `razorstrike-v1-IQ4_XS.gguf` — 4.37 BPW IQ4_XS text model, converted with llama.cpp from a raw-HF-compatible reconstruction of the source above. - `mmproj.gguf` — matching multimodal projector for image input. - `RAZORSTRIKE_V1_GGUF_MANIFEST.json` — build and verification notes. The bf16 GGUF intermediate was generated and smoke-tested locally but is not uploaded because it is ~65 GiB. ## Fix in this build Previous GGUF attempts generated corrupted text because the source checkpoint was MLX-derived. MLX stores non-linear-attention text RMSNorm weights with the `+1.0` offset already baked in, while the llama.cpp Qwen converter expects raw HF-style weights and applies its own `+1.0` transform for those tensors. This build reconstructs the HF-naming checkpoint with: - `model.language_model.*norm.weight` shifted by `-1.0`, except `linear_attn.norm.weight`. - `linear_attn.norm.weight` kept unchanged. - vision tower norms kept unchanged. - MLX `switch_mlp.{gate,up}_proj.weight` merged into raw HF `experts.gate_up_proj`. - MLX `switch_mlp.down_proj.weight` renamed to raw HF `experts.down_proj`. - llama.cpp conversion run with `--no-mtp` to avoid false extra-layer metadata. ## Local verification Verified on the fixed IQ4_XS GGUF through `llama-server`: ```bash llama-server \ -m razorstrike-v1-IQ4_XS.gguf \ --mmproj mmproj.gguf \ -c 4096 \ -fit off ``` Text smoke test: ```text System: <|think_off|> User: What is 17 times 24? Answer directly. Assistant: 408 ``` Image smoke test with a generated PNG containing a green square, yellow circle, and `TEST-42`: ```text Assistant: - Green square - Yellow circle - Text: "TEST-42" ``` ## License Apache-2.0, matching the Qwen3.6 lineage and the current RazorStrike-v1 model card.