File size: 2,949 Bytes
174afb7
 
2b174f0
 
 
 
 
 
 
174afb7
874da54
 
0d5f9a7
2b174f0
 
0d5f9a7
2b174f0
0d5f9a7
174afb7
 
0d5f9a7
874da54
2b174f0
 
 
874da54
0d5f9a7
874da54
2b174f0
0d5f9a7
 
874da54
0d5f9a7
874da54
0d5f9a7
874da54
0d5f9a7
874da54
0d5f9a7
 
 
 
 
 
874da54
0d5f9a7
874da54
0d5f9a7
874da54
 
0d5f9a7
 
 
 
 
174afb7
874da54
0d5f9a7
b6c7dbd
0d5f9a7
 
 
 
174afb7
874da54
0d5f9a7
306e7a4
0d5f9a7
 
 
 
 
306e7a4
 
174afb7
874da54
0d5f9a7
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
---
license: apache-2.0
base_model:
- Qwen/Qwen3.6-35B-A3B
- nightmedia/Qwen3.6-35B-A3B-Holo3-Qwopus-AgentWorld-qx64-hi-mlx
- Hcompany/Holo3-35B-A3B
- samuelcardillo/Qwopus-MoE-35B-A3B
- Qwen/Qwen-AgentWorld-35B-A3B
- llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved
tags:
- gguf
- llama.cpp
- qwen3_5_moe
- moe
- merge
- vision-language
- security
- image-text-to-text
---

# RazorStrike-v1 GGUF

GGUF build of `lancejames221b/razorstrike-v1`. **Same lineage as the current MLX repo**, not the older DARE-TIES build documented in `lancejames221b/razorstrike-v1-bf16` (that lineage was superseded 2026-07-21 and has no GGUF/MLX quant under this org).

**Base**: a 4-bit-equivalent quantization of `nightmedia/Qwen3.6-35B-A3B-Holo3-Qwopus-AgentWorld-qx64-hi-mlx` (itself a merge of `Hcompany/Holo3-35B-A3B`, `samuelcardillo/Qwopus-MoE-35B-A3B`, and `llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved` with `Qwen/Qwen-AgentWorld-35B-A3B`) with the Qwen3.6-35B-A3B vision tower spliced back in — see `lancejames221b/razorstrike-v1`'s README for the full provenance and known-issues notes (including the inherited repetition/looping caveat).

## Files

- `razorstrike-v1-IQ4_XS.gguf` — 4.37 BPW IQ4_XS text model, converted with llama.cpp from a raw-HF-compatible reconstruction of the source above.
- `mmproj.gguf` — matching multimodal projector for image input.
- `RAZORSTRIKE_V1_GGUF_MANIFEST.json` — build and verification notes.

The bf16 GGUF intermediate was generated and smoke-tested locally but is not uploaded because it is ~65 GiB.

## Fix in this build

Previous GGUF attempts generated corrupted text because the source checkpoint was MLX-derived. MLX stores non-linear-attention text RMSNorm weights with the `+1.0` offset already baked in, while the llama.cpp Qwen converter expects raw HF-style weights and applies its own `+1.0` transform for those tensors. This build reconstructs the HF-naming checkpoint with:

- `model.language_model.*norm.weight` shifted by `-1.0`, except `linear_attn.norm.weight`.
- `linear_attn.norm.weight` kept unchanged.
- vision tower norms kept unchanged.
- MLX `switch_mlp.{gate,up}_proj.weight` merged into raw HF `experts.gate_up_proj`.
- MLX `switch_mlp.down_proj.weight` renamed to raw HF `experts.down_proj`.
- llama.cpp conversion run with `--no-mtp` to avoid false extra-layer metadata.

## Local verification

Verified on the fixed IQ4_XS GGUF through `llama-server`:

```bash
llama-server \
  -m razorstrike-v1-IQ4_XS.gguf \
  --mmproj mmproj.gguf \
  -c 4096 \
  -fit off
```

Text smoke test:

```text
System: <|think_off|>
User: What is 17 times 24? Answer directly.
Assistant: 408
```

Image smoke test with a generated PNG containing a green square, yellow circle, and `TEST-42`:

```text
Assistant:
- Green square
- Yellow circle
- Text: "TEST-42"
```

## License

Apache-2.0, matching the Qwen3.6 lineage and the current RazorStrike-v1 model card.