apetersson commited on
Commit
26e1323
·
verified ·
1 Parent(s): 98a692e

Add files using upload-large-folder tool

Browse files
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf filter=lfs diff=lfs merge=lfs -text
37
+ DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf filter=lfs diff=lfs merge=lfs -text
BUILD_MANIFEST.json ADDED
@@ -0,0 +1,1046 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema": "deepseek-model-tools.build128.v1",
3
+ "artifact": "ds4-quality128",
4
+ "output": "DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128",
5
+ "engine": "ds4",
6
+ "dspark": true,
7
+ "created_at": "2026-08-01T13:29:21.384605+00:00",
8
+ "source": {
9
+ "path": "apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8",
10
+ "only_weight_source": true,
11
+ "config_sha256": "6c8f3d2d3b48707541b88f32f22ef3f0f8a6b57d8523281e2b8d3cdb0ae9a023",
12
+ "index_sha256": "98efab455cf08dfbbbaaba6f570e1bf10bf927d2b4c3c453a59c2f6f0e3be92b",
13
+ "abliteration_manifest_sha256": "8e6f40f3d92542720aabdafb728f90e3397fba68c6a8d7ba91772b5dd41bc767",
14
+ "edited_tensors": 36,
15
+ "upstream_model": "deepseek-ai/DeepSeek-V4-Flash-0731",
16
+ "upstream_revision": "9e165c30e2704aec5d9d593cce3eebd58bbef1cb",
17
+ "weight_shards": [
18
+ {
19
+ "filename": "model-00001-of-00048.safetensors",
20
+ "bytes": 1059061856,
21
+ "sha256": "f3668ba4cccf1ca6a7eb84e888fb92c1cdc7204d472ba9db771e6fd3abf6b874",
22
+ "upload_mode": "lfs"
23
+ },
24
+ {
25
+ "filename": "model-00002-of-00048.safetensors",
26
+ "bytes": 3566321192,
27
+ "sha256": "77b26c939a0e25b3113c8d6bb04e1901a748bd4a7d2589e3bfdaabdf1e9bba14",
28
+ "upload_mode": "lfs"
29
+ },
30
+ {
31
+ "filename": "model-00003-of-00048.safetensors",
32
+ "bytes": 3566321192,
33
+ "sha256": "412abf4c906faadc221ef0cb50f90fe20bde8454a08ad4dc2364b6b79e7fda5c",
34
+ "upload_mode": "lfs"
35
+ },
36
+ {
37
+ "filename": "model-00004-of-00048.safetensors",
38
+ "bytes": 3596229272,
39
+ "sha256": "9610f56bc587fb0ff9a8b68a60299482ee8c433fe5b5587e4257aca98add4a2e",
40
+ "upload_mode": "lfs"
41
+ },
42
+ {
43
+ "filename": "model-00005-of-00048.safetensors",
44
+ "bytes": 3568768976,
45
+ "sha256": "f87a5ac7b8becc31f9c3169afd3a6f33fb82b4af9e21022e3755a10bc28f0180",
46
+ "upload_mode": "lfs"
47
+ },
48
+ {
49
+ "filename": "model-00006-of-00048.safetensors",
50
+ "bytes": 3590024776,
51
+ "sha256": "4a4f3764e3fc772b9fba67f0a44ef68e18f178b6f00faa80b75db549e51894cd",
52
+ "upload_mode": "lfs"
53
+ },
54
+ {
55
+ "filename": "model-00007-of-00048.safetensors",
56
+ "bytes": 3568768976,
57
+ "sha256": "df81bb80e27a689e01fa579eebd6499f86e0b6105f7fea18961aa5eebbbee9bc",
58
+ "upload_mode": "lfs"
59
+ },
60
+ {
61
+ "filename": "model-00008-of-00048.safetensors",
62
+ "bytes": 3590024776,
63
+ "sha256": "224968d2b27f8669365ec08657a768dfec40da0585f85f302a31495931f6a526",
64
+ "upload_mode": "lfs"
65
+ },
66
+ {
67
+ "filename": "model-00009-of-00048.safetensors",
68
+ "bytes": 3568768976,
69
+ "sha256": "04d69ef1071fff8721c62968c200a5583122b59b015e9ef9b2978bfed271b2b7",
70
+ "upload_mode": "lfs"
71
+ },
72
+ {
73
+ "filename": "model-00010-of-00048.safetensors",
74
+ "bytes": 3590024776,
75
+ "sha256": "627145f4ebeb1cc3f5bdd03416b8cb7370b3c96974853cbeb8e5516ad5713e49",
76
+ "upload_mode": "lfs"
77
+ },
78
+ {
79
+ "filename": "model-00011-of-00048.safetensors",
80
+ "bytes": 3568768976,
81
+ "sha256": "e4b8e601dcbebe902e0102e7b098b670a121cb8b9564dd719fc41d782c8416e0",
82
+ "upload_mode": "lfs"
83
+ },
84
+ {
85
+ "filename": "model-00012-of-00048.safetensors",
86
+ "bytes": 3590026352,
87
+ "sha256": "d200bcd9186f15344b8271f845bd81f782509c5997ee493b1a96079e91a1a8ac",
88
+ "upload_mode": "lfs"
89
+ },
90
+ {
91
+ "filename": "model-00013-of-00048.safetensors",
92
+ "bytes": 3568770544,
93
+ "sha256": "9c36e454195177bafdb866bbfd052f69dd727fafbe328c0d5d4d659ab49389b2",
94
+ "upload_mode": "lfs"
95
+ },
96
+ {
97
+ "filename": "model-00014-of-00048.safetensors",
98
+ "bytes": 3590026352,
99
+ "sha256": "ce44ce426be0ed8b21141ac1659ca108d2d0705b5b442a7691511bc38831bf9e",
100
+ "upload_mode": "lfs"
101
+ },
102
+ {
103
+ "filename": "model-00015-of-00048.safetensors",
104
+ "bytes": 3568770544,
105
+ "sha256": "f9b40ffad55e01cbfc012c737640f882ad7161026ea1c6d878d272c655e74b02",
106
+ "upload_mode": "lfs"
107
+ },
108
+ {
109
+ "filename": "model-00016-of-00048.safetensors",
110
+ "bytes": 3590026352,
111
+ "sha256": "acf5cbaf18a2d2e005358b0a81527ed70bc92f352ab46f658172a0505983bb2c",
112
+ "upload_mode": "lfs"
113
+ },
114
+ {
115
+ "filename": "model-00017-of-00048.safetensors",
116
+ "bytes": 3568770544,
117
+ "sha256": "d67716007b48845ed1de1473ca0ab15908a9dff8b592bb62ab95137cc67cf198",
118
+ "upload_mode": "lfs"
119
+ },
120
+ {
121
+ "filename": "model-00018-of-00048.safetensors",
122
+ "bytes": 3590026352,
123
+ "sha256": "ac550ef17db3b8c95c2a7858fefa7ba08c94b5e6feafa8bfdb3f9763612fd3b1",
124
+ "upload_mode": "lfs"
125
+ },
126
+ {
127
+ "filename": "model-00019-of-00048.safetensors",
128
+ "bytes": 3568770544,
129
+ "sha256": "c31a222ce5c7e0d02b8c2614badd16781cb3979f7e516d5293ab976ee670f34f",
130
+ "upload_mode": "lfs"
131
+ },
132
+ {
133
+ "filename": "model-00020-of-00048.safetensors",
134
+ "bytes": 3590026352,
135
+ "sha256": "a27655e314a1919d9c7042b1c7f842fa8c8f1bf77df4c0a9586371b4ab3029d9",
136
+ "upload_mode": "lfs"
137
+ },
138
+ {
139
+ "filename": "model-00021-of-00048.safetensors",
140
+ "bytes": 3568770544,
141
+ "sha256": "4068e1c0a59dcd4b7a22c09cf870e86439407f1eb8c8669b20d81d3294e4ca60",
142
+ "upload_mode": "lfs"
143
+ },
144
+ {
145
+ "filename": "model-00022-of-00048.safetensors",
146
+ "bytes": 3590026352,
147
+ "sha256": "eb87aef24cbf17c98f8ab4803590606db89acc69c83b9d4f44919bac328d830d",
148
+ "upload_mode": "lfs"
149
+ },
150
+ {
151
+ "filename": "model-00023-of-00048.safetensors",
152
+ "bytes": 3568770544,
153
+ "sha256": "4b8d427c9d2f285986d18b465f3ffcb73c12a109e1dedc8b0a3b6ac83e0e1ab7",
154
+ "upload_mode": "lfs"
155
+ },
156
+ {
157
+ "filename": "model-00024-of-00048.safetensors",
158
+ "bytes": 3590026352,
159
+ "sha256": "2cf277a128b4f51dcd13c130db18e325006962eec0a1f1e90596d32085980daa",
160
+ "upload_mode": "lfs"
161
+ },
162
+ {
163
+ "filename": "model-00025-of-00048.safetensors",
164
+ "bytes": 3568770544,
165
+ "sha256": "dd813e4837baa77bd5fac84d5dce225e6e949c7e781cf40621c06081b6bfcbe1",
166
+ "upload_mode": "lfs"
167
+ },
168
+ {
169
+ "filename": "model-00026-of-00048.safetensors",
170
+ "bytes": 3590026352,
171
+ "sha256": "a4c8f53dbf43374c5a277e992973b0f5b9e916022dd17cecc024c4290165eae8",
172
+ "upload_mode": "lfs"
173
+ },
174
+ {
175
+ "filename": "model-00027-of-00048.safetensors",
176
+ "bytes": 3568770544,
177
+ "sha256": "70050f03122a0050bad34b03eb0bb88def42f2c3ea99c525c1af776409efb76e",
178
+ "upload_mode": "lfs"
179
+ },
180
+ {
181
+ "filename": "model-00028-of-00048.safetensors",
182
+ "bytes": 3590026352,
183
+ "sha256": "51113c3fce251a429452829adea388497184a80e0c0023febbb88c4068f65842",
184
+ "upload_mode": "lfs"
185
+ },
186
+ {
187
+ "filename": "model-00029-of-00048.safetensors",
188
+ "bytes": 3568770544,
189
+ "sha256": "f7b9198bb1e6c05ab5e88bf22fbe78d3f67f789a494d4c1be763787ee11f5f1e",
190
+ "upload_mode": "lfs"
191
+ },
192
+ {
193
+ "filename": "model-00030-of-00048.safetensors",
194
+ "bytes": 3590026352,
195
+ "sha256": "51e9984913ff3abbb7649d8738e764360c165401a9c31d9b53715c1fbe313061",
196
+ "upload_mode": "lfs"
197
+ },
198
+ {
199
+ "filename": "model-00031-of-00048.safetensors",
200
+ "bytes": 3568770544,
201
+ "sha256": "d31d4f707391eca341eaa6025f3e8ecf21ee254340bd8f11d494200aa8923447",
202
+ "upload_mode": "lfs"
203
+ },
204
+ {
205
+ "filename": "model-00032-of-00048.safetensors",
206
+ "bytes": 3590026352,
207
+ "sha256": "b7ecad0122f051d1e2a0953312640ed53ac7da3875766e1e70ff0d0666c28087",
208
+ "upload_mode": "lfs"
209
+ },
210
+ {
211
+ "filename": "model-00033-of-00048.safetensors",
212
+ "bytes": 3568770544,
213
+ "sha256": "7bf5a6c9a2774698ebf819f3abf20604376fd75daecb6ecb6499df964e4447af",
214
+ "upload_mode": "lfs"
215
+ },
216
+ {
217
+ "filename": "model-00034-of-00048.safetensors",
218
+ "bytes": 3590026352,
219
+ "sha256": "b0de49f523078e8766c5deb61be4fc73da35271f3bd354bb3e7c71e975cd8be5",
220
+ "upload_mode": "lfs"
221
+ },
222
+ {
223
+ "filename": "model-00035-of-00048.safetensors",
224
+ "bytes": 3568770544,
225
+ "sha256": "471a793bb28aa16c95d0cad6e9194191fd9c2e0657c9e35917bd7856ddbf8187",
226
+ "upload_mode": "lfs"
227
+ },
228
+ {
229
+ "filename": "model-00036-of-00048.safetensors",
230
+ "bytes": 3590026352,
231
+ "sha256": "002bdca78f533708438ed04ca42139fb67f0325571adb29645f5ef38be54b066",
232
+ "upload_mode": "lfs"
233
+ },
234
+ {
235
+ "filename": "model-00037-of-00048.safetensors",
236
+ "bytes": 3568770544,
237
+ "sha256": "71f148a710c0ff1b505d6e285b1e782eef87d14d5fa4f5411e1faa4c06210708",
238
+ "upload_mode": "lfs"
239
+ },
240
+ {
241
+ "filename": "model-00038-of-00048.safetensors",
242
+ "bytes": 3590026352,
243
+ "sha256": "ab90e440639d4b0de9a9e2f7a48238e7ccdc08c37d40cf91d1a380ba26c1938a",
244
+ "upload_mode": "lfs"
245
+ },
246
+ {
247
+ "filename": "model-00039-of-00048.safetensors",
248
+ "bytes": 3568770544,
249
+ "sha256": "26f3acc9d23d426edabef80260ac4030aad030553356c47b1d5d47667c0813c3",
250
+ "upload_mode": "lfs"
251
+ },
252
+ {
253
+ "filename": "model-00040-of-00048.safetensors",
254
+ "bytes": 3590026352,
255
+ "sha256": "549d59fc74aa7bba55678da85e5a369ebe8017763c4d5f83359e15323659afcf",
256
+ "upload_mode": "lfs"
257
+ },
258
+ {
259
+ "filename": "model-00041-of-00048.safetensors",
260
+ "bytes": 3568770544,
261
+ "sha256": "1868317b4ff748d1388805c530fc7b0310397fdef8ac1c0bd7f034f14abe83ce",
262
+ "upload_mode": "lfs"
263
+ },
264
+ {
265
+ "filename": "model-00042-of-00048.safetensors",
266
+ "bytes": 3590026352,
267
+ "sha256": "11e81dc83b2f6ab688b9205552a147cf781d73ee791a8d54c763e0a461006def",
268
+ "upload_mode": "lfs"
269
+ },
270
+ {
271
+ "filename": "model-00043-of-00048.safetensors",
272
+ "bytes": 3568770544,
273
+ "sha256": "bc1dc626c7ec234f58807443ecbdec2a5927be771339697cbac3dca94c6ed9a0",
274
+ "upload_mode": "lfs"
275
+ },
276
+ {
277
+ "filename": "model-00044-of-00048.safetensors",
278
+ "bytes": 3590026352,
279
+ "sha256": "a20dfe1e58e2744ed8c82cdca63326ed62c2f024361a072ecb827edd9e0287bc",
280
+ "upload_mode": "lfs"
281
+ },
282
+ {
283
+ "filename": "model-00045-of-00048.safetensors",
284
+ "bytes": 1059332516,
285
+ "sha256": "a5be6aed7b84fc87ec42b5d24ba0b0d67f253a3906fcd99c13f4f7be5958fc00",
286
+ "upload_mode": "lfs"
287
+ },
288
+ {
289
+ "filename": "model-00046-of-00048.safetensors",
290
+ "bytes": 3610455184,
291
+ "sha256": "da1e0108b463653b292a648aef58b4dbbeac03a8c213c53829b99c6d84d941c9",
292
+ "upload_mode": "lfs"
293
+ },
294
+ {
295
+ "filename": "model-00047-of-00048.safetensors",
296
+ "bytes": 3560111960,
297
+ "sha256": "b8b41236c1ccf4378a3ae2ccc22a784ca1a5681f13a1d64c262836b788a0f093",
298
+ "upload_mode": "lfs"
299
+ },
300
+ {
301
+ "filename": "model-00048-of-00048.safetensors",
302
+ "bytes": 3692775244,
303
+ "sha256": "d8b68325b29a33f7ee43cafcaa69c5de8c1de7c1f6007758961ad4ed11aa74db",
304
+ "upload_mode": "lfs"
305
+ }
306
+ ],
307
+ "weight_shard_bytes": 166886535336,
308
+ "weight_shard_hashes_verified": true,
309
+ "weight_shard_hashes_verified_at": "2026-08-01T11:53:46.973992+00:00",
310
+ "post_hash_stat_snapshot": {
311
+ "model-00001-of-00048.safetensors": {
312
+ "bytes": 1059061856,
313
+ "mtime_ns": 1785506342100000000,
314
+ "ctime_ns": 1785506342100000000,
315
+ "device": 16777242,
316
+ "inode": 10534358
317
+ },
318
+ "model-00002-of-00048.safetensors": {
319
+ "bytes": 3566321192,
320
+ "mtime_ns": 1785506344340000000,
321
+ "ctime_ns": 1785506344340000000,
322
+ "device": 16777242,
323
+ "inode": 10538402
324
+ },
325
+ "model-00003-of-00048.safetensors": {
326
+ "bytes": 3566321192,
327
+ "mtime_ns": 1785506346560000000,
328
+ "ctime_ns": 1785506346560000000,
329
+ "device": 16777242,
330
+ "inode": 10552010
331
+ },
332
+ "model-00004-of-00048.safetensors": {
333
+ "bytes": 3596229272,
334
+ "mtime_ns": 1785506349630000000,
335
+ "ctime_ns": 1785506349630000000,
336
+ "device": 16777242,
337
+ "inode": 10565618
338
+ },
339
+ "model-00005-of-00048.safetensors": {
340
+ "bytes": 3568768976,
341
+ "mtime_ns": 1785506351960000000,
342
+ "ctime_ns": 1785506351960000000,
343
+ "device": 16777242,
344
+ "inode": 10579340
345
+ },
346
+ "model-00006-of-00048.safetensors": {
347
+ "bytes": 3590024776,
348
+ "mtime_ns": 1785506354260000000,
349
+ "ctime_ns": 1785506354260000000,
350
+ "device": 16777242,
351
+ "inode": 10592957
352
+ },
353
+ "model-00007-of-00048.safetensors": {
354
+ "bytes": 3568768976,
355
+ "mtime_ns": 1785506356530000000,
356
+ "ctime_ns": 1785506356530000000,
357
+ "device": 16777242,
358
+ "inode": 10606655
359
+ },
360
+ "model-00008-of-00048.safetensors": {
361
+ "bytes": 3590024776,
362
+ "mtime_ns": 1785506358750000000,
363
+ "ctime_ns": 1785506358750000000,
364
+ "device": 16777242,
365
+ "inode": 10620272
366
+ },
367
+ "model-00009-of-00048.safetensors": {
368
+ "bytes": 3568768976,
369
+ "mtime_ns": 1785506362250000000,
370
+ "ctime_ns": 1785506362250000000,
371
+ "device": 16777242,
372
+ "inode": 10633970
373
+ },
374
+ "model-00010-of-00048.safetensors": {
375
+ "bytes": 3590024776,
376
+ "mtime_ns": 1785506365890000000,
377
+ "ctime_ns": 1785506365890000000,
378
+ "device": 16777242,
379
+ "inode": 10647587
380
+ },
381
+ "model-00011-of-00048.safetensors": {
382
+ "bytes": 3568768976,
383
+ "mtime_ns": 1785506369620000000,
384
+ "ctime_ns": 1785506369620000000,
385
+ "device": 16777242,
386
+ "inode": 10661285
387
+ },
388
+ "model-00012-of-00048.safetensors": {
389
+ "bytes": 3590026352,
390
+ "mtime_ns": 1785506377140000000,
391
+ "ctime_ns": 1785506377140000000,
392
+ "device": 16777242,
393
+ "inode": 10674902
394
+ },
395
+ "model-00013-of-00048.safetensors": {
396
+ "bytes": 3568770544,
397
+ "mtime_ns": 1785506386820000000,
398
+ "ctime_ns": 1785506386820000000,
399
+ "device": 16777242,
400
+ "inode": 10688602
401
+ },
402
+ "model-00014-of-00048.safetensors": {
403
+ "bytes": 3590026352,
404
+ "mtime_ns": 1785506396120000000,
405
+ "ctime_ns": 1785506396120000000,
406
+ "device": 16777242,
407
+ "inode": 10702221
408
+ },
409
+ "model-00015-of-00048.safetensors": {
410
+ "bytes": 3568770544,
411
+ "mtime_ns": 1785506405630000000,
412
+ "ctime_ns": 1785506405630000000,
413
+ "device": 16777242,
414
+ "inode": 10715921
415
+ },
416
+ "model-00016-of-00048.safetensors": {
417
+ "bytes": 3590026352,
418
+ "mtime_ns": 1785506415210000000,
419
+ "ctime_ns": 1785506415210000000,
420
+ "device": 16777242,
421
+ "inode": 10729540
422
+ },
423
+ "model-00017-of-00048.safetensors": {
424
+ "bytes": 3568770544,
425
+ "mtime_ns": 1785506424840000000,
426
+ "ctime_ns": 1785506424840000000,
427
+ "device": 16777242,
428
+ "inode": 10743240
429
+ },
430
+ "model-00018-of-00048.safetensors": {
431
+ "bytes": 3590026352,
432
+ "mtime_ns": 1785506432280000000,
433
+ "ctime_ns": 1785506432280000000,
434
+ "device": 16777242,
435
+ "inode": 10756859
436
+ },
437
+ "model-00019-of-00048.safetensors": {
438
+ "bytes": 3568770544,
439
+ "mtime_ns": 1785506439960000000,
440
+ "ctime_ns": 1785506439960000000,
441
+ "device": 16777242,
442
+ "inode": 10770559
443
+ },
444
+ "model-00020-of-00048.safetensors": {
445
+ "bytes": 3590026352,
446
+ "mtime_ns": 1785506449550000000,
447
+ "ctime_ns": 1785506449550000000,
448
+ "device": 16777242,
449
+ "inode": 10784178
450
+ },
451
+ "model-00021-of-00048.safetensors": {
452
+ "bytes": 3568770544,
453
+ "mtime_ns": 1785506459270000000,
454
+ "ctime_ns": 1785506459270000000,
455
+ "device": 16777242,
456
+ "inode": 10797878
457
+ },
458
+ "model-00022-of-00048.safetensors": {
459
+ "bytes": 3590026352,
460
+ "mtime_ns": 1785506468940000000,
461
+ "ctime_ns": 1785506468940000000,
462
+ "device": 16777242,
463
+ "inode": 10811497
464
+ },
465
+ "model-00023-of-00048.safetensors": {
466
+ "bytes": 3568770544,
467
+ "mtime_ns": 1785506478610000000,
468
+ "ctime_ns": 1785506478610000000,
469
+ "device": 16777242,
470
+ "inode": 10825197
471
+ },
472
+ "model-00024-of-00048.safetensors": {
473
+ "bytes": 3590026352,
474
+ "mtime_ns": 1785506488130000000,
475
+ "ctime_ns": 1785506488130000000,
476
+ "device": 16777242,
477
+ "inode": 10838816
478
+ },
479
+ "model-00025-of-00048.safetensors": {
480
+ "bytes": 3568770544,
481
+ "mtime_ns": 1785506495770000000,
482
+ "ctime_ns": 1785506495770000000,
483
+ "device": 16777242,
484
+ "inode": 10852516
485
+ },
486
+ "model-00026-of-00048.safetensors": {
487
+ "bytes": 3590026352,
488
+ "mtime_ns": 1785506505150000000,
489
+ "ctime_ns": 1785506505150000000,
490
+ "device": 16777242,
491
+ "inode": 10866135
492
+ },
493
+ "model-00027-of-00048.safetensors": {
494
+ "bytes": 3568770544,
495
+ "mtime_ns": 1785506514650000000,
496
+ "ctime_ns": 1785506514650000000,
497
+ "device": 16777242,
498
+ "inode": 10879835
499
+ },
500
+ "model-00028-of-00048.safetensors": {
501
+ "bytes": 3590026352,
502
+ "mtime_ns": 1785506524000000000,
503
+ "ctime_ns": 1785506524000000000,
504
+ "device": 16777242,
505
+ "inode": 10893454
506
+ },
507
+ "model-00029-of-00048.safetensors": {
508
+ "bytes": 3568770544,
509
+ "mtime_ns": 1785506533430000000,
510
+ "ctime_ns": 1785506533430000000,
511
+ "device": 16777242,
512
+ "inode": 10907154
513
+ },
514
+ "model-00030-of-00048.safetensors": {
515
+ "bytes": 3590026352,
516
+ "mtime_ns": 1785506540820000000,
517
+ "ctime_ns": 1785506540820000000,
518
+ "device": 16777242,
519
+ "inode": 10920773
520
+ },
521
+ "model-00031-of-00048.safetensors": {
522
+ "bytes": 3568770544,
523
+ "mtime_ns": 1785506550210000000,
524
+ "ctime_ns": 1785506550210000000,
525
+ "device": 16777242,
526
+ "inode": 10934473
527
+ },
528
+ "model-00032-of-00048.safetensors": {
529
+ "bytes": 3590026352,
530
+ "mtime_ns": 1785506559630000000,
531
+ "ctime_ns": 1785506559630000000,
532
+ "device": 16777242,
533
+ "inode": 10948092
534
+ },
535
+ "model-00033-of-00048.safetensors": {
536
+ "bytes": 3568770544,
537
+ "mtime_ns": 1785506567220000000,
538
+ "ctime_ns": 1785506567220000000,
539
+ "device": 16777242,
540
+ "inode": 10961792
541
+ },
542
+ "model-00034-of-00048.safetensors": {
543
+ "bytes": 3590026352,
544
+ "mtime_ns": 1785506576510000000,
545
+ "ctime_ns": 1785506576510000000,
546
+ "device": 16777242,
547
+ "inode": 10975411
548
+ },
549
+ "model-00035-of-00048.safetensors": {
550
+ "bytes": 3568770544,
551
+ "mtime_ns": 1785506585920000000,
552
+ "ctime_ns": 1785506585920000000,
553
+ "device": 16777242,
554
+ "inode": 10989111
555
+ },
556
+ "model-00036-of-00048.safetensors": {
557
+ "bytes": 3590026352,
558
+ "mtime_ns": 1785506595350000000,
559
+ "ctime_ns": 1785506595350000000,
560
+ "device": 16777242,
561
+ "inode": 11002730
562
+ },
563
+ "model-00037-of-00048.safetensors": {
564
+ "bytes": 3568770544,
565
+ "mtime_ns": 1785506602780000000,
566
+ "ctime_ns": 1785506602780000000,
567
+ "device": 16777242,
568
+ "inode": 11016430
569
+ },
570
+ "model-00038-of-00048.safetensors": {
571
+ "bytes": 3590026352,
572
+ "mtime_ns": 1785506611980000000,
573
+ "ctime_ns": 1785506611980000000,
574
+ "device": 16777242,
575
+ "inode": 11030049
576
+ },
577
+ "model-00039-of-00048.safetensors": {
578
+ "bytes": 3568770544,
579
+ "mtime_ns": 1785506621720000000,
580
+ "ctime_ns": 1785506621720000000,
581
+ "device": 16777242,
582
+ "inode": 11043749
583
+ },
584
+ "model-00040-of-00048.safetensors": {
585
+ "bytes": 3590026352,
586
+ "mtime_ns": 1785506630860000000,
587
+ "ctime_ns": 1785506630860000000,
588
+ "device": 16777242,
589
+ "inode": 11057368
590
+ },
591
+ "model-00041-of-00048.safetensors": {
592
+ "bytes": 3568770544,
593
+ "mtime_ns": 1785506638120000000,
594
+ "ctime_ns": 1785506638120000000,
595
+ "device": 16777242,
596
+ "inode": 11071068
597
+ },
598
+ "model-00042-of-00048.safetensors": {
599
+ "bytes": 3590026352,
600
+ "mtime_ns": 1785506647820000000,
601
+ "ctime_ns": 1785506647820000000,
602
+ "device": 16777242,
603
+ "inode": 11084687
604
+ },
605
+ "model-00043-of-00048.safetensors": {
606
+ "bytes": 3568770544,
607
+ "mtime_ns": 1785506657140000000,
608
+ "ctime_ns": 1785506657140000000,
609
+ "device": 16777242,
610
+ "inode": 11098387
611
+ },
612
+ "model-00044-of-00048.safetensors": {
613
+ "bytes": 3590026352,
614
+ "mtime_ns": 1785506666700000000,
615
+ "ctime_ns": 1785506666700000000,
616
+ "device": 16777242,
617
+ "inode": 11112006
618
+ },
619
+ "model-00045-of-00048.safetensors": {
620
+ "bytes": 1059332516,
621
+ "mtime_ns": 1785506668200000000,
622
+ "ctime_ns": 1785506668200000000,
623
+ "device": 16777242,
624
+ "inode": 11125706
625
+ },
626
+ "model-00046-of-00048.safetensors": {
627
+ "bytes": 3610455184,
628
+ "mtime_ns": 1785506675660000000,
629
+ "ctime_ns": 1785506675660000000,
630
+ "device": 16777242,
631
+ "inode": 11129751
632
+ },
633
+ "model-00047-of-00048.safetensors": {
634
+ "bytes": 3560111960,
635
+ "mtime_ns": 1785506683250000000,
636
+ "ctime_ns": 1785506683250000000,
637
+ "device": 16777242,
638
+ "inode": 11143529
639
+ },
640
+ "model-00048-of-00048.safetensors": {
641
+ "bytes": 3692775244,
642
+ "mtime_ns": 1785506690860000000,
643
+ "ctime_ns": 1785506690860000000,
644
+ "device": 16777242,
645
+ "inode": 11157115
646
+ }
647
+ }
648
+ },
649
+ "profile": {
650
+ "path": "embedded-profile",
651
+ "source_profile_sha256": "6c25185500c07d4419516d30b0902e457c99bd70543c4f5bea026756d1804eb4",
652
+ "content": {
653
+ "name": "dsv4-ds4-quality128-native-mxfp4",
654
+ "engine": "ds4",
655
+ "description": "Quality-first 128 GB DS4 profile preserving exact native MXFP4 routed experts on ten sensitivity-selected layers, with IQ2_XXS/Q2_K elsewhere and native-MXFP4 DSpark down projections.",
656
+ "source": "apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8",
657
+ "output": "DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128",
658
+ "ds4_revision": "45e6cd76e849d86a8280e24ef43fa149ce9e330a",
659
+ "toolchain": {
660
+ "root": "antirez/ds4@ds4f-mxfp4",
661
+ "upstream_base_revision": "4893e0c40fba03dbc85555faeb035799aa04e0b6",
662
+ "quantizer_sha256": "f0a381f4ada808ea2afa740d964354fa327fc1235ba7cebf50874eb89fb97ac5",
663
+ "runtime_sha256": "2aaf20469b9918d6d6ab8787a02811c11228547cd979787879a03dba8a9e7824"
664
+ },
665
+ "main": {
666
+ "experts": "iq2_xxs",
667
+ "routed_w2": "q2_k",
668
+ "attention_proj": "q8_0",
669
+ "shared": "q8_0",
670
+ "preserve_indexer_f16": true,
671
+ "tensor_types": [
672
+ "output.weight=q8_0"
673
+ ],
674
+ "native_mxfp4_layers": [
675
+ 10,
676
+ 14,
677
+ 30,
678
+ 34,
679
+ 37,
680
+ 38,
681
+ 39,
682
+ 40,
683
+ 41,
684
+ 42
685
+ ],
686
+ "native_mxfp4_layer_selection": "Antirez 37-42 seed plus cross-recipe sensitivity controls 10,14,30,34; source packed I8 codes and F8_E8M0 scales are preserved exactly.",
687
+ "estimated_bytes": 102826238912
688
+ },
689
+ "dspark": {
690
+ "experts": "iq2_xxs",
691
+ "routed_w2": "mxfp4",
692
+ "block_size": 5,
693
+ "target_layers": [
694
+ 40,
695
+ 41,
696
+ 42
697
+ ],
698
+ "markov_rank": 256,
699
+ "noise_token_id": 128799,
700
+ "calibration": "target-layer proxy aliases"
701
+ },
702
+ "verification": {
703
+ "require_imatrix_strict": true,
704
+ "require_native_byte_compare": true,
705
+ "main": {
706
+ "bytes": 102826238912,
707
+ "tensors": 1328,
708
+ "tensor_types": {
709
+ "i32": 3,
710
+ "iq2_xxs": 66,
711
+ "q2_k": 33,
712
+ "mxfp4": 30,
713
+ "f16": 359,
714
+ "f32": 492,
715
+ "q8_0": 345
716
+ }
717
+ },
718
+ "dspark": {
719
+ "bytes": 7297737120,
720
+ "tensors": 81,
721
+ "tensor_types": {
722
+ "f32": 34,
723
+ "f16": 7,
724
+ "q8_0": 31,
725
+ "iq2_xxs": 6,
726
+ "mxfp4": 3
727
+ }
728
+ }
729
+ },
730
+ "calibration": {
731
+ "repo_id": "ox-ox/DeepSeek-V4-Flash-0731-GGUF",
732
+ "revision": "6d58a3a36030c3ccb969bb5759fc6ae08cd299f8",
733
+ "filename": "imatrix/DeepSeek-V4-Flash-0731-chat-v2-routed-moe-ds4-1p5m.dat",
734
+ "bytes": 450892654,
735
+ "sha256": "6fce7674df701de544e5d3351aab04e67602eddeafeb48cf70e77ebe47239eb4",
736
+ "coverage_policy": "exact 129 routed names and vector dimensions; strict mode checks only imatrix-consuming output types"
737
+ },
738
+ "template": {
739
+ "role": "metadata-tokenizer-tensor-order-only",
740
+ "repo_id": "antirez/deepseek-v4-gguf",
741
+ "revision": "1cd7b564460821938add0475a60b942c409295e0",
742
+ "filename": "DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf",
743
+ "remote_bytes": 86720111488,
744
+ "lfs_sha256": "ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0",
745
+ "xet_hash": "7da16e1025c1b856490c29c341f4e467d15cb195389c70383dade5e6108799ac",
746
+ "header_sha256": "f0e1d5e8f3b008402aa6eb32cada3873dd926c8bf5e7d00d7788eec65f09dd6d"
747
+ },
748
+ "recipe_control": {
749
+ "repo_id": "antirez/deepseek-v4-gguf",
750
+ "revision": "1cd7b564460821938add0475a60b942c409295e0",
751
+ "dspark_filename": "DeepSeek-V4-Flash-DSpark-support.gguf",
752
+ "weights_used": false
753
+ }
754
+ }
755
+ },
756
+ "calibration": {
757
+ "repo_id": "ox-ox/DeepSeek-V4-Flash-0731-GGUF",
758
+ "filename": "imatrix/DeepSeek-V4-Flash-0731-chat-v2-routed-moe-ds4-1p5m.dat",
759
+ "revision": "6d58a3a36030c3ccb969bb5759fc6ae08cd299f8",
760
+ "bytes": 450892654,
761
+ "sha256": "6fce7674df701de544e5d3351aab04e67602eddeafeb48cf70e77ebe47239eb4",
762
+ "coverage_validation": {
763
+ "entries": 129,
764
+ "scope": "all routed gate/up/down tensors",
765
+ "policy": "exact-name-and-vector-length",
766
+ "global_ds4_imatrix_strict": true
767
+ }
768
+ },
769
+ "template": {
770
+ "repo_id": "antirez/deepseek-v4-gguf",
771
+ "filename": "DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf",
772
+ "revision": "1cd7b564460821938add0475a60b942c409295e0",
773
+ "role": "metadata-tokenizer-tensor-order-and-shapes-only",
774
+ "weights_used": false,
775
+ "remote_bytes": 86720111488,
776
+ "lfs_sha256": "ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0",
777
+ "xet_hash": "7da16e1025c1b856490c29c341f4e467d15cb195389c70383dade5e6108799ac",
778
+ "downloaded_header_bytes": 67108864,
779
+ "header_sha256": "f0e1d5e8f3b008402aa6eb32cada3873dd926c8bf5e7d00d7788eec65f09dd6d",
780
+ "resolved": {
781
+ "repo_id": "antirez/deepseek-v4-gguf",
782
+ "filename": "DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf",
783
+ "role": "metadata-tokenizer-tensor-order-and-shapes-only",
784
+ "weights_used": false,
785
+ "remote_bytes": 86720111488,
786
+ "lfs_sha256": "ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0",
787
+ "xet_hash": "7da16e1025c1b856490c29c341f4e467d15cb195389c70383dade5e6108799ac",
788
+ "resolved_etag": "7da16e1025c1b856490c29c341f4e467d15cb195389c70383dade5e6108799ac",
789
+ "range": "bytes=0-67108863",
790
+ "header_bytes": 67108864,
791
+ "header_sha256": "f0e1d5e8f3b008402aa6eb32cada3873dd926c8bf5e7d00d7788eec65f09dd6d"
792
+ }
793
+ },
794
+ "dspark_calibration_alias": {
795
+ "schema": "deepseek-model-tools.dspark-imatrix-alias.v1",
796
+ "function": "extend_dspark_legacy_imatrix",
797
+ "source": {
798
+ "repo_id": "ox-ox/DeepSeek-V4-Flash-0731-GGUF",
799
+ "revision": "6d58a3a36030c3ccb969bb5759fc6ae08cd299f8",
800
+ "filename": "imatrix/DeepSeek-V4-Flash-0731-chat-v2-routed-moe-ds4-1p5m.dat",
801
+ "bytes": 450892654,
802
+ "sha256": "6fce7674df701de544e5d3351aab04e67602eddeafeb48cf70e77ebe47239eb4"
803
+ },
804
+ "output_sha256": "689b446ed2e2657ebcb69a6516781ee0444e07641aeb6125233fb4e6fe7cbce3",
805
+ "mapping": [
806
+ {
807
+ "stage": 0,
808
+ "target_layer": 40,
809
+ "alias": "mtp.0.ffn_gate_exps.weight",
810
+ "source": "blk.40.ffn_gate_exps.weight"
811
+ },
812
+ {
813
+ "stage": 0,
814
+ "target_layer": 40,
815
+ "alias": "mtp.0.ffn_up_exps.weight",
816
+ "source": "blk.40.ffn_up_exps.weight"
817
+ },
818
+ {
819
+ "stage": 0,
820
+ "target_layer": 40,
821
+ "alias": "mtp.0.ffn_down_exps.weight",
822
+ "source": "blk.40.ffn_down_exps.weight"
823
+ },
824
+ {
825
+ "stage": 1,
826
+ "target_layer": 41,
827
+ "alias": "mtp.1.ffn_gate_exps.weight",
828
+ "source": "blk.41.ffn_gate_exps.weight"
829
+ },
830
+ {
831
+ "stage": 1,
832
+ "target_layer": 41,
833
+ "alias": "mtp.1.ffn_up_exps.weight",
834
+ "source": "blk.41.ffn_up_exps.weight"
835
+ },
836
+ {
837
+ "stage": 1,
838
+ "target_layer": 41,
839
+ "alias": "mtp.1.ffn_down_exps.weight",
840
+ "source": "blk.41.ffn_down_exps.weight"
841
+ },
842
+ {
843
+ "stage": 2,
844
+ "target_layer": 42,
845
+ "alias": "mtp.2.ffn_gate_exps.weight",
846
+ "source": "blk.42.ffn_gate_exps.weight"
847
+ },
848
+ {
849
+ "stage": 2,
850
+ "target_layer": 42,
851
+ "alias": "mtp.2.ffn_up_exps.weight",
852
+ "source": "blk.42.ffn_up_exps.weight"
853
+ },
854
+ {
855
+ "stage": 2,
856
+ "target_layer": 42,
857
+ "alias": "mtp.2.ffn_down_exps.weight",
858
+ "source": "blk.42.ffn_down_exps.weight"
859
+ }
860
+ ],
861
+ "tool": {
862
+ "repository_revision": "d97ba6b63ed857a5456f92a5190381fd7e45ca9f",
863
+ "importer_sha256": "b39987ce5ce35564bb4053ec11265c0ff41be233595ae17a6e039318a52dd4da"
864
+ },
865
+ "coverage_validation": {
866
+ "entries": 138,
867
+ "trunk_entries": 129,
868
+ "dspark_entries": 9,
869
+ "coverage": "exact-name-and-vector-length"
870
+ }
871
+ },
872
+ "tools": {
873
+ "ds4_revision": "d516d4eeb82c454aeb2831af1b1961801d6b571b",
874
+ "ds4_root": "antirez/ds4@ds4f-mxfp4",
875
+ "ds4_quantizer_sha256": "f0a381f4ada808ea2afa740d964354fa327fc1235ba7cebf50874eb89fb97ac5",
876
+ "ds4_runtime_sha256": "2aaf20469b9918d6d6ab8787a02811c11228547cd979787879a03dba8a9e7824",
877
+ "omlx_revision": "76352ed2363e42b2146243463875756e683a640d",
878
+ "deepseek_model_tools_revision": "e24746463a0a3e79036dd9c0472deac6ce704f08",
879
+ "deepseek_model_tools_worktree": {
880
+ "dirty": true,
881
+ "status_sha256": "59efc8a4b51d3dc33f267055db94edc9ac8e2392aa9f6e7033b3842a060b4ad2",
882
+ "tracked_patch_sha256": "658a58301aad033c5027f62ee1f1719ab0b01141c4b06b2d296f5f05c8f76aea"
883
+ },
884
+ "build_driver_sha256": "e9048f7535fe360756ce4505edae8f5d10f901e3fcc7759a37abbfbfc3949b4e"
885
+ },
886
+ "commands": [
887
+ [
888
+ "deepseek4-quantize",
889
+ "--hf",
890
+ "apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8",
891
+ "--template",
892
+ "antirez/deepseek-v4-gguf/ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0.header-only.sparse.gguf",
893
+ "--experts",
894
+ "iq2_xxs",
895
+ "--routed-w2",
896
+ "q2_k",
897
+ "--attention-proj",
898
+ "q8_0",
899
+ "--shared",
900
+ "q8_0",
901
+ "--threads",
902
+ "8",
903
+ "--imatrix",
904
+ "ox-ox/DeepSeek-V4-Flash-0731-GGUF/imatrix/DeepSeek-V4-Flash-0731-chat-v2-routed-moe-ds4-1p5m.dat",
905
+ "--tensor-type",
906
+ "output.weight=q8_0",
907
+ "--tensor-type",
908
+ "blk.2.indexer.attn_q_b.weight=f16",
909
+ "--tensor-type",
910
+ "blk.4.indexer.attn_q_b.weight=f16",
911
+ "--tensor-type",
912
+ "blk.6.indexer.attn_q_b.weight=f16",
913
+ "--tensor-type",
914
+ "blk.8.indexer.attn_q_b.weight=f16",
915
+ "--tensor-type",
916
+ "blk.10.indexer.attn_q_b.weight=f16",
917
+ "--tensor-type",
918
+ "blk.12.indexer.attn_q_b.weight=f16",
919
+ "--tensor-type",
920
+ "blk.14.indexer.attn_q_b.weight=f16",
921
+ "--tensor-type",
922
+ "blk.16.indexer.attn_q_b.weight=f16",
923
+ "--tensor-type",
924
+ "blk.18.indexer.attn_q_b.weight=f16",
925
+ "--tensor-type",
926
+ "blk.20.indexer.attn_q_b.weight=f16",
927
+ "--tensor-type",
928
+ "blk.22.indexer.attn_q_b.weight=f16",
929
+ "--tensor-type",
930
+ "blk.24.indexer.attn_q_b.weight=f16",
931
+ "--tensor-type",
932
+ "blk.26.indexer.attn_q_b.weight=f16",
933
+ "--tensor-type",
934
+ "blk.28.indexer.attn_q_b.weight=f16",
935
+ "--tensor-type",
936
+ "blk.30.indexer.attn_q_b.weight=f16",
937
+ "--tensor-type",
938
+ "blk.32.indexer.attn_q_b.weight=f16",
939
+ "--tensor-type",
940
+ "blk.34.indexer.attn_q_b.weight=f16",
941
+ "--tensor-type",
942
+ "blk.36.indexer.attn_q_b.weight=f16",
943
+ "--tensor-type",
944
+ "blk.38.indexer.attn_q_b.weight=f16",
945
+ "--tensor-type",
946
+ "blk.40.indexer.attn_q_b.weight=f16",
947
+ "--tensor-type",
948
+ "blk.42.indexer.attn_q_b.weight=f16",
949
+ "--tensor-type",
950
+ "blk.10.ffn_gate_exps.weight=mxfp4",
951
+ "--tensor-type",
952
+ "blk.10.ffn_up_exps.weight=mxfp4",
953
+ "--tensor-type",
954
+ "blk.10.ffn_down_exps.weight=mxfp4",
955
+ "--tensor-type",
956
+ "blk.14.ffn_gate_exps.weight=mxfp4",
957
+ "--tensor-type",
958
+ "blk.14.ffn_up_exps.weight=mxfp4",
959
+ "--tensor-type",
960
+ "blk.14.ffn_down_exps.weight=mxfp4",
961
+ "--tensor-type",
962
+ "blk.30.ffn_gate_exps.weight=mxfp4",
963
+ "--tensor-type",
964
+ "blk.30.ffn_up_exps.weight=mxfp4",
965
+ "--tensor-type",
966
+ "blk.30.ffn_down_exps.weight=mxfp4",
967
+ "--tensor-type",
968
+ "blk.34.ffn_gate_exps.weight=mxfp4",
969
+ "--tensor-type",
970
+ "blk.34.ffn_up_exps.weight=mxfp4",
971
+ "--tensor-type",
972
+ "blk.34.ffn_down_exps.weight=mxfp4",
973
+ "--tensor-type",
974
+ "blk.37.ffn_gate_exps.weight=mxfp4",
975
+ "--tensor-type",
976
+ "blk.37.ffn_up_exps.weight=mxfp4",
977
+ "--tensor-type",
978
+ "blk.37.ffn_down_exps.weight=mxfp4",
979
+ "--tensor-type",
980
+ "blk.38.ffn_gate_exps.weight=mxfp4",
981
+ "--tensor-type",
982
+ "blk.38.ffn_up_exps.weight=mxfp4",
983
+ "--tensor-type",
984
+ "blk.38.ffn_down_exps.weight=mxfp4",
985
+ "--tensor-type",
986
+ "blk.39.ffn_gate_exps.weight=mxfp4",
987
+ "--tensor-type",
988
+ "blk.39.ffn_up_exps.weight=mxfp4",
989
+ "--tensor-type",
990
+ "blk.39.ffn_down_exps.weight=mxfp4",
991
+ "--tensor-type",
992
+ "blk.40.ffn_gate_exps.weight=mxfp4",
993
+ "--tensor-type",
994
+ "blk.40.ffn_up_exps.weight=mxfp4",
995
+ "--tensor-type",
996
+ "blk.40.ffn_down_exps.weight=mxfp4",
997
+ "--tensor-type",
998
+ "blk.41.ffn_gate_exps.weight=mxfp4",
999
+ "--tensor-type",
1000
+ "blk.41.ffn_up_exps.weight=mxfp4",
1001
+ "--tensor-type",
1002
+ "blk.41.ffn_down_exps.weight=mxfp4",
1003
+ "--tensor-type",
1004
+ "blk.42.ffn_gate_exps.weight=mxfp4",
1005
+ "--tensor-type",
1006
+ "blk.42.ffn_up_exps.weight=mxfp4",
1007
+ "--tensor-type",
1008
+ "blk.42.ffn_down_exps.weight=mxfp4",
1009
+ "--imatrix-strict",
1010
+ "--out",
1011
+ "DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf"
1012
+ ],
1013
+ [
1014
+ "deepseek4-quantize",
1015
+ "--hf",
1016
+ "apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8",
1017
+ "--template",
1018
+ "antirez/deepseek-v4-gguf/ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0.header-only.sparse.gguf",
1019
+ "--dspark-support",
1020
+ "--experts",
1021
+ "iq2_xxs",
1022
+ "--routed-w2",
1023
+ "mxfp4",
1024
+ "--attention-proj",
1025
+ "q8_0",
1026
+ "--shared",
1027
+ "q8_0",
1028
+ "--dspark-block-size",
1029
+ "5",
1030
+ "--dspark-markov-rank",
1031
+ "256",
1032
+ "--dspark-noise-token-id",
1033
+ "128799",
1034
+ "--dspark-target-layers",
1035
+ "40,41,42",
1036
+ "--threads",
1037
+ "8",
1038
+ "--imatrix",
1039
+ "ox-ox/DeepSeek-V4-Flash-0731-GGUF/imatrix/DeepSeek-V4-Flash-0731-dspark-alias.dat",
1040
+ "--out",
1041
+ "DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf",
1042
+ "--imatrix-strict"
1043
+ ]
1044
+ ],
1045
+ "appledouble_excluded": true
1046
+ }
BUILD_PLAN.md ADDED
@@ -0,0 +1,73 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Build and verification record
2
+
3
+ ## Objective
4
+
5
+ Create a reproducible DS4 package that maximizes target-model quality within a
6
+ 128 GB unified-memory envelope. Preserve exact native MXFP4 routed experts on
7
+ the ten sensitivity-selected layers and retain the proven low-bit/Q8 policy
8
+ everywhere else.
9
+
10
+ ## Inputs and toolchain
11
+
12
+ - All weights were regenerated from
13
+ `apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8`.
14
+ - The sparse reference GGUF supplied metadata, tokenizer, tensor order and
15
+ shapes only.
16
+ - All 48 source-shard hashes, source metadata and the abliteration manifest
17
+ were verified before conversion.
18
+ - The routed-expert imatrix supplied exact coverage for the 129 tensors whose
19
+ output types consume calibration data.
20
+ - DS4 source revision: `d516d4eeb82c454aeb2831af1b1961801d6b571b`.
21
+ - Upstream `ds4f-mxfp4` base: `4893e0c40fba03dbc85555faeb035799aa04e0b6`.
22
+ - Quantizer SHA-256: `f0a381f4ada808ea2afa740d964354fa327fc1235ba7cebf50874eb89fb97ac5`.
23
+ - Runtime SHA-256: `2aaf20469b9918d6d6ab8787a02811c11228547cd979787879a03dba8a9e7824`.
24
+
25
+ ## Quantization policy
26
+
27
+ - Native MXFP4 gate/up/down routed experts on layers
28
+ `10, 14, 30, 34, 37, 38, 39, 40, 41, 42`.
29
+ - IQ2_XXS gate/up and Q2_K down routed experts on the other 33 MoE layers.
30
+ - Q8 attention, shared-expert and output tensors.
31
+ - Protected F16 indexer and auxiliary tensors.
32
+ - DSpark support with IQ2_XXS gate/up and native MXFP4 down projections for
33
+ target layers 40, 41 and 42.
34
+
35
+ ## Conversion gates
36
+
37
+ - Strict imatrix mode passed for main and DSpark conversions.
38
+ - Main GGUF contains exactly 1,328 tensors and the intended 30 MXFP4 tensors.
39
+ - DSpark support contains exactly 81 tensors and three MXFP4 down aggregates.
40
+ - Every native tensor comparison required an explicit `byte_compare: OK`.
41
+ - All 30 main and three DSpark MXFP4 tensors independently reproduced the
42
+ source FP4 codes and scale bytes.
43
+ - `ds4 --cpu --inspect --dspark-strict` reported zero missing tensors, invalid
44
+ bindings or metadata errors.
45
+ - The final GGUF hashes are recorded in `SHA256SUMS`.
46
+
47
+ ## Result
48
+
49
+ | Component | Bytes | GiB |
50
+ | --- | ---: | ---: |
51
+ | Main GGUF | `102,826,238,912` | `95.7644` |
52
+ | DSpark support | `7,297,737,120` | `6.7965` |
53
+ | Combined | `110,123,976,032` | `102.5609` |
54
+
55
+ Native MXFP4 improves fidelity over a second Q4_K requantization and uses 4.25
56
+ bits per weight rather than 4.50. Native MXFP4 Metal and the mixed DSpark path
57
+ are newer than mature Q4_K kernels, so throughput should be measured on the
58
+ target system.
59
+
60
+ ## One-million-token context
61
+
62
+ Recommended target-only mode:
63
+
64
+ ```sh
65
+ ds4 --metal \
66
+ -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
67
+ --ctx 1048576 --prefill-chunk 2048
68
+ ```
69
+
70
+ Estimated residency is at most `110.30 GiB`. DSpark can be tested with a 1,024
71
+ token prefill chunk; estimated residency is `114.02 GiB`. A 4,096-token chunk
72
+ with DSpark is estimated at `123.24 GiB` and is not a reliably resident mode on
73
+ a machine whose Metal recommended working set is approximately `121.60 GiB`.
DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cd8593a232c9feebc4c91855d5ab486b17250fc8bc2f294bc80401f93b371566
3
+ size 7297737120
DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2cfc36b761b59ea43531e7cdb02a690436a330e42ad57cb162726b385914df59
3
+ size 102826238912
PROVENANCE.md ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Provenance
2
+
3
+ This artifact was regenerated from the abliterated FP8 checkpoint only. No
4
+ model tensors were copied from a public GGUF or another quantized model.
5
+
6
+ - Weight source: `apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8`
7
+ - Upstream base model: `deepseek-ai/DeepSeek-V4-Flash-0731`
8
+ - Upstream revision: `9e165c30e2704aec5d9d593cce3eebd58bbef1cb`
9
+ - Abliteration manifest SHA-256: `8e6f40f3d92542720aabdafb728f90e3397fba68c6a8d7ba91772b5dd41bc767`
10
+ - Calibration: `ox-ox/DeepSeek-V4-Flash-0731-GGUF/imatrix/DeepSeek-V4-Flash-0731-chat-v2-routed-moe-ds4-1p5m.dat`
11
+ - Calibration SHA-256: `6fce7674df701de544e5d3351aab04e67602eddeafeb48cf70e77ebe47239eb4`
12
+ - DSpark calibration proxy SHA-256: `689b446ed2e2657ebcb69a6516781ee0444e07641aeb6125233fb4e6fe7cbce3`
13
+ - Metadata template: `antirez/deepseek-v4-gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf`
14
+ - Template LFS SHA-256: `ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0`
15
+ - Template Xet object: `7da16e1025c1b856490c29c341f4e467d15cb195389c70383dade5e6108799ac`
16
+ - Template 64 MiB header SHA-256: `f0e1d5e8f3b008402aa6eb32cada3873dd926c8bf5e7d00d7788eec65f09dd6d`
17
+ - Engine: `ds4`
18
+ - DSpark support: yes
19
+
20
+ The source abliteration modified 36 attention `wo_b` tensors. It did not modify
21
+ routed-expert codes or scales, allowing the selected MXFP4 experts to be
22
+ repacked losslessly from their original I8 codes and F8_E8M0 scales.
README.md ADDED
@@ -0,0 +1,128 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # DeepSeek V4 Flash Abliterated — DS4 Quality128
2
+
3
+ > **Exact MXFP4 experts, maximum resident quality.**
4
+
5
+ > **Runtime requirement:** until native MXFP4 support is merged into DS4's
6
+ > default branch, use the
7
+ > official [`antirez/ds4` `ds4f-mxfp4` branch](https://github.com/antirez/ds4/tree/ds4f-mxfp4).
8
+
9
+ This is a quality-first DS4 package designed to keep DeepSeek V4 Flash resident
10
+ on a 128 GB M1 Ultra while preserving the most sensitive routed experts in
11
+ their exact native MXFP4 representation.
12
+
13
+ All weights were regenerated from the abliterated FP8 checkpoint
14
+ [`apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8`](https://huggingface.co/apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8).
15
+
16
+ No model tensor was copied from another GGUF. The source abliteration edited
17
+ only 36 attention `wo_b` tensors; routed-expert codes and scales were untouched.
18
+
19
+ ## Artifacts
20
+
21
+ | File | Bytes | GiB | SHA-256 |
22
+ | --- | ---: | ---: | --- |
23
+ | `DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf` | `102,826,238,912` | `95.7644` | `2cfc36b761b59ea43531e7cdb02a690436a330e42ad57cb162726b385914df59` |
24
+ | `DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf` | `7,297,737,120` | `6.7965` | `cd8593a232c9feebc4c91855d5ab486b17250fc8bc2f294bc80401f93b371566` |
25
+ | **Combined** | **`110,123,976,032`** | **`102.5609`** | — |
26
+
27
+ ## Quantization profile
28
+
29
+ - Exact native MXFP4 gate/up/down routed experts on layers
30
+ `10, 14, 30, 34, 37, 38, 39, 40, 41, 42`.
31
+ - IQ2_XXS gate/up and Q2_K down routed experts on the other 33 MoE layers.
32
+ - Q8 attention, shared-expert and output paths.
33
+ - F16 protected indexer and auxiliary tensors.
34
+ - The routed-expert imatrix applies only to genuinely requantized IQ2/Q2
35
+ tensors; preserved MXFP4 needs no imatrix.
36
+ - DSpark support uses IQ2_XXS gate/up and exact native MXFP4 down projections
37
+ for target layers 40, 41 and 42.
38
+
39
+ Main GGUF type histogram (1,328 tensors): F32 492, F16 359, I32 3, Q8_0 345,
40
+ IQ2_XXS 66, Q2_K 33 and MXFP4 30. DSpark support histogram (81 tensors): F32
41
+ 34, F16 7, Q8_0 31, IQ2_XXS 6 and MXFP4 3.
42
+
43
+ ## Runtime requirement
44
+
45
+ Native MXFP4 and this DSpark layout require the official
46
+ [`antirez/ds4` `ds4f-mxfp4` branch](https://github.com/antirez/ds4/tree/ds4f-mxfp4),
47
+ or DS4's default branch after those changes are merged. The artifact was
48
+ qualified against `ds4f-mxfp4` revision
49
+ `4893e0c40fba03dbc85555faeb035799aa04e0b6` with the DSpark generation fix
50
+ tracked in [antirez/ds4#642](https://github.com/antirez/ds4/issues/642).
51
+
52
+ - Quantizer SHA-256: `f0a381f4ada808ea2afa740d964354fa327fc1235ba7cebf50874eb89fb97ac5`
53
+ - Runtime SHA-256: `2aaf20469b9918d6d6ab8787a02811c11228547cd979787879a03dba8a9e7824`
54
+
55
+ Do not assume an older DS4 executable can load or execute this package
56
+ correctly merely because it accepts the GGUF container.
57
+
58
+ ## Launch examples
59
+
60
+ Normal target-only inference:
61
+
62
+ ```sh
63
+ ds4 --metal \
64
+ -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf
65
+ ```
66
+
67
+ Greedy DSpark inference:
68
+
69
+ ```sh
70
+ ds4 --metal \
71
+ -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
72
+ --mtp DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf \
73
+ --dspark --temp 0
74
+ ```
75
+
76
+ ## One-million-token context
77
+
78
+ `--ctx 1048576` counts prompt and completion together. The safest maximum-quality
79
+ resident mode omits DSpark and uses a 2,048-token prefill chunk:
80
+
81
+ ```sh
82
+ ds4 --metal \
83
+ -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
84
+ --ctx 1048576 --prefill-chunk 2048
85
+ ```
86
+
87
+ Estimated residency is at most `110.30 GiB`, leaving at least `11.30 GiB`
88
+ below Metal's approximately `121.60 GiB` recommended working set. Omitting
89
+ DSpark does not reduce target-model quality; it only forgoes speculative decode.
90
+
91
+ After measuring peak memory on this machine, DSpark can be enabled with the
92
+ smaller chunk:
93
+
94
+ ```sh
95
+ ds4 --metal \
96
+ -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
97
+ --ctx 1048576 --prefill-chunk 1024 \
98
+ --mtp DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf \
99
+ --dspark --temp 0
100
+ ```
101
+
102
+ That mode is estimated at `114.02 GiB`, about `7.58 GiB` below the recommended
103
+ working set. DSpark with chunk 4096 is estimated at `123.24 GiB` and is not a
104
+ reliably resident configuration.
105
+
106
+ ## Verification and provenance
107
+
108
+ The build passed strict main and DSpark planning, exact size/type/name-set
109
+ checks, strict imatrix coverage, source validation, and byte reproduction for
110
+ all 30 main plus three DSpark MXFP4 tensors. `SHA256SUMS` binds the final GGUFs
111
+ and documentation. See:
112
+
113
+ - [`BUILD_MANIFEST.json`](BUILD_MANIFEST.json) for pinned source shards, tools,
114
+ commands and publication metadata.
115
+ - [`PROVENANCE.md`](PROVENANCE.md) for the compact lineage record.
116
+ - [`BUILD_PLAN.md`](BUILD_PLAN.md) for the completed build gates and benchmark
117
+ matrix.
118
+ - [`ds4-upstream-issues.md`](ds4-upstream-issues.md) for remaining DS4 runtime
119
+ and memory improvements.
120
+
121
+ The sibling MLX package
122
+ `DeepSeek-V4-Flash-0731-Abliterated-MLX-Quality128-suboptimal` is intentionally
123
+ retained for speed comparison; it is not the canonical quality artifact.
124
+
125
+ Native MXFP4 Metal kernels and the mixed DSpark path are newer than the mature
126
+ Q4_K path. Correctness and throughput should therefore be benchmarked against
127
+ the archived DS4 v1 and retained MLX comparator before changing everyday launch
128
+ defaults.
SHA256SUMS ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ 27e3fbe612341ffe8f0060014831d88b51b2f1ca986f820d98296e4e6b7c85a7 BUILD_MANIFEST.json
2
+ b52972dd4cf914a8abe0c224b39ac5975c808450595e87945bc7a00c37eac649 BUILD_PLAN.md
3
+ cd8593a232c9feebc4c91855d5ab486b17250fc8bc2f294bc80401f93b371566 DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf
4
+ 2cfc36b761b59ea43531e7cdb02a690436a330e42ad57cb162726b385914df59 DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf
5
+ ea2f1f2bdb10266e509dbe376a5f5b51eba6433ebecc43c8a0d1b3c87d2aaa02 PROVENANCE.md
6
+ 6b8c9bc266d71ac941ded6bd3750d7716f4158507446b8ac42a517b9872a9049 README.md
7
+ afd8c0883b69a9280ccbd6e191d6f9abfe757498e1cfbd1c9f949edf4ef1a95b ds4-upstream-issues.md
ds4-upstream-issues.md ADDED
@@ -0,0 +1,163 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # DS4 upstream issues and improvement candidates
2
+
3
+ Audit base: the official
4
+ [`antirez/ds4` `ds4f-mxfp4` branch](https://github.com/antirez/ds4/tree/ds4f-mxfp4)
5
+ at revision `4893e0c40fba03dbc85555faeb035799aa04e0b6`.
6
+
7
+ This is an engineering backlog, not a claim that issues have already been filed upstream. Items distinguish blockers for the planned artifact from optional improvements to DS4 itself.
8
+
9
+ ## P0 — required for the Quality128 package
10
+
11
+ ### 1. DSpark planner rejects preserved MXFP4 routed tensors
12
+
13
+ **Status:** fixed on the local PR branch and tracked upstream as
14
+ [antirez/ds4#642](https://github.com/antirez/ds4/issues/642). The real strict
15
+ support dry-run now reports 81 tensors, three MXFP4 down aggregates, and
16
+ `7,297,737,120` bytes.
17
+
18
+ **Observed:**
19
+
20
+ ```text
21
+ --dspark-support --routed-w2 mxfp4 --dry-run
22
+ error: unsupported DSpark planned tensor type
23
+ ```
24
+
25
+ The main planner recognizes a narrowly scoped preserved-MXFP4 expert exception around `gguf-tools/deepseek4-quantize.c:1734`. The DSpark sizing path at approximately `gguf-tools/deepseek4-quantize.c:2232` accepts I32 or ordinary quantizable targets but does not mirror that exception.
26
+
27
+ The DSpark generator already routes `mtp.*` expert tensors through the common native repacker. This appears to be a planning omission, not a missing conversion implementation.
28
+
29
+ **Proposed change:** permit MXFP4 only when the DSpark plan entry is a packed routed expert with the required block alignment and source dtype/shape contract. Do not make MXFP4 a generic quantizable type.
30
+
31
+ **Acceptance:**
32
+
33
+ - Strict DSpark dry-run succeeds with `routed_w2=mxfp4`.
34
+ - Exactly three DSpark down aggregates become MXFP4.
35
+ - Expected support size is `7,297,737,120` bytes unless reviewed metadata changes explain otherwise.
36
+ - Non-expert or incompatible source tensors are still rejected.
37
+
38
+ ### 2. Missing mixed IQ2/IQ2/MXFP4 DSpark Metal coverage
39
+
40
+ **Status:** covered on the local PR branch by the hardened mixed-path Metal
41
+ fixture for token counts one through five and the native-C quantizer regression
42
+ fixture. Full support-model runtime smoke remains an artifact acceptance gate.
43
+
44
+ Current MXFP4 Metal coverage tests all-MXFP4 gate/up/down at limited batch sizes. The planned sidecar uses IQ2_XXS gate/up and MXFP4 down with DSpark proposal blocks of one through five tokens.
45
+
46
+ **Proposed test matrix:**
47
+
48
+ - Token counts `1, 2, 3, 4, 5`.
49
+ - IQ2_XXS fused gate/up plus MXFP4 down.
50
+ - Fused sum path for one through four tokens.
51
+ - Generic MV plus expert-summation path for five tokens.
52
+ - CPU/scalar reference comparison with explicit tolerances.
53
+ - At least one full DSpark verifier smoke test using the real support binding.
54
+
55
+ **Acceptance:** deterministic pass on M1 Ultra Metal with no out-of-bounds access and bounded numerical error versus the reference.
56
+
57
+ ### 3. Byte-compare failure should fail the process
58
+
59
+ **Status:** still an upstream hardening opportunity. The Quality128 build
60
+ wrapper closes the production risk by requiring the literal
61
+ `byte_compare: OK` for every one of the 30 main and three DSpark native tensors.
62
+
63
+ `gguf-tools/deepseek4-quantize.c` prints `byte_compare: FAIL` but the comparison path can still exit successfully. A production wrapper can parse for an explicit `byte_compare: OK`, but the quantizer should provide reliable process semantics itself.
64
+
65
+ **Proposed change:** return nonzero when any byte mismatch is found or when a requested comparison cannot be completed.
66
+
67
+ **Acceptance:** a deliberately corrupted reference tensor produces a nonzero exit; an exact native repack produces zero and `byte_compare: OK`.
68
+
69
+ ## P1 — high-value Metal/runtime improvements
70
+
71
+ ### 4. Right-size the one-million-context `comp_mask` allocation
72
+
73
+ At `ctx=1,048,576` and prefill chunk `4096`, both `indexer_scores_by_tier` and `comp_mask_by_tier` are allocated as `comp_cap * prefill_cap * sizeof(float)` near `ds4.c:17158`.
74
+
75
+ Each allocation is roughly `4 GiB`. The normal indexed batch path consumes compact top-k selected indices, while `comp_mask` appears to be used principally by one-row/fallback attention and small top-value scratch paths. If confirmed across all backends, allocating a full context-by-prefill mask is unnecessary.
76
+
77
+ **Potential improvement:** size `comp_mask` to its maximum actual live row count, or give the large batch path a separate compact scratch allocation.
78
+
79
+ **Expected benefit:** reclaim close to `4 GiB` at 1M/chunk4096, potentially moving the full resident Quality128 package plus DSpark from approximately `123.24` to `119.24 GiB` without changing weights.
80
+
81
+ **Required proof:** audit every Metal/CUDA use, add allocation-bound assertions, run indexed-prefill/decode consistency tests, and demonstrate bit-identical or tolerance-identical logits before and after the change.
82
+
83
+ ### 5. Tile indexed-attention scoring independently of the MoE prefill chunk
84
+
85
+ `indexer_scores` genuinely scales with compressed-cache rows times the global prefill batch. Reducing `--prefill-chunk` from 4096 to 1024 cuts total planned memory from about `123.24` to `114.02 GiB`, but also makes every prefill stage use smaller batches.
86
+
87
+ **Potential improvement:** keep a larger MoE/dense prefill batch while processing indexer score/top-k work in bounded subtiles. This separates the attention-scratch memory decision from the MoE throughput decision.
88
+
89
+ **Acceptance:** materially lower peak memory at 1M context, no retrieval/logit regression, and better prefill throughput than globally reducing the chunk to the same scratch footprint.
90
+
91
+ ### 6. Phase-separate prefill workspace and DSpark residency
92
+
93
+ DSpark does not accelerate prefill, yet its approximately `6.80 GiB` support model and verifier state coexist with the largest prefill workspace. The prefill workspace is also retained after prefill for possible later session work.
94
+
95
+ **Potential improvement:**
96
+
97
+ 1. Prefill before mapping/requesting residency for the DSpark support model.
98
+ 2. Release or shrink prefill-only tensors when entering decode.
99
+ 3. Map DSpark and allocate verifier state for speculative decode.
100
+ 4. Reverse the transition safely if a later prompt extension needs batched prefill.
101
+
102
+ This is a lifecycle/allocator change, not a model-quality trade-off.
103
+
104
+ **Acceptance:** chunk4096 prefill and DSpark decode both remain under the Metal working-set limit in their respective phases; session extension, cancellation and cleanup tests show no leaks or stale views.
105
+
106
+ ### 7. Startup memory reporting omits important persistent workspace
107
+
108
+ The high-level planned-memory log accounts for model span, KV and its `scratch_bytes` estimate, but does not make all batch-prefill tensors, DSpark verifier snapshots, logits/host buffers and driver transients obvious. At 1M this can understate the actionable footprint by multiple GiB.
109
+
110
+ **Potential improvement:** have the session allocator report:
111
+
112
+ - Model and support-model mapped/resident bytes.
113
+ - Raw and compressed KV.
114
+ - Indexed-attention scratch.
115
+ - Other prefill workspace.
116
+ - Decode and DSpark verifier workspace.
117
+ - Runtime tensor live/peak totals.
118
+ - Metal recommended working set and remaining margin.
119
+
120
+ The same complete calculation should gate allocation before a predictable out-of-memory failure.
121
+
122
+ ### 8. SSD streaming is incompatible with `--mtp`, and mixed-size caching is weak
123
+
124
+ Current DS4 rejects `--ssd-streaming` with `--mtp` around `ds4.c:56160`. Separately, the streaming cache uses the dominant routed-expert size class. In this hybrid, the 33 IQ2/Q2 layers define that slab while the ten larger MXFP4 layers bypass it through mapped-model views.
125
+
126
+ **Potential improvements:**
127
+
128
+ - Make support-model residency compatible with main-model expert streaming.
129
+ - Account for DSpark and long-context buffers when choosing an automatic cache budget.
130
+ - Support two expert-cache size classes or a byte-addressed allocator so native-MXFP4 layers can participate.
131
+
132
+ **Acceptance:** explicit memory caps are respected, native expert bytes remain exact, and cache-miss/decode benchmarks show predictable behavior on an external SSD.
133
+
134
+ ## P2 — build and release hardening
135
+
136
+ ### 9. Include the native Metal MXFP4 test in the default relevant test suite
137
+
138
+ The scalar MXFP4 test is covered by the general test target, while `test-mxfp4-metal` is separate. A source revision can therefore pass routine tests without exercising the new native Metal kernel.
139
+
140
+ **Proposed change:** add a documented aggregate target suitable for Metal release qualification, or include the test when Metal is available.
141
+
142
+ ### 10. Make runtime Metal-source provenance explicit
143
+
144
+ DS4 compiles/loads Metal sources dynamically and can depend on the working directory, `--chdir`, or `DS4_METAL_MOE_SOURCE`. A current executable paired with stale or missing Metal source is a reproducibility risk.
145
+
146
+ **Potential improvement:** embed a Metal-source revision/hash in the executable and print the resolved source path plus hash at startup. Release verification should fail on a mismatch when native MXFP4 is required.
147
+
148
+ ### 11. Expose machine-readable plan and inspection output
149
+
150
+ Dry-run and inspection text is currently parsed by external tooling. A stable JSON mode would make it safer to assert tensor counts, types, byte sizes, imatrix coverage, support binding and memory planning.
151
+
152
+ **Acceptance:** JSON schema/version is explicit and text output remains available for humans.
153
+
154
+ ## Suggested upstream submission order
155
+
156
+ 1. DSpark preserved-MXFP4 planner fix plus mixed-path test.
157
+ 2. Byte-compare exit semantics.
158
+ 3. `comp_mask` allocation audit and minimal right-sizing patch.
159
+ 4. Complete memory-reporting/accounting patch.
160
+ 5. Tiled indexer scratch and phase-separated residency as separately benchmarked changes.
161
+ 6. Streaming/dual-size-cache work only after resident Quality128 is stable.
162
+
163
+ Each submission should be small, independently testable and accompanied by before/after memory or throughput measurements rather than bundled into the model-build patch.