Qwen 3.8 27B Hereticโ€‘Ara IQ4_XS ้‡ๅŒ–ๆจกๅž‹๏ผˆ้€‚้… 16GB ๆ˜พๅญ˜๏ผ‰

ๆœฌๆจกๅž‹ๅŸบไบŽ Qwen 3.8 27B Hereticโ€‘Ara BF16 ่ฟ›่กŒ IQ4_XS ้‡ๅŒ–๏ผˆ4โ€‘bit๏ผ‰๏ผŒๆ–‡ไปถไฝ“็งฏไธบ 12.7-13.3 GiB๏ผŒไธ“ไธบ 16GB ๆ˜พๅญ˜็š„ๆ˜พๅกไผ˜ๅŒ–

8/22 ๆ›ดๆ–ฐไฟฎๅคๆ€่€ƒ็š„้—ฎ้ข˜๏ผŒๆจกๅž‹ๆ€ง่ƒฝๆฒกๆœ‰ๅ˜ๅŒ– ๅคงๆฆ‚่ฟ˜้œ€่ฆ1ๅคฉๆˆ‘ไผšๆ›ดๆ–ฐMTP็š„็‰ˆๆœฌ

8/22 Updated and fixed the issues with thinking; model performance remains unchanged It will probably take about another day for me to update the MTP version.

8/23 ๆ›ดๆ–ฐ๏ผŒๅฏนๆจกๅž‹ๆœฌ่บซ็š„ๆ€ง่ƒฝ่ฟ›่กŒไธ€ๅฎšไผ˜ๅŒ–๏ผŒไฝ“็งฏ็•ฅๅพฎๅ˜ๅคง๏ผŒๆ— MTP็‰ˆๆœฌ13Gib๏ผŒMTP็‰ˆๆœฌ13.3Gib 8/23 Update: The model's performance has been slightly optimized, resulting in a slight increase in size The non-MTP version is 13GiB, while the MTP version is 13.3GiB

ไธŽๅŒไฝ“็งฏ็š„ Hereticโ€‘Araโ€‘Q3_K_M๏ผˆ12.4 GiB๏ผ‰้‡ๅŒ–ๆ–นๆกˆ่ฟ›่กŒไบ†ๅ…จ้ขๅฏนๆฏ”

ๆœฌๆจกๅž‹ไฝฟ็”จHeretic Arbitrary-Rank Ablationๅšๅˆฐ็š„ๆ— ๅฎกๆŸฅ

๐Ÿ“Š ้‡ๅŒ–่ดจ้‡ๅฏนๆฏ”

่ฏ„ไผฐๆŒ‡ๆ ‡ Heretic-Ara BF16 (base) IQ4_XS-3.0 IQ4_XS-2.0 Heretic-Ara-Q3_K_M
ๆ–‡ไปถๅคงๅฐ 50.1 GiB 13 GiB (ๅธฆMTP 13.3 GiB) 12.7 GiB 12.4 GiB
้‡ๅŒ–็ฒพๅบฆ BF16 IQ4_XS (4โ€‘bit) IQ4_XS (4โ€‘bit) Q3_K_M (็บฆ 3โ€‘bit)
ๆจกๅž‹ๅ›ฐๆƒ‘ๅบฆ (Mean PPL) 7.008212 ยฑ 0.045362 7.046980 ยฑ 0.045498 7.102940 ยฑ 0.046017 7.403971 ยฑ 0.048924
ไธŽๅŸบๅบงๆจกๅž‹ PPL ็›ธๅ…ณๆ€ง 100% 99.34% 99.26% 98.31%
ๅนณๅ‡ KL ๆ•ฃๅบฆ (Mean KLD) 0 0.027832 ยฑ 0.000324 0.033398 ยฑ 0.000308 0.076034 ยฑ 0.000554
ๆœ€ๅคง KL ๆ•ฃๅบฆ (Max KLD) 0 18.317436 15.094215 17.866985
99.9% KL ๅˆ†ไฝๆ•ฐ 0 1.162850 1.130034 2.448278
Topโ€‘1 ไธ€่‡ด็އ (Same top p) 100% 92.867% ยฑ 0.067% 91.619% ยฑ 0.072% 88.152% ยฑ 0.084%
ๅนณๅ‡ๆฆ‚็އๅ˜ๅŒ– (Mean ฮ”p) 0% -0.243% ยฑ 0.012% -0.306% ยฑ 0.013% -0.490% ยฑ 0.020%
RMS ๆฆ‚็އๅ˜ๅŒ– (RMS ฮ”p) 0% 4.538% ยฑ 0.045% 4.952% ยฑ 0.041% 7.560% ยฑ 0.054%

ๆณจ๏ผšๅŸบๅบง๏ผˆBF16๏ผ‰็š„ KL ๆ•ฃๅบฆใ€ฮ”p ็ญ‰ๆŒ‡ๆ ‡ๅ‡ไธบ 0๏ผˆ่‡ช่บซๅฏนๆฏ”๏ผ‰๏ผŒไธ€่‡ด็އไธบ 100%ใ€‚

ๅœจไธๅฏ็”จ MTP ็š„ๆƒ…ๅ†ตไธ‹๏ผŒIQ4_XS ๆจกๅž‹ๅœจ 16 GiB ๆ— ๆ˜พๅญ˜ๅ ็”จ๏ผˆไธไฝœไธบ Windows ๆ˜พ็คบๆ˜พๅก๏ผ‰ไธ‹ๅฏๆ”ฏๆŒ็บฆ 110k ไธŠไธ‹ๆ–‡

ๅผ€ๅฏ MTP ๅŽ็บฆไธบ 80k

Qwen 3.8 27B Hereticโ€‘Ara IQ4_XS Quantized Model (Optimized for 16 GB VRAM) This model is quantized from Qwen 3.8 27B Hereticโ€‘Ara BF16 using the IQ4_XS scheme (4โ€‘bit), with a file size of 12.8 GiB, specifically designed for graphics cards with 16 GB of VRAM.

It has been comprehensively compared against the similarly sized Hereticโ€‘Araโ€‘Q3_K_M (12.4 GiB) quantization variant.

This model achieves uncensored behavior through Heretic's Arbitrary-Rank Ablation.

๐Ÿ“Š Quantization Quality Comparison

Evaluation Metric Heretic-Ara BF16 (base) IQ4_XS-3.0 IQ4_XS-2.0 Heretic-Ara-Q3_K_M (comparison)
File Size 50.1 GiB 13 GiB (with MTP 13.3 GiB) 12.7 GiB 12.4 GiB
Quantization Precision BF16 IQ4_XS (4โ€‘bit) IQ4_XS (4โ€‘bit) Q3_K_M (~3โ€‘bit)
Model Perplexity (Mean PPL) 7.008212 ยฑ 0.045362 7.046980 ยฑ 0.045498 7.102940 ยฑ 0.046017 7.403971 ยฑ 0.048924
PPL Correlation with Base Model 100% 99.34% 99.26% 98.31%
Mean KL Divergence (Mean KLD) 0 0.027832 ยฑ 0.000324 0.033398 ยฑ 0.000308 0.076034 ยฑ 0.000554
Max KL Divergence (Max KLD) 0 18.317436 15.094215 17.866985
99.9% KL Quantile 0 1.162850 1.130034 2.448278
Topโ€‘1 Agreement Rate (Same top p) 100% 92.867% ยฑ 0.067% 91.619% ยฑ 0.072% 88.152% ยฑ 0.084%
Mean Probability Change (Mean ฮ”p) 0% -0.243% ยฑ 0.012% -0.306% ยฑ 0.013% -0.490% ยฑ 0.020%
RMS Probability Change (RMS ฮ”p) 0% 4.538% ยฑ 0.045% 4.952% ยฑ 0.041% 7.560% ยฑ 0.054%

Note: For the base model (BF16), KL divergence, ฮ”p, etc. are all 0 (selfโ€‘comparison), and the agreement rate is 100%.

With MTP (Multiโ€‘Token Prediction) disabled, the IQ4_XS model supports approximately 110k context length on a 16 GiB GPU with no VRAM reserved for display (i.e., not used as the primary display adapter). With MTP enabled, the context length is approximately 80k.

Downloads last month
4,459
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Bucoid/Qwen3.8-27B-Heretic-Ara-16GB-VRAM-IQ4-XS-MTP-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(776)
this model