ๆญคๆจกๅž‹ๅฏ่ƒฝ็ŸญๆœŸๆ›ดๆ–ฐ๏ผŒๆˆ‘ๅ‘ๅธƒไบ†ไธ€ไธชๅŸบไบŽHeretic Arbitrary-Rank Ablation็š„ๆ— ๅฎกๆŸฅ็‰ˆๆœฌๆจกๅž‹๏ผŒๆ€ง่ƒฝๆ›ดๅฅฝไธ”ไฝ“็งฏๆ›ดๅฐ ้“พๆŽฅ๏ผš https://huggingface.co/Bucoid/Qwen3.8-27B-Heretic-Ara-IQ4-XS-16GB-VRAM-GGUF ่ฟ™ไธชๆจกๅž‹ๅฏ่ƒฝ่ฟ‡ไธ€ๆฎตๆ—ถ้—ดๆˆ‘ไผšๆ›ดๆ–ฐ่ฎฉไป–ไธ้‚ฃไนˆ่œ๏ผŒๅฆ‚ๆžœไฝ ้œ€่ฆๆ— ๅฎกๆŸฅ็‰ˆๆœฌ็š„ๆจกๅž‹๏ผŒๅŸบไบŽไธ‹่ฝฝ่ฟ™ไธชAra็š„

This model may receive short-term updates. I have released an uncensored version based on Heretic Arbitrary-Rank Ablation, which offers better performance and a smaller file size. Link: https://huggingface.co/Bucoid/Qwen3.8-27B-Heretic-Ara-IQ4-XS-16GB-VRAM-GGUF This model may be updated in a while to make it less underwhelming If you need an uncensored version, please download this Ara-based one instead.

Qwen3.8-27B Uncensored IQ4_XS ้‡ๅŒ–ๆจกๅž‹๏ผˆ้€‚้… 16GB ๆ˜พๅญ˜๏ผ‰ ๆœฌๆจกๅž‹ๅŸบไบŽ Qwen3.8-27B Uncensored ่ฟ›่กŒ IQ4_XS ้‡ๅŒ–๏ผˆ4โ€‘bit๏ผ‰๏ผŒๆ–‡ไปถไฝ“็งฏไธบ 12.9 GiB๏ผŒไธ“ไธบ 16GB ๆ˜พๅญ˜ ็š„ๆ˜พๅกไผ˜ๅŒ–๏ผŒๅœจไฟๆŒ่พƒไฝŽๅ›ฐๆƒ‘ๅบฆ็š„ๅŒๆ—ถ๏ผŒๅ…ผ้กพๆŽจ็†้€Ÿๅบฆๅ’Œๆ˜พๅญ˜ๅ ็”จใ€‚

ไธŽๅŒไฝ“็งฏ็š„ UD_IQ3_K_XL๏ผˆ12.5 GiB๏ผ‰้‡ๅŒ–ๆ–นๆกˆ่ฟ›่กŒไบ†ๅ…จ้ขๅฏนๆฏ”๏ผŒ่ฏ„ไผฐๆŒ‡ๆ ‡ๅฆ‚ไธ‹ใ€‚

๐Ÿ“Š ้‡ๅŒ–่ดจ้‡ๅฏนๆฏ”

่ฏ„ไผฐๆŒ‡ๆ ‡ IQ4_XS (ๆœฌๆจกๅž‹) UD_IQ3_K_XL (ๅฏนๆฏ”)
ๆ–‡ไปถๅคงๅฐ 12.9 GB 12.5 GB
้‡ๅŒ–็ฒพๅบฆ IQ4_XS (4โ€‘bit) UD_IQ3_K_XL (็บฆ 3โ€‘bit?)
้‡ๅŒ–ๆจกๅž‹ๅ›ฐๆƒ‘ๅบฆ (Mean PPL) 7.1481 ยฑ 0.0465 7.1117 ยฑ 0.0459
ไธŽๅŸบๅบงๆจกๅž‹ PPL ็›ธๅ…ณๆ€ง 99.28% 99.31%
ๅนณๅ‡ KL ๆ•ฃๅบฆ (Mean KLD) 0.03268 ยฑ 0.00030 0.03130 ยฑ 0.00032
ๆœ€ๅคง KL ๆ•ฃๅบฆ (Max KLD) 16.017๏ผˆๆ›ดๅฐ๏ผ‰ 21.409
99.9% KL ๅˆ†ไฝๆ•ฐ 1.075 1.219
Topโ€‘1 ไธ€่‡ด็އ (Same top p) 91.655% ยฑ 0.072% 92.419% ยฑ 0.069%
ๅนณๅ‡ๆฆ‚็އๅ˜ๅŒ– (Mean ฮ”p) -0.343% ยฑ 0.013%๏ผˆๆ›ดๆŽฅ่ฟ‘ 0๏ผ‰ -0.738% ยฑ 0.013%
RMS ๆฆ‚็އๅ˜ๅŒ– (RMS ฮ”p) 4.986% ยฑ 0.039%๏ผˆๆ›ดๅฐ๏ผ‰ 5.120% ยฑ 0.046%

ๅœจไธๅฏ็”จMTP็š„ๆƒ…ๅ†ตไธ‹ๅฏไปฅๅšๅˆฐ16GiBๅ‡€็ฉบVRAM๏ผˆไธไฝœไธบWindows็š„ๆ˜พ็คบๆ˜พๅก๏ผ‰็š„ๆƒ…ๅ†ตไธ‹110kไธŠไธ‹ๆ–‡

ๅผ€ๅฏMTPๅคงๆฆ‚80kไธŠไธ‹ๆ–‡ใ€‚


license: apache-2.0 base_model: - Qwen/Qwen3.8-27B

Qwen3.8-27B Uncensored IQ4_XS Quantized Model (Optimized for 16GB VRAM) This model is based on Qwen3.8-27B Uncensored and quantized with IQ4_XS (4โ€‘bit), with a file size of 12.9 GiB. It is tailored for GPUs with 16GB VRAM, balancing low perplexity, inference speed, and memory usage.

We conducted a comprehensive comparison against the UD_IQ3_K_XL quantization scheme (12.5 GiB, roughly 3โ€‘bit) of the same model size. The evaluation metrics are as follows.

๐Ÿ“Š Quantization Quality Comparison

Metric IQ4_XS (this model) UD_IQ3_K_XL (baseline)
File size 12.9 GB 12.5 GB
Quantization precision IQ4_XS (4โ€‘bit) UD_IQ3_K_XL (~3โ€‘bit)
Mean perplexity (quantized) 7.1481 ยฑ 0.0465 7.1117 ยฑ 0.0459
Correlation with base model PPL 99.28% 99.31%
Mean KL divergence 0.03268 ยฑ 0.00030 0.03130 ยฑ 0.00032
Maximum KL divergence 16.017 (lower) 21.409
99.9% KL quantile 1.075 1.219
Topโ€‘1 agreement rate 91.655% ยฑ 0.072% 92.419% ยฑ 0.069%
Mean probability change (Mean ฮ”p) -0.343% ยฑ 0.013% (closer to 0) -0.738% ยฑ 0.013%
RMS probability change (RMS ฮ”p) 4.986% ยฑ 0.039% (lower) 5.120% ยฑ 0.046%

With MTP disabled, the model can achieve ~110k context length while keeping ~16 GiB free VRAM (when not used as the primary display GPU on Windows). With MTP enabled, the context length is around 80k.

Downloads last month
9,697
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Bucoid/Qwen3.8-27B-Uncensored-IQ4-XS-MTP-16GB-VRAM-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(776)
this model