Qwen2.5-Coder-3B-HXQ

Qwen2.5-Coder-3B-Instruct compressed with HXQ (HelixCode vector quantization).

Available as both HuggingFace safetensors (via helix-substrate) and native GGUF (via llama.cpp HXQ fork).

GGUF Runtime Benchmark (RTX 3090)

Benchmarked against standard GGUF K-quants on RTX 3090, full GPU offload (-ngl 99), using the hxq-affine-type branch at commit 580e9a2.

Decode Speed (tg128, 3 runs)

Format Size bpw tok/s vs Q4 vs Q6
Q4_K_M 1.95 GB 4.5 245.03 100% 120%
Q5_K_M 2.07 GB 5.5 229.00 93.5% 112%
HXQ_AF6 2.26 GB 6.25 226.53 92.4% 110.6%
Q6_K 2.36 GB 6.56 204.86 83.6% 100%

Perplexity (WikiText-2, 50 chunks, ctx=512)

Format bpw PPL vs Q4
HXQ_AF6 6.25 9.954 -0.118 (best)
Q6_K 6.56 9.964 -0.108
Q5_K_M 5.5 10.004 -0.068
Q4_K_M 4.5 10.072 baseline

HumanEval (evalplus v0.3.1, greedy, pass@1)

Format HumanEval base HumanEval+
HXQ_AF6 84.1% 78.0%
Q5_K_M 83.5% 75.6%
Q6_K 83.5% 78.0%
Q4_K_M 82.3% 78.0%

Prefill (pp512, 3 runs)

Format tok/s vs Q4
Q4_K_M 8974 100%
Q5_K_M 8543 95.2%
Q6_K 8173 91.1%
HXQ_AF6 7430 82.8%

Summary: HXQ_AF6 has the lowest perplexity, the highest HumanEval base pass@1, and decodes 10.6% faster than Q6_K while being smaller (2.26 vs 2.36 GB). It trades ~7.6% decode speed vs Q4_K_M for better quality preservation at 6.25 bpw.

Reproducibility

All claims are within-run comparisons using the same dataset, llama.cpp commit, and hardware. Do not compare these PPL numbers with numbers from other runs using different model variants, dataset files, or build configurations.

Full receipt with SHA256 artifact hashes, exact commands, and raw HumanEval outputs: bench_receipt.json

Install and Run

Option 1: Native GGUF (llama.cpp)

# Build llama.cpp with HXQ support
git clone -b hxq-affine-type https://github.com/echo313unfolding/llama.cpp.git
cd llama.cpp && mkdir build && cd build
cmake .. -DGGML_CUDA=ON && make -j$(nproc) llama-cli

# Run
./bin/llama-cli -m qwen2.5-coder-3b-instruct-hxq-affine6.gguf \
  -ngl 99 -p "def fibonacci(n):" -n 128

Option 2: HuggingFace (Python)

pip install "helix-substrate[hf]"
import helix_substrate  # registers the HXQ quantizer with HuggingFace
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("EchoLabs33/qwen2.5-coder-3b-hxq")
tokenizer = AutoTokenizer.from_pretrained("EchoLabs33/qwen2.5-coder-3b-hxq")

inputs = tokenizer("def fibonacci(n):", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Safetensors Benchmark

Dense (BF16) HXQ (safetensors)
Size 6.2 GB 3.84 GB
Perplexity (WikiText-2, 2048 ctx) 6.113 6.230 (+1.92%)
Compression ratio 1x 1.6x
Compressed modules 0 252 HelixLinear layers

Note: The safetensors PPL (6.230) and GGUF PPL (9.954) use different evaluation configurations (ctx=2048/stride=512 vs ctx=512/50 chunks). They are not directly comparable.

Good to Know

  • GPU and CPU supported — runs on any CUDA GPU or CPU via standard PyTorch. Native GGUF runs via llama.cpp.
  • Fine-tunable via LoRA — compressed weights remain frozen, but LoRA adapters attach to each HelixLinear layer via HelixLinearSTE. See helix-substrate for training infrastructure.
  • Requires helix-substrate for safetensors path — the quantizer is not built into transformers. You need pip install "helix-substrate[hf]".
  • Requires llama.cpp HXQ fork for GGUF path — standard llama.cpp does not have HXQ type support yet.
  • Tied embeddingslm_head shares embed_tokens, stored at full precision.

What is HXQ?

HXQ is a weight compression codec based on vector quantization with per-group affine correction:

  • Each weight matrix is replaced by a 256-entry codebook + uint8 index matrix + per-group affine scale/offset
  • The compressed form is the executable — codebook[indices] * scale + offset during matmul, no decompression step
  • Works on any nn.Linear regardless of architecture (Transformer, Mamba, MLP)
  • No calibration data required — codebooks are fit from the weights alone via k-means
  • 6.25 bits per weight in the GGUF affine-6 format

Companion Models

Same codec, multiple architectures:

Model Architecture GGUF Safetensors
qwen2.5-7b-instruct-hxq Transformer Yes Yes
qwen2.5-3b-instruct-hxq Transformer Yes Yes
qwen2.5-coder-1.5b-hxq Transformer (code) Yes Yes
qwen2.5-14b-instruct-hxq Transformer Yes Yes
qwen2.5-sentinel-3b-hxq Transformer (security) Yes

Citation

@software{hxq_2026,
  title={HXQ: Vector Quantization with Per-Group Affine Correction for Neural Network Weight Compression},
  author={Echo Labs},
  year={2026},
  url={https://github.com/echo313unfolding/helix-substrate}
}

License

Apache 2.0 (inherited from Qwen/Qwen2.5-Coder-3B-Instruct).

Downloads last month
53
Safetensors
Model size
3B params
Tensor type
I64
·
F32
·
F16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EchoLabs33/qwen2.5-coder-3b-hxq

Base model

Qwen/Qwen2.5-3B
Quantized
(117)
this model

Collection including EchoLabs33/qwen2.5-coder-3b-hxq

Evaluation results