xv0y5ncu/SmolLM3-3B-trellis-3inst-4bpw-kernel
Text Generation • 1.0B • Updated • 835
One curated pick per model family - current codebooks, true-to-label footprints. pip install glq
Note Flagship. Trellis (TCQ) codebook, v0.7+: bf16-class decode speed at 4 bpw.
Note MoE. Trellis (TCQ) 4 bpw at 13.95 GB - the smallest 26B GLQ checkpoint - and the only one with the fused grouped MoE decode: 91.6 tok/s at B=1 under a CUDA graph (RTX PRO 6000). Needs glq >= 0.8.1. AIME-2026 avg@8 85.4%.
Note 31B on a single GPU: 16.5 GiB vs 57.9 GiB bf16.
Note 12B, mixed 3-8 bpw allocation.
Note Compact Gemma-4 for 24 GB cards.
Note Code model, 24B at 4 bpw.
Note Block-diagonal RHT: footprint tracks the bpw budget.
Note Tiny; good for CI and demos.
Note Smallest checkpoint; CI smoke tests.