SenseNova U1.5 8B MoT - GGUF Quantizations

Optimized GGUF quantization formats for SenseNova-U1.5-8B-MoT, tailored for lower-VRAM setups (such as 8GB NVIDIA GPUs) via ComfyUI.

The tracking tags above link directly back to the main repository model page to aggregate download statistics.

Available Quants

Files are engineered to balance performance and memory footprints using high-precision F16 tensor swaps for sensitive layers (patch_embedding, dense_embedding, and fm_head) to preserve generation quality:

  • Q8_0 (~18.6 GB) β€” Near-lossless quality.
  • Q6_K (~14.5 GB) β€” High fidelity.
  • Q5_K_M (~12.5 GB) β€” Balanced tier.
  • Q4_K_M (~10.5 GB) β€” Standard compressed tier.
  • Q3_K_M (~8.5 GB) β€” Recommended tight VRAM profile for 8GB cards.
  • Q2_K (~6.4 GB) β€” Maximum compression.

Required Custom Node Fork

To run these GGUF files natively without structural AttributeError crashes or shape mismatches, use the dedicated node fork:

πŸ‘‰ GitHub: ComfyUI_SenseNova_U1_REBEL

Installation & Usage

  1. Place your chosen .gguf file into your ComfyUI models directory (e.g., ComfyUI/models/gguf/).
  2. Ensure you have installed the ComfyUI_SenseNova_U1_REBEL custom node suite.
  3. Load the model using the SenseNova_SM_Model loader node and hook it up to the SenseNova_SM_Sampler.
Downloads last month
-
GGUF
Model size
18B params
Architecture
wan
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for realrebelai/SenseNova-U1.5-8B_GGUFs

Quantized
(1)
this model