--- base_model: nvidia/Nemotron-Cascade-2-30B-A3B license: other license_name: nvidia-open-model-license license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/ library_name: mlx pipeline_tag: text-generation tags: - mlx - quantized - turboquant - moe --- # Nemotron-Cascade-2-30B-A3B — TurboQuant MLX 5bit [`nvidia/Nemotron-Cascade-2-30B-A3B`](https://huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B) quantized pack, published as `Nemotron-Cascade-2-30B-A3B-TurboQuant-MLX-5bit`. ## Method MLX quantization via mlx_lm (5bit, group_size 64). ## Release line Released under the **TurboQuant** line. RotorQuant and TurboQuant are this project's release labels for this pack, not distinct quantization algorithms — both brand repos carry byte-identical weights. No brand-specific speedup is claimed or measured. ## Modality `pipeline_tag: text-generation`. This is a Mixture-of-Experts (MoE) model — a subset of experts is active per token; total and active parameter counts differ. ## License This pack is a **derivative** of [`nvidia/Nemotron-Cascade-2-30B-A3B`](https://huggingface.co/nvidia/Nemotron-Cascade-2-30B-A3B); all credit for the original model, training, and weights belongs to the upstream authors. This repo republishes a quantized conversion of those weights only. Governed by the [nvidia-open-model-license](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/). See the upstream repo and the linked license for the full terms — no license text is reproduced here.