Configuration Parsing Warning:Invalid JSON for config file config.json

KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively — use -ctk q8_0 -ctv q8_0 (half KV memory, negligible quality loss: perplexity +0.002–0.05) or -ctk q4_0 -ctv q4_0 (quarter memory, ≈7.6% perplexity increase). In Ollama: OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1. Keep K and V types symmetric to stay on the fast fused Flash-Attention path. Since April 2026, mainline llama.cpp also applies Hadamard rotation to KV activations (PR #21038), which greatly improves low-bit KV quality (opt-out: LLAMA_ATTN_ROT_DISABLE=1).

The RotorQuant/TurboQuant fork flow below is experimental/legacy: the TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork is unmaintained relative to mainline. It is NOT required to use this model.

Nemotron-3-Nano-30B-A3B - RotorQuant MLX 4-bit

4-bit weight-quantized MLX version of nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 with the legacy RotorQuant KV-cache fork (superseded by upstream llama.cpp KV options). Optimized for Apple Silicon inference via the MLX framework. A good balance between model quality and memory efficiency. Only 3.2B parameters are active per token despite 30.7B total, making this model significantly more efficient at inference time than its parameter count suggests. The hybrid Mamba-2 + Transformer MoE architecture supports up to 1M context length.

Approximate model size: ~17 GB

Model Specifications

Property Value
Base Model nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Parameters 30.7 billion total (3.2 billion active per token)
Architecture Hybrid Mamba-2 + Transformer MoE (3.2B active per token)
Context Length 1,048,576 tokens (1M)
License NVIDIA Open Model License (commercial use OK)
Weight Quantization 4-bit (~17 GB)
KV-Cache Quantization RotorQuant
Framework MLX (Apple Silicon)

Quickstart

from mlx_lm import load, generate
from rotorquant import IsoQuantCache

model, tokenizer = load("majentik/Nemotron-3-Nano-30B-A3B-RotorQuant-MLX-4bit")

prompt = "Explain the theory of relativity."
response = generate(model, tokenizer, prompt=prompt, max_tokens=512)
print(response)

About the RotorQuant / TurboQuant labels

RotorQuant and TurboQuant are this project's release labels, not distinct quantization algorithms — for any given tier, both brand repos carry byte-identical weights produced with the standard MLX / llama.cpp quantizers. No brand-specific speedup is claimed or measured. The KV-cache fork these labels originally referred to is legacy; for KV-cache memory savings use the upstream options described above (-ctk/-ctv q8_0, OLLAMA_KV_CACHE_TYPE).

KV-Cache Quantization Comparison

Method Prefill Speed Decode Speed Memory Savings Reference
TurboQuant 1x (baseline) 1x (baseline) High arXiv: 2504.19874

Memory Estimates (Nemotron-3-Nano-30B-A3B)

Precision Approximate Size MLX Variant
BF16 (original) ~60 GB --
8-bit quantized ~30 GB RotorQuant-MLX-8bit
4-bit quantized ~17 GB This model
2-bit quantized ~9 GB RotorQuant-MLX-2bit

Hardware Requirements

This model requires approximately 17 GB of unified memory. Recommended hardware:

  • Apple M2 Pro (24 GB+)
  • Apple M3 Pro (24 GB+)
  • Apple M4 Pro (24 GB+)
  • Any Apple Silicon Mac with 24 GB+ unified memory

See Also

Downloads last month
77
Safetensors
Model size
32B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for majentik/Nemotron-3-Nano-30B-A3B-RotorQuant-MLX-4bit

Finetuned
(56)
this model

Collection including majentik/Nemotron-3-Nano-30B-A3B-RotorQuant-MLX-4bit

Paper for majentik/Nemotron-3-Nano-30B-A3B-RotorQuant-MLX-4bit