---
license: other
license_name: nvidia-open-model-license
pipeline_tag: text-generation
provider: NVIDIA
tool_calling_supported: true
required_cli_args: ['--reasoning-parser nemotron_v3', '--tool-call-parser qwen3_coder']
chat_template_file_name: None
chat_template_path: None
tool_call_parser: qwen3_coder
validated_tasks:
- tool-calling
tasks:
- text-to-text
- text-generation
- reasoning
- tool-calling
base_model:
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
name: RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16
description: This model was obtained by quantizing the weights of nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 to INT4 (W4A16) data type.
readme: https://huggingface.co/RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16/blob/main/README.md
tags:
- int4
- w4a16
- quantized
- llm-compressor
- compressed-tensors
- red hat
validated_on:
- RHOAI 3.5
- RHAIIS 3.5
- vLLM 0.24.0
---
NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16 
## Model Overview
- **Model Architecture:** NemotronHForCausalLM
- **Input:** Text
- **Output:** Text
- **Total Parameters:** 550B
- **Active Parameters:** 55B
- **Model Optimizations:**
- **Weight quantization:** INT4 (W4A16, group size 128)
- **Intended Use Cases:**
- Reasoning and complex problem solving.
- Mathematics and science.
- Code generation.
- Instruction following.
- **Out-of-scope:** Use in any manner that violates applicable laws or regulations (including trade compliance laws).
- **Release Date:** 06/04/2025
- **Version:** 1.0
- **Model Developers:** Red Hat
Quantized version of [nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16).
### Model Optimizations
This model was obtained by quantizing the weights of [nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16) to INT4 data type.
This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%.
Only the weights of the linear operators within transformer blocks are quantized.
Weights are quantized using an asymmetric per-group scheme with group size 128.
The [llm-compressor](https://github.com/vllm-project/llm-compressor) library is used for quantization.
## Deployment
### Use with vLLM
This model can be deployed efficiently using the [vLLM](https://docs.vllm.ai/en/latest/) backend.
**Install dependencies:**
```bash
uv pip install git+https://github.com/vllm-project/vllm.git
uv pip install llmcompressor
```
**Launch the vLLM server:**
```bash
vllm serve RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16 \
--host 0.0.0.0 --port 8088 \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--max-num-seqs 32 \
--max-num-batched-tokens 32768 \
--enable-chunked-prefill \
--enable-prefix-caching \
--reasoning-parser nemotron_v3 \
--mamba-ssm-cache-dtype float16 \
--mamba-backend flashinfer \
--enable-mamba-cache-stochastic-rounding \
--mamba-cache-philox-rounds 5 \
--speculative-config '{"method": "nemotron_h_mtp", "num_speculative_tokens": 5}' \
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 96}' \
--trust-remote-code
```
**Send requests:**
```python
from openai import OpenAI
openai_api_key = "EMPTY"
openai_api_base = "http://localhost:8088/v1"
client = OpenAI(
api_key=openai_api_key,
base_url=openai_api_base,
)
model = "RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16"
messages = [
{"role": "user", "content": "Solve for x: 2x + 5 = 13"},
]
outputs = client.chat.completions.create(
model=model,
messages=messages,
)
generated_text = outputs.choices[0].message.content
print(generated_text)
```
## Creation
This model was quantized using the [llm-compressor](https://github.com/vllm-project/llm-compressor) library as shown below.
```python
from llmcompressor import model_free_ptq
MODEL_ID = "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16"
SAVE_DIR = MODEL_ID.rstrip("/").split("/")[-1] + "-W4A16-G128"
model_free_ptq(
model_stub=MODEL_ID,
save_directory=SAVE_DIR,
scheme="W4A16",
ignore=[
"re:.*gate$",
"lm_head",
"model.embed_tokens",
"re:.*mixer.conv1d.*",
"re:.*norm_f*",
"re:.*bias$",
"re:.*embed_tokens$",
"backbone.embeddings"
],
max_workers=15,
device="cuda:0",
)
```
## Evaluation
The model was evaluated on reasoning tasks using [lighteval](https://github.com/huggingface/lighteval).
[vLLM](https://docs.vllm.ai/en/stable/) was used as the serving backend for all evaluations.
**Install dependencies:**
```bash
uv pip install git+https://github.com/vllm-project/vllm.git
uv pip install lighteval==0.13.0
uv pip install "litellm[caching]>=1.66.0"
```
**Launch the vLLM server:**
```bash
vllm serve RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16 \
--host 0.0.0.0 --port 8088 \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--max-num-seqs 32 \
--max-num-batched-tokens 32768 \
--enable-chunked-prefill \
--enable-prefix-caching \
--reasoning-parser nemotron_v3 \
--mamba-ssm-cache-dtype float16 \
--mamba-backend flashinfer \
--enable-mamba-cache-stochastic-rounding \
--mamba-cache-philox-rounds 5 \
--speculative-config '{"method": "nemotron_h_mtp", "num_speculative_tokens": 5}' \
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 96}' \
--trust-remote-code
```
**AIME 2025:**
```bash
lighteval endpoint litellm \
"model_name=hosted_vllm/RedHatAI__NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16,provider=hosted_vllm,base_url=http://127.0.0.1:8088/v1,timeout=3600,concurrent_requests=32,generation_parameters={temperature:1.0,top_p:0.95,max_new_tokens:32768}" \
"aime25|0" \
--output-dir results --save-details
```
**GPQA Diamond:**
```bash
lighteval endpoint litellm \
"model_name=hosted_vllm/RedHatAI__NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16,provider=hosted_vllm,base_url=http://127.0.0.1:8088/v1,timeout=3600,concurrent_requests=32,generation_parameters={temperature:1.0,top_p:0.95,max_new_tokens:32768}" \
"gpqa:diamond|0" \
--output-dir results --save-details
```
### Accuracy
| Benchmark |
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 |
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 |
RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-Dynamic |
RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-BLOCK |
RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16 (this model) |
| AIME 2025 (pass@1) |
90.00 |
90.00 (100.0%) |
93.33 (103.7%) |
86.67 (96.3%) |
86.67 (96.3%) |
| GPQA Diamond (pass@1) |
78.79 |
84.85 (107.7%) |
82.32 (104.5%) |
81.31 (103.2%) |
81.82 (103.8%) |
| Average |
84.39 |
87.42 (103.6%) |
87.83 (104.1%) |
83.99 (99.5%) |
84.24 (99.8%) |
### Extended Evaluation
The model was evaluated across instruct, reasoning, and coding tasks. Scores are averaged over multiple seeds (3 seeds for most benchmarks, 8 for AIME 2025). Recovery is computed as the ratio of the quantized model score to the base model score.
| Category |
Benchmark |
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 |
RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16 (this model) |
Recovery |
| Instruct |
MMLU-CoT (5-shot) |
91.09 |
90.95 |
99.85% |
| Instruct |
GSM8K Platinum (5-shot) |
98.59 |
98.59 |
100.00% |
| Instruct |
IFEval (0-shot) |
91.00 |
89.03 |
97.84% |
| Instruct |
MATH-500 |
83.60 |
83.67 |
100.08% |
| Reasoning |
AIME 2025 |
64.17 |
61.67 |
96.10% |
| Reasoning |
MATH-500 |
85.80 |
85.93 |
100.15% |
| Reasoning |
GSM8K Platinum (0-shot) |
96.77 |
96.58 |
99.80% |
| Reasoning |
IFEval (0-shot) |
94.58 |
94.02 |
99.41% |
| Coding |
LCB CodeGen v6 |
51.81 |
48.19 |
93.01% |
### Tool Calling Evaluation
The model was evaluated on tool calling tasks using the [Berkeley Function-Calling Leaderboard v4 (BFCLv4)](https://gorilla.cs.berkeley.edu/leaderboard.html).
[vLLM](https://docs.vllm.ai/en/stable/) was used as the serving backend for all evaluations.
| Category |
Benchmark |
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 |
RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-quantized.w4a16 (this model) |
Recovery |
| Overall |
BFCLv4 Overall Acc |
55.44 |
53.95 |
97.31% |
| Single Turn |
Non-Live Acc |
45.00 |
44.58 |
99.07% |
| Single Turn |
Live Acc |
71.65 |
71.87 |
100.31% |
| Multi-Turn |
Multi-Turn Acc |
42.12 |
42.12 |
100.00% |
| Agentic |
Agentic Acc |
57.64 |
54.24 |
94.10% |