File size: 2,823 Bytes
f5acbf0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1c35f60
 
f5acbf0
 
 
 
 
 
 
 
 
 
 
 
1c35f60
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f5acbf0
 
 
1c35f60
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
---
license: apache-2.0
base_model: InternScience/Agents-A1-4B
base_model_relation: quantized
tags:
- nvfp4
- compressed-tensors
- llm-compressor
- vllm
- qwen3_5
- agent
---

# Agents-A1-4B-NVFP4

NVFP4 (W4A4, compressed-tensors) quantization of [InternScience/Agents-A1-4B](https://huggingface.co/InternScience/Agents-A1-4B) — a 4B agentic model on the Qwen3.5 dense architecture (`Qwen3_5ForConditionalGeneration`: hybrid DeltaNet linear attention + full attention every 4 layers, Qwen3-VL vision tower, 262K context).

- **Size:** 4.21 GB (from ~8 GB bf16), incl. 0.24 GB bf16 MTP draft head
- **Bonus — grafted MTP:** the source checkpoint declares `mtp_num_hidden_layers: 1` but ships no `mtp.*` weights. This repo grafts the 15-tensor bf16 MTP draft head from the base model [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (identical text-config dims), enabling speculative decoding: **63–66% draft acceptance, ~+20% single-stream decode** (measured, see below)
- **Scheme:** NVFP4 W4A4, group size 16, via llm-compressor 0.11.0 / compressed-tensors 0.16.0
- **Kept bf16:** `lm_head`, vision tower (`model.visual*`), DeltaNet `conv1d`
- **Calibration:** 32 samples × 8192 seq, `neuralmagic/calibration` (pure-CPU calibration — no GPU used in the bake)

## Serve with vLLM

```bash
vllm serve sakamakismile/Agents-A1-4B-NVFP4
```

Quantization is auto-detected — no `--quantization` flag needed. Requires a GPU with FP4 support (SM120 Blackwell) for the NVFP4 kernels.

With speculative decoding (grafted MTP draft, ~+20% single-stream):

```bash
vllm serve sakamakismile/Agents-A1-4B-NVFP4 \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'
```

## Measured throughput — single GPU (vLLM 0.22.0, RTX PRO 2000 Blackwell SM120, KV fp8, 512-tok decode)

| Config | single-stream (c1) | aggregate 8-way (c8) | vs base |
|---|---|---|---|
| base (no MTP) | 73.6 t/s | 492.9 t/s | — |
| **MTP n=3 (grafted)** | **88.6 t/s** | 474.5 t/s | **+20.4% c1** / −3.7% c8 |

SpecDecoding metrics (vLLM): mean acceptance length ~2.9, per-position acceptance 0.83 / 0.66 / 0.48, avg draft acceptance **63–66%** — remarkable for a draft head grafted from the pre-fine-tune base model across InternScience's 3-stage agentic distillation. As usual, MTP is a latency win (single-stream / low concurrency); at saturation the extra draft forward costs slightly more than it saves.

Speculative decoding is lossless — outputs are identical to base decoding.

## Recipe

Baked (pure-CPU, llm-compressor) with the same validated recipe as [Qwen3.6-27B-MTP-pi-tune-NVFP4](https://huggingface.co/sakamakismile/Qwen3.6-27B-MTP-pi-tune-NVFP4) and [ThinkingCap-Qwen3.6-27B-NVFP4](https://huggingface.co/sakamakismile/ThinkingCap-Qwen3.6-27B-NVFP4) (same `qwen3_5` architecture family).