--- license: apache-2.0 base_model: InternScience/Agents-A1-4B base_model_relation: quantized tags: - nvfp4 - compressed-tensors - llm-compressor - vllm - qwen3_5 - agent --- # Agents-A1-4B-NVFP4 NVFP4 (W4A4, compressed-tensors) quantization of [InternScience/Agents-A1-4B](https://huggingface.co/InternScience/Agents-A1-4B) — a 4B agentic model on the Qwen3.5 dense architecture (`Qwen3_5ForConditionalGeneration`: hybrid DeltaNet linear attention + full attention every 4 layers, Qwen3-VL vision tower, 262K context). - **Size:** 4.21 GB (from ~8 GB bf16), incl. 0.24 GB bf16 MTP draft head - **Bonus — grafted MTP:** the source checkpoint declares `mtp_num_hidden_layers: 1` but ships no `mtp.*` weights. This repo grafts the 15-tensor bf16 MTP draft head from the base model [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (identical text-config dims), enabling speculative decoding: **63–66% draft acceptance, ~+20% single-stream decode** (measured, see below) - **Scheme:** NVFP4 W4A4, group size 16, via llm-compressor 0.11.0 / compressed-tensors 0.16.0 - **Kept bf16:** `lm_head`, vision tower (`model.visual*`), DeltaNet `conv1d` - **Calibration:** 32 samples × 8192 seq, `neuralmagic/calibration` (pure-CPU calibration — no GPU used in the bake) ## Serve with vLLM ```bash vllm serve sakamakismile/Agents-A1-4B-NVFP4 ``` Quantization is auto-detected — no `--quantization` flag needed. Requires a GPU with FP4 support (SM120 Blackwell) for the NVFP4 kernels. With speculative decoding (grafted MTP draft, ~+20% single-stream): ```bash vllm serve sakamakismile/Agents-A1-4B-NVFP4 \ --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}' ``` ## Measured throughput — single GPU (vLLM 0.22.0, RTX PRO 2000 Blackwell SM120, KV fp8, 512-tok decode) | Config | single-stream (c1) | aggregate 8-way (c8) | vs base | |---|---|---|---| | base (no MTP) | 73.6 t/s | 492.9 t/s | — | | **MTP n=3 (grafted)** | **88.6 t/s** | 474.5 t/s | **+20.4% c1** / −3.7% c8 | SpecDecoding metrics (vLLM): mean acceptance length ~2.9, per-position acceptance 0.83 / 0.66 / 0.48, avg draft acceptance **63–66%** — remarkable for a draft head grafted from the pre-fine-tune base model across InternScience's 3-stage agentic distillation. As usual, MTP is a latency win (single-stream / low concurrency); at saturation the extra draft forward costs slightly more than it saves. Speculative decoding is lossless — outputs are identical to base decoding. ## Recipe Baked (pure-CPU, llm-compressor) with the same validated recipe as [Qwen3.6-27B-MTP-pi-tune-NVFP4](https://huggingface.co/sakamakismile/Qwen3.6-27B-MTP-pi-tune-NVFP4) and [ThinkingCap-Qwen3.6-27B-NVFP4](https://huggingface.co/sakamakismile/ThinkingCap-Qwen3.6-27B-NVFP4) (same `qwen3_5` architecture family).