Update README.md
Browse files
README.md
CHANGED
|
@@ -144,7 +144,7 @@ This abliterated checkpoint is a drop-in replacement for the original weights
|
|
| 144 |
|
| 145 |
For inference guidance specific to the **NVIDIA RTX PRO 6000 Blackwell** (TP2/TP4, the `lucifer-default` / `lucifer-cutlass` / `b12x` backends, and the native DSpark `method=dspark` speculative-decoding path with `num_speculative_tokens=5`), see the community guide:
|
| 146 |
|
| 147 |
-
👉 **https://github.com/local-inference-lab/rtx6kpro/blob/master/models/ds4dspark-
|
| 148 |
|
| 149 |
That page documents the validated Docker image, launch helpers, and full TP2/TP4 throughput sweep (decode + prefill) for this exact checkpoint on RTX PRO 6000, including DSpark speculative decoding which gives a ~41% single-stream throughput speedup on this abliterated release.
|
| 150 |
|
|
|
|
| 144 |
|
| 145 |
For inference guidance specific to the **NVIDIA RTX PRO 6000 Blackwell** (TP2/TP4, the `lucifer-default` / `lucifer-cutlass` / `b12x` backends, and the native DSpark `method=dspark` speculative-decoding path with `num_speculative_tokens=5`), see the community guide:
|
| 146 |
|
| 147 |
+
👉 **https://github.com/local-inference-lab/rtx6kpro/blob/master/models/ds4dspark-v9.md**
|
| 148 |
|
| 149 |
That page documents the validated Docker image, launch helpers, and full TP2/TP4 throughput sweep (decode + prefill) for this exact checkpoint on RTX PRO 6000, including DSpark speculative decoding which gives a ~41% single-stream throughput speedup on this abliterated release.
|
| 150 |
|