RedHatAI/Qwen3.6-35B-A3B-speculator.dspark

This is a DSpark speculator model for Qwen/Qwen3.6-35B-A3B.

Training Details

This model was trained using the Speculators library library on a subset of Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered and the train_sft split of HuggingFaceH4/ultrachat_200k. Responses were regenerated by Qwen3.6-35B-A3B. DSpark extends DFlash with a Markov head (intra-block token dependency) and a confidence head (per-position acceptance prediction). Training compute for this model was generously provided by Lambda, a leading cloud platform for AI training and inference.

Commands

Using the Speculators library and the helper scripts provided in the repo.

Prepare data

# In virtual environment with speculators installed
python scripts/prepare_data.py \
  --model Qwen/Qwen3.6-35B-A3B \
  --data ./data/qwen3_6_35b_a3b_regen/magpie_output.jsonl \
  --data ./data/qwen3_6_35b_a3b_regen/ultrachat_output.jsonl \
  --output ./output/dspark_qwen3_6_35b_a3b \
  --seq-length 16000

Launch vLLM

# In (separate) virtual environment with vllm installed
CUDA_VISIBLE_DEVICES=0,1,2,3 python scripts/launch_vllm.py \
  Qwen/Qwen3.6-35B-A3B \
  --target-layer-ids 2 10 20 30 37 \
  -- --port 8000 \
  --tensor-parallel-size 4 \
  --max-model-len 16002 \
  --enforce-eager

Launch training

Must be run once vLLM has finished launching and is running in the background.

# In virtual environment with speculators installed
CUDA_VISIBLE_DEVICES=4,5,6,7 torchrun \
  --standalone \
  --nproc_per_node 4 \
  scripts/train.py \
  --verifier-name-or-path Qwen/Qwen3.6-35B-A3B \
  --speculator-type dspark \
  --data-path ./output/dspark_qwen3_6_35b_a3b \
  --vllm-endpoint http://localhost:8000/v1 \
  --save-path ./output/dspark_qwen3_6_35b_a3b/checkpoints \
  --block-size 8 \
  --max-anchors 3072 \
  --target-layer-ids 2 10 20 30 37 \
  --num-layers 5 \
  --markov-rank 256 \
  --markov-head-type vanilla \
  --enable-confidence-head \
  --confidence-head-with-markov \
  --loss-fn '{"ce": 0.1, "tv": 0.9}' \
  --confidence-head-alpha 1.0 \
  --epochs 5 \
  --lr 3e-4 \
  --total-seq-len 16000 \
  --on-missing generate \
  --on-generate delete \
  --checkpoint-freq 0.1

Model Specifications

Base Model Qwen/Qwen3.6-35B-A3B
Chat Template Qwen/Qwen3.6-35B-A3B (use /chat/completions endpoint)
Format Safetensors
License Apache 2.0
Validation Hardware Nvidia H100

Deployment

# Deploy with speculative decoding
vllm serve Qwen/Qwen3.6-35B-A3B \
    --tensor-parallel-size 2 \
    --speculative-config '{
        "model": "RedHatAI/Qwen3.6-35B-A3B-speculator.dspark",
        "num_speculative_tokens": 8,
        "method": "dspark"
    }'

Acceptance Rates

Per-position token acceptance rates across datasets:

Dataset Pos 0 Pos 1 Pos 2 Pos 3 Pos 4 Pos 5 Pos 6 Pos 7 Avg. Length
HumanEval 82.0% 66.2% 54.1% 44.1% 36.4% 29.9% 24.7% 20.4% 4.58
math_reasoning 84.0% 70.3% 59.8% 51.2% 43.5% 37.0% 31.2% 26.5% 5.03
qa 71.8% 50.9% 37.4% 28.0% 20.9% 16.0% 12.1% 9.2% 3.46
question 73.5% 53.1% 39.5% 30.4% 24.2% 19.5% 15.6% 12.7% 3.68
rag 78.2% 57.1% 44.1% 34.7% 27.3% 21.2% 16.4% 13.1% 3.92
summarization 74.2% 53.6% 40.4% 30.6% 23.5% 17.8% 13.7% 10.5% 3.64
tool_call 71.7% 50.8% 36.1% 26.3% 19.5% 14.7% 11.2% 8.6% 3.39
translation 70.7% 51.2% 37.8% 28.4% 21.2% 15.7% 11.9% 8.8% 3.46
writing 74.4% 54.3% 40.8% 31.6% 25.2% 20.3% 16.5% 13.4% 3.76
Downloads last month
2,182
Safetensors
Model size
1.0B params
Tensor type
I64
·
BF16
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Qwen3.6-35B-A3B-speculator.dspark

Finetuned
(223)
this model

Collection including RedHatAI/Qwen3.6-35B-A3B-speculator.dspark