Solar-Open2 250B + Speculative Draft (combined, 500K draft)

This repository bundles Upstage Solar Open2 250B-A15B (original weights at the repository root, unchanged — see LICENSE and the tech report) together with our block-draft speculative-decoding draft model in draft/.

  • Root: Solar Open2 250B-A15B — 48 layers (12 NoPE GQA + 36 linear-attention), 320-expert MoE top-8, vocab 196,608. Original Upstage distribution files.
  • draft/: 5-layer draft (core 0.90B, total 2.51B with frozen target embedding/LM head), conditioned on five taps of the target's hidden states (layers 3/15/27/39/47). Proposes up to γ=7 tokens per step; the target verifies the block in one forward.

Draft results (500K-scale, 2 epochs)

Metric Value
val accept_len (γ=7) 3.431
Serving accept_len γ=3 / γ=7 2.75 / 3.54
gsm8k (5-shot greedy, spec on) 0.930–0.940 (lossless band)

Training: 500K-scale multi-domain collection (415K-prompt manifest), ~1.2B tokens through the target for hidden-state distillation. 2 epochs, chunked KL, 2048 stratified anchors per batch.

Usage

Serve the root model as usual for non-speculative inference. For speculative decoding, point your serving stack's block-draft speculative option at the draft/ folder (block size 4–7 recommended, draft unquantized). Requires a serving build with block-draft speculative decoding support for this draft format.

License

Root model weights follow the Upstage Solar Open2 license (LICENSE). The draft contains the target's frozen embedding/LM-head weights and follows the same terms.

Downloads last month
8
Safetensors
Model size
250B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support