nemotron-super-120b-cc-mt-curriculum-nodr-500m

Nemotron 3 Super 120B-A12B continued-pretraining (midtraining) checkpoint from the constitutional-curriculum midtraining study (cc-mt).

Training

  • Base checkpoint: NVIDIA-Nemotron-3-Super-120B-A12B-Base-Chat-Init-BF16 (pure Base weights; 1,188 chat-scaffolding embedding rows grafted from Instruct. No post-training behavior included.)
  • Data: 1:1 blend of cho-ai/constitutional-curriculum-mt-data data/curriculum_noDR_500M.jsonl (~509M tokens, consumed in exact file order, no-deliberative-reasoning 500M variant) and shuffled pretraining replay (geodesic-research/Nemotron-Pretraining-Specialized)
  • Data order: FILE-ORDER curriculum (boundaries at iters 243/486/729)
  • Recipe ("dyad-1" Base-CPT stability stack): lr 1e-6 cosine w/ 10% warmup, FP32 optimizer states, GBS 128, seq 8192, BF16, 971 iterations = 1.018B tokens (1 epoch). Tokenizer: geodesic-research/nemotron-base-tokenizer (EOD=</s>=id 2). All data filtered against the 1,188 zero-embedding token ids of Super-Base (0 docs dropped).
  • Topology: TP=1 EP=4 PP=22 ETP=1 (parallel folding), 22 nodes / 88 GH200 GPUs on Isambard-AI.
  • W&B: geodesic/megatron_training/w9j966ts

Usage notes

Base-style model (pretraining-format CPT; no instruction tuning) — use completion-style prompting. Loads with native transformers NemotronHForCausalLM; tokenizer PreTrainedTokenizerFast. Coherence-checked post-export (8/8 non-empty, fluent completions).

SFT companions: nemotron-super-120b-cc-mt-curriculum-nodr-500m-sr-sft (reasoning) and nemotron-super-120b-cc-mt-curriculum-nodr-500m-200k-sft (non-reasoning).

Downloads last month
6
Safetensors
Model size
124B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for geodesic-research/nemotron-super-120b-cc-mt-curriculum-nodr-500m

Finetuned
(9)
this model
Finetunes
2 models