GPT2Medium EN/IT NanoChat — continual-pretraining checkpoint step_34000

This is an ordinary checkpoint release from the main medium continual-pretraining run. It is not a separately launched decay-only branch. It is published as a reference candidate for the future 1gpu-llm-medium-v2 release; it is not yet that definitive release.

Released checkpoint

  • checkpoint: step_34000.pt
  • training run: 20260703_continual-pretraining-gpt2medium-gpt2preln-k20-step14700-lr5e5-w500-s18500-d2000-final1e5-webwiki
  • source family base: step_14700
  • training mode: continual pretraining, main CPT run (non-decay-only branch)
  • languages: English + Italian
  • context window: 2500 tokens
  • architecture: GPT-2-style decoder with pre-layernorm blocks
  • architecture identifiers: architecture: gpt2, block_type: gpt2_prelayernorm
  • parameter count: approximately 337.7M native training parameters; approximately 337.6M in the Transformers export
  • hardware: trained on a single RTX 4060 Ti 16GB

Why this checkpoint

step_34000 is currently the scalar champion of the evaluated medium checkpoint family:

  • val_loss_mixed = 4.4401
  • step_23100 22k decay branch: 4.4675
  • step_36000 d2000 branch from 34k: 4.4493

The checkpoint is therefore the current best base candidate for the medium v2 investigation, while the final 1gpu-llm-medium-v2 promotion remains intentionally open pending the remaining decoding/behavior comparison.

Training data

The model was trained on the bilingual EN/IT web + wiki corpus:

  • English FineWeb-HQ (epfml/FineWeb-HQ)
  • Italian FineWeb2-HQ (epfml/FineWeb2-HQ)
  • English and Italian Wiki40B (google/wiki40b)
  • local dataset: 202605141153_fineweb50_wiki50_50en_50it_score100_2500context_5Btokens_tok_20260515_en50it50_webwiki_stratified_500M

Quick start

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

repo_id = "nazdef/20260703_continual-pretraining-gpt2medium-step34000-nodecay"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)

prompt = "La capitale d'Italia è"
prompt_ids = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
bos = torch.tensor([[tokenizer.bos_token_id]], dtype=prompt_ids["input_ids"].dtype)
input_ids = torch.cat([bos, prompt_ids["input_ids"]], dim=1)
attention_mask = torch.ones_like(input_ids)

outputs = model.generate(
    input_ids=input_ids,
    attention_mask=attention_mask,
    do_sample=True,
    max_new_tokens=64,
    temperature=0.8,
    top_k=50,
    top_p=0.95,
    repetition_penalty=1.1,
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

This is a base pretraining checkpoint, not an instruction-tuned chat model.

License

This release uses CC-BY-SA-4.0 as the practical downstream posture for the mixed training corpus. The corpus combines FineWeb-HQ/FineWeb2-HQ web data and Wiki40B slices, whose upstream terms and attribution/share-alike obligations may apply to downstream use and redistribution. Users are responsible for checking that their intended use and derivative packaging comply with the upstream dataset terms.

Downloads last month
67
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train nazdef/20260703_continual-pretraining-gpt2medium-step34000-nodecay