File size: 772 Bytes
7e8f6cf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
---
language: en
tags:
  - loracle
  - lora-oracle
  - mechinterp
  - llama-3-3-70b
  - drgrpo
license: mit
base_model: meta-llama/Llama-3.3-70B-Instruct
---

# llamacle_drgrpo_v1_step10 — DrGRPO RL on top of llamacle_v6_clean

Continuation of `ceselder/llamacle_v6_clean_step1875` via online Dr. GRPO RL
on the 2,500 held-out FineWeb LoRAs from v6 pretrain. 32 prompts/cycle x K=16
rollouts, lr=7e-6, eps_low/high=0.2/0.28, NF4-quantized base + DDP across 6 B200s,
sub-batched K=4x4 decode, forward_ckpt_inject for backward.

This is step 10 of 80 (early checkpoint).

## Score progression (judge 1-10, mean over kept rollouts)
cycle 1: 4.08
cycle 2: 3.85
cycle 3: 4.45
cycle 4: 4.95
cycle 5: 4.95
cycle 6: 5.17
cycle 7: 5.77
cycle 8: 4.21
cycle 9: 4.82
cycle 10: 4.73