{ "checkpoint": "gemma-27b_layer58_maghead_mlp_h2048x2_t0.05_a2.0.pt", "arch": "maghead_mlp (MLP dir head 5376->2048->2048->1024 GELU + linear mag head)", "base_model": "gemma-3-27b", "data": "nickypro/fineweb-gemma27b-residuals (16 shards)", "layer": 58, "d_res": 5376, "data_shards": 16, "val_rank1": 0.658, "val_dir_rank1": 0.676, "val_cos": 0.43, "val_mag_mse": 0.7, "decode_cos_pred_sbert": 0.343, "decode_cos_upper_sbert": 0.801, "decode_cos_sonar_shortcut": 0.431, "date": "2026-06-06", "note": "RETRIEVAL SOTA across all parascope probes: rank1 0.658 / dir_rank1 0.676 (> llama-3b champion 0.624/0.638) with only 16 shards. Bigger base model -> better DIRECTION prediction. Magnitude head is the wall (mag_mse 0.70); decode ~0.43 of this corpus's roundtrip ceiling 0.80 (same fraction as llama-3b's 0.42). Different corpus (fineweb) than the llama-3b probes, so absolute decode not directly comparable." }