Chess GPT strong-winner phase-MoE policy 0005

This browser-native tournament candidate retrains the 0003 phase-MoE architecture under an exactly matched compute budget. It uses only positions where the eventual winner is to move, from decisive games whose winner has ELO at least 1600. The loser's rating is unrestricted.

Load the public repository in the ChessGPT arena as:

peterwooden/chess-gpt-phase-moe-strong-winner-0005@6bae33a48f207aa1519bbca620094f77ace61dfb

The tournament runner reads browser/manifest.json. The package reconstructs the board from SAN history, deterministically selects its visible-board phase output, scores 4,272 move identities, masks to the supplied legal SAN moves, and returns deterministic argmax SAN.

Architecture, data, and compute

  • Six-layer Transformer encoder, width 336, eight heads, three phase experts, 12,397,296 parameters.
  • Frozen January 2026 Lichess standard-rated games only for optimization.
  • 210,285 games scanned, 120,000 accepted, 90,285 filtered, zero invalid.
  • 4,149,869 available winner-side positions; exactly 3,358,828 processed.
  • Seed 20260729, AdamW, batch size 128, float32 Apple M4 MPS.
  • 11,074,728,541,459,968 ratified training FLOPs, exactly matching model 0003 and charging all evaluated expert branches.
  • No pretrained weights, engine labels, synthetic data, or parent checkpoint.

On 174,404 separately held-out filtered April positions, validation loss was 2.88798, raw top-1 accuracy was 27.030%, legal-masked top-1 accuracy was 28.944%, and the legal-move rate was 100%.

Playing-strength evidence

In a 100-game paired validation match from 50 frozen unfiltered April openings with colors reversed, model 0005 beat model 0003 by 10 wins to 4 with 86 draws, scoring 53–47. The preregistered prediction was an 80% all-game win rate; the observed win rate was 10%, so the direction was supported but the magnitude was not.

This limited local match is not a precise Elo estimate and does not use the unrevealed official tournament openings. The predicted failure mode—poor play against poor opponents—was not tested.

Integrity

The canonical browser package is 49,841,122 bytes, below the 100,000,000-byte limit. The checkpoint SHA-256 is bc23e98f9423ead630ee2850d2facc5e561035475da754fa11629704b6e596ae. Revision 6bae33a48f207aa1519bbca620094f77ace61dfb was downloaded cleanly, compared byte-for-byte with the local export, loaded with ONNX Runtime Web 1.27.0, and completed 40 legal SAN self-play plies.

The exact experiment specification, metrics, full loss log, and paired-match artifact are included under training/ and evaluation/.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support