95% CI: 2,190–2,601
R2 v2 · epoch 14 · 4,000 simulations
Calibrated Stockfish at UCI 1,800
world-models / a chess learning study
Stockfish supplies the examples. An AlphaZero-style network learns from them. Tree search turns its predictions into moves.
95% CI: 2,190–2,601
R2 v2 · epoch 14 · 4,000 simulations
Calibrated Stockfish at UCI 1,800
95% CI: 2,084–2,235
R2 v2 · epoch 4 · 8,000 simulations
Calibrated Stockfish at UCI 2,000
The peak interval spans 2,300. These are engine-relative estimates under the stated evaluation settings, not a tournament rating or a head-to-head result against AlphaZero.
Evaluation details →01 / the method
Stockfish labels board positions with candidate moves and scores. Recorded game outcomes supply value targets.
A residual network learns a move distribution and a value estimate through supervised training.
At inference, Monte Carlo Tree Search uses those predictions to explore moves and choose an action.
The main 20 × 256 network has about 24 million parameters. The project also tests a 40-block variant, self-play follow-ons and learned-model search.
02 / the experiments
On the baseline checkpoint, increasing inference-time search from 800 to 4,000 simulations raised the recorded Elo estimate.
Search ablation →The published data-scale comparison increased training positions from 5 million to 30 million.
Data ablation →Doubling depth at the 5-million-position setting did not improve the reported headline estimate.
Capacity ablation →These summaries belong to different experiments. Each linked write-up records its setup, measurements and limitations.
03 / keep exploring
Related network architecture, different training signal and compute budget.
What happened when self-play followed supervised distillation.
Compare search with a learned model to search with known game rules.
A second game, with KataGo as the teacher and a smaller network.
Failed hypotheses, implementation bugs and what changed afterward.
Dated experiments, follow-up evaluations and open questions.