world-models / a chess learning study

Learning chess
from a strong teacher.

Stockfish supplies the examples. An AlphaZero-style network learns from them. Tree search turns its predictions into moves.

Highest point estimate2,301 Elo

95% CI: 2,190–2,601

R2 v2 · epoch 14 · 4,000 simulations
Calibrated Stockfish at UCI 1,800

Tighter interval, separate evaluation2,153 Elo

95% CI: 2,084–2,235

R2 v2 · epoch 4 · 8,000 simulations
Calibrated Stockfish at UCI 2,000

How to read the result

A measured estimate.

The peak interval spans 2,300. These are engine-relative estimates under the stated evaluation settings, not a tournament rating or a head-to-head result against AlphaZero.

Evaluation details →

01 / the method

From examples to decisions

Full method →
01

Collect

Stockfish labels board positions with candidate moves and scores. Recorded game outcomes supply value targets.

02

Distill

A residual network learns a move distribution and a value estimate through supervised training.

03

Search

At inference, Monte Carlo Tree Search uses those predictions to explore moves and choose an action.

input19 × 8 × 8conv 3×3256 chres block× 20policy head4,672 logitsvalue headtanh → [-1,+1]~ 23.7M paramssame as AZpolicy + valuejointly trainedcross-entropy+ MSE loss

The main 20 × 256 network has about 24 million parameters. The project also tests a 40-block variant, self-play follow-ons and learned-model search.

02 / the experiments

What changed the results?

All experiments →
Search budget

1,807 → 2,084

On the baseline checkpoint, increasing inference-time search from 800 to 4,000 simulations raised the recorded Elo estimate.

Search ablation →
Training data

1,810 → 2,009

The published data-scale comparison increased training positions from 5 million to 30 million.

Data ablation →
Network capacity

20 → 40 blocks

Doubling depth at the 5-million-position setting did not improve the reported headline estimate.

Capacity ablation →

These summaries belong to different experiments. Each linked write-up records its setup, measurements and limitations.