Tic-tac-toe through self-play
Loading the trained model and exact minimax comparison…
Loading model…Perfect-play value for X: draw
Network vs. perfect play
The bot’s last decision will appear here after it moves.
Network scores are learned estimates between −1 and +1. Exact values come from exhaustively solving the remaining game: +1 win, 0 draw, −1 loss under perfect play.
What is running here?
The model was trained through 2,500 updates on batches of 2,048 self-play games. Its three dense layers were exported once and this page reproduces their matrix operations in your browser. No Python server or retraining is needed when you play.
Built while following Joe Antognini’s JAX tic-tac-toe tutorial ↗.