saved
Scalable decision-making for games of imperfect information
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Samuel Sokota, Gabriele Farina and coauthors report Ataraxos, a Stratego AI that beat the strongest human 15–1 with four draws. Ataraxos pairs self-play reinforcement learning with a belief model of hidden pieces and a test-time search step that mimics one more policy update. Trained for a few thousand dollars, it also reached superhuman Barrage Stratego and new state-of-the-art results on Hanabi and dou dizhu.
ideas
- Stratego had stayed beyond multimillion-dollar AI efforts. Prior public-information transformations scale with hidden information, and Stratego has more than 10^33 possible piece layouts, so human experts still led before Ataraxos.
- The design pattern is policy–value, belief, and one search update. Self-play trains a policy–value network; a belief network models opponent piece types from self-play; test-time search samples beliefs, rolls out candidates, then applies one more damped policy update.
- Dynamic damping stabilizes imperfect-information learning. Stronger regularization and larger updates early, then weaker regularization and smaller updates later, are meant to avoid chaotic self-play dynamics without heavy reweighting.
- Ataraxos beat Pim Niemeijer 15–1–4 on modest compute. The 20-game series gave an 85% effective win rate while training cost a few thousand dollars, versus prior work that spent millions and did not reach top humans.
- The same pattern transferred across game types. The authors report superhuman Barrage Stratego against three multi-time world champions, new Hanabi state of the art for two- to five-player variants, and wins over PerfectDou and DouZero at dou dizhu.
quotes
“the best Stratego player ever”
“Ataraxos achieved this result while costing only a few thousand dollars to train.”
“reinforcement learning and search are no longer precluded from high performance by the presence of large amounts of hidden information.”
“consumed roughly 1/500th of the compute cost, 1/30th of the self-play games and 1/100th of the training examples.”