hraness
Theme
Appearance

saved

Scalable decision-making for games of imperfect information

by Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter and Gabriele FarinaNaturepublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Samuel Sokota, Gabriele Farina and coauthors report Ataraxos, a Stratego AI that beat the strongest human 15–1 with four draws. Ataraxos pairs self-play reinforcement learning with a belief model of hidden pieces and a test-time search step that mimics one more policy update. Trained for a few thousand dollars, it also reached superhuman Barrage Stratego and new state-of-the-art results on Hanabi and dou dizhu.

ideas

  • Stratego had stayed beyond multimillion-dollar AI efforts. Prior public-information transformations scale with hidden information, and Stratego has more than 10^33 possible piece layouts, so human experts still led before Ataraxos.
  • The design pattern is policy–value, belief, and one search update. Self-play trains a policy–value network; a belief network models opponent piece types from self-play; test-time search samples beliefs, rolls out candidates, then applies one more damped policy update.
  • Dynamic damping stabilizes imperfect-information learning. Stronger regularization and larger updates early, then weaker regularization and smaller updates later, are meant to avoid chaotic self-play dynamics without heavy reweighting.
  • Ataraxos beat Pim Niemeijer 15–1–4 on modest compute. The 20-game series gave an 85% effective win rate while training cost a few thousand dollars, versus prior work that spent millions and did not reach top humans.
  • The same pattern transferred across game types. The authors report superhuman Barrage Stratego against three multi-time world champions, new Hanabi state of the art for two- to five-player variants, and wins over PerfectDou and DouZero at dou dizhu.

quotes

“the best Stratego player ever”

George Franka, quoted by Sokota et al.

“Ataraxos achieved this result while costing only a few thousand dollars to train.”

Samuel Sokota, Gabriele Farina, and coauthors

“reinforcement learning and search are no longer precluded from high performance by the presence of large amounts of hidden information.”

Samuel Sokota, Gabriele Farina, and coauthors

“consumed roughly 1/500th of the compute cost, 1/30th of the self-play games and 1/100th of the training examples.”

Samuel Sokota, Gabriele Farina, and coauthors