A new benchmark finds policy gradients beat game-theory algorithms
A benchmark presented at ICLR by MIT researchers and collaborators finds that policy gradient methods, a general-purpose family of reinforcement learning tools that date to the early 1990s, can outperform specialized game-theory algorithms in two-player zero-sum imperfect-information games. The result, the source reports, is less