In 1997, a machine beat a world chess champion, and the story of artificial intelligence and games seemed, for a while, settled. Deep Blue's win over Garry Kasparov was treated as a milestone in computing, and it was. But it was also a slightly misleading one, because chess, for all its depth, is a game where both players see the entire board at all times. There is no hidden information to reason about, only a very large number of moves to search through.
Poker does not offer that comfort. Two players at a table know their own hole cards and nothing more about each other's. The game is built on incomplete information, on bluffing, on the fact that the correct decision in a given spot depends on what an opponent believes you might be holding, which depends on what they think you believe, and so on. For decades, this made poker a much harder problem for computer scientists than chess or, later, Go — and a more interesting one.
That difficulty is exactly why poker became one of the defining testbeds in artificial intelligence research over the past twenty-five years. Long before large language models made machine learning a household topic, researchers were quietly using poker to solve problems that chess could never pose.
The limits of perfect information
Game theorists divide games into two broad families. In games of perfect information — chess, checkers, Go — every player can see the full state of the game at every point. The challenge is computational: search enough of the game tree, evaluate positions well enough, and you converge on strong play. This is what let IBM beat Kasparov and, in 2016, what let DeepMind's AlphaGo beat Lee Sedol, a feat many researchers had expected to take another decade.
Poker sits in the other family, games of imperfect information, alongside negotiation, auctions and most real economic decisions. In these games there is no single best move that exists independent of what your opponent might do or believe. A hand played face-up would be trivial; the entire game rests on what stays hidden. Building a machine that could play well under that condition meant building a machine that could reason about uncertainty and deception, not just calculate outcomes.
Alberta's long game
The University of Alberta's Computer Poker Research Group, formed in the 1990s under Jonathan Schaeffer — the same computer scientist who had built Chinook, the checkers program that would eventually solve that game completely — spent years building poker-playing programs with names like Loki, Poki and eventually Polaris. In 2007 and 2008, Polaris played exhibition matches against professional players in Vancouver, close enough to show that the gap between machine and human play in poker was narrowing fast.
The group's most consequential result came later. In 2015, a program called Cepheus, built by Michael Bowling and colleagues at Alberta, was described in the journal Science as having essentially solved heads-up limit hold'em — meaning its strategy was close enough to game-theoretically optimal that no opponent, however skilled, could expect to beat it by a meaningful margin over a very large number of hands. It remains one of the few games of real strategic depth ever solved to that degree.
Regret, minimized
None of this was possible without a shift in algorithmic thinking. In 2007, researchers including Martin Zinkevich, Michael Johanson and Michael Bowling introduced counterfactual regret minimization, an algorithm — now usually called CFR — for finding approximate equilibrium strategies in large games with hidden information.
The idea, stripped of its mathematics, is straightforward: a program plays itself over and over, and after each hand it asks what it would have done differently in hindsight, at every decision point, given how the hand actually unfolded. It nudges its strategy away from choices it regrets and toward ones that would have performed better. Repeated over millions of hands, this converges toward a strategy that is very hard to exploit. CFR and its later variants became the backbone of almost every serious poker-playing program that followed, and the technique has since been adapted well outside poker.
Libratus, Pluribus and the multiplayer problem
Solving heads-up limit hold'em was a landmark, but limit hold'em, with its fixed bet sizes, is a comparatively contained game. No-limit hold'em, where a player can bet any amount at any time, expands the decision tree enormously. In 2017, a program called Libratus, built by Tuomas Sandholm and his then-PhD student Noam Brown at Carnegie Mellon University, played a twenty-day heads-up no-limit hold'em match against four professional players in Pittsburgh and won decisively enough that the result was not in serious statistical doubt.
Two years later, Sandholm and Brown went further with Pluribus, a program that played six-player no-limit hold'em — a format with far more strategic complexity than heads-up play, since equilibrium concepts that work cleanly for two players do not extend neatly to five opponents. Pluribus played against strong professional players, among them Darren Elias and Chris Ferguson, and its results were again published in Science. What made Pluribus notable to computer scientists was less the win than the resources behind it: it ran on comparatively modest hardware, a sign that the underlying techniques, not brute computational force, were doing the work.
- Robotic and human patrol scheduling for airport and infrastructure security, drawing on game-theoretic models first refined for poker research
- Negotiation and bargaining agents that must act without full knowledge of the other party's constraints
- Auction design and bidding strategy, where hidden valuations mirror hidden hole cards
- Cybersecurity resource allocation, treating attacker and defender as players with incomplete information about each other
- Multi-agent reinforcement learning more broadly, where CFR-descended methods are used well outside any card game
Beyond the table
It is worth being precise about what this research does and does not show. None of it means a machine, or a person, can turn poker into a reliable source of income; the games these programs mastered were narrow, fully specified contests, played under conditions no live cash game or tournament replicates, and the variance inherent to poker does not disappear because an algorithm plays well on average. What the research shows is something narrower and, for computer science, more useful: that a game built on bluffing and hidden information could be formalized well enough for machines to reason about it rigorously.
That turns out to matter for a wide range of problems that have nothing to do with cards. Security agencies scheduling patrols, negotiators structuring an offer, market makers pricing under uncertain information — all of these are, in the language of game theory, imperfect-information games, and the algorithms first stress-tested at a poker table have found their way into all of them.
Poker gave computer science its clearest laboratory for a problem chess never had to face: how to act well when you cannot see everything.
There is a certain irony in a card game associated, fairly or not, with bluster and bravado turning out to be one of the more rigorous proving grounds in modern computer science. But the irony is only surface-level. Poker forces a decision-maker to reason honestly about what they don't know, and that is precisely the kind of problem worth building better machines to think about — whether or not any of us is holding a hand at the table.
If this is your lane, these go deeper: pot odds and equity and reading your opponents. For more poker writing in English and Spanish, follow us on Facebook.