I made 6 AI models play poker against each other. The 1.2B model has a gambling problem and it keeps winning.
Summary
An experiment where six AI models played Texas Hold'em against each other, with a tiny 1.2B model winning twice by being too reckless to fold. A community tournament is being organized, inviting participants to submit model personas and formats.
Similar Articles
I Made LLMs Play Texas Hold’em. The Smallest Model Beat a ~1T Model by Being Too Dumb to Fold
An experiment where six LLMs played Texas Hold'em poker; a tiny 1.2B model won twice due to its aggressive 'never fold' strategy, highlighting how format can favor simpler models. The author built a poker engine and agent framework called Hive, and invites community feedback.
I gave the same AI 6 different personalities and made them play poker 100 times.
An experiment giving the same 1.2B language model six different personalities and playing 100 poker tournaments reveals drastic behavioral differences: a 'Grinder' never wins but never loses, a 'Tilter' wins big or busts, and a 'Shark' dominates. The results highlight how personality prompts can profoundly shape LLM decision-making.
I gave 6 AI models a challenge they could only win with a partner. They found their own allies, cut deals in private, and faced off as three rival teams — including two that only paired up because no one else would have them.
Six AI models were tasked with forming alliances to win a funding proposal challenge. They independently negotiated partnerships and created three rival teams, demonstrating autonomous coordination and strategic negotiation.
I think poker is an underrated benchmark for AI agents
The author argues that poker is an underrated benchmark for AI agents because it tests reasoning under uncertainty, adaptation, and risk management, and describes an upcoming AI poker arena where builders can submit bots to compete.
I built a site that lets you watch, wager, and prompt inject agents playing games
A developer built a site where users can watch AI agents play games, wager fake coins, and use winnings to prompt inject agents. The author shares observations about model performance, noting that smaller models struggle while Qwen3 235B excels.