Contents

What board games can tell us about frontier labs' posttraining.

Early Motivations for Boardgames & AI

A good deal of AI investment has gone into AI for math and coding. At this point there’re so many coding tasks in the distribution of what AI has seen and what it can do that it can feel hard to find tasks that the best LLMs can’t do. Of course, companies internally still have unique problems that you won’t find people solving on StackOverflow, but the shape of such proprietary knowledge isn’t openly published on the internet. “Even the most recent models, including Fable and GPT-5.6, do not produce systems with design patterns that are commonplace at HFTs.”

I was wondering earlier this year if there were perhaps obvious things that AI wasn’t being heavily trained on. One thought that came to mind was board games (aka tabletop games). Key developments in AI history have centered around such games: Chess with DeepBlue, and Go with AlphaGo. I’ve always liked playing tabletop games as they can exercise your planning, reasoning, and decision-making skills, while also being a great way to socialize.

I vibe-coded the early workings of a project in which I could play some simple games sitting at a virtual table of LLMs. It was both fun and quite instructive for seeing how LLMs themselves fit into the agents that were playing the games. Then as I read more about RLHF and posttraining, I wondered if such games could be used for evaluation. Doing a bit of searching, I saw preexisting work of people testing LLMs on Chess, Poker, and in one instance, Risk. However, these directions didn’t interest me – for one they just weren’t games that I was interested in playing myself. Also, I didn’t want to feel like I was just playing a chess bot or poker bot (both dry experiences I’ve tried before). With a proper chat interface, I could have potentially create a rich environment with socialization, coordination, and betrayal between players.

Anthropic’s Secret Evals?

I had a few games in mind, one of which was Catan. While diving into it a bit more, I ran across an interesting tidbit in one of Anthropic’s videos on context management where they demonstrated their work by having Claude play Catan. Separately, it was observed that Claude Mythos used “illegible reasoning” which people later realized was just a shorthand notation for a card game like FreeCell or Solitaire. Anthropic even mentions in the Fable 5 release that the model outperformed Opus 4.8 in Slay the Spire, which despite being a videogame plays out more like a board game. (Slay the Spire is a deckbuilding game involving turns of discrete choices. It doesn’t test your reflexes or hand-eye coordination in any way.) The space of possible games is immensely rich, and each one could derive insights for how an AI model explores a new problem. Clearly I wasn’t the only one who believed in the interestingness of using boardgames as evals.

So seeing that Anthropic likely has quite a few evals based on board games, I felt quite validated in exploring this idea further. Is this something other frontier labs are doing? What about the open-source ones? Would general capability correlate well with performance across board games? Would the results clearly show out-of-distribution performance for some models but not others? At this point I was excited to vibe-code a dozen implementations of my favorite board games, but figured it was best to start small.

The Games

The full repo has several functioning games, but for AI evaluation specifically I chose these two for simplicity and interestingness.

Game 1: Coup

Coup is a fairly straightforward game involving strategy and bluffing.

There was an online implementation of it that I’d play with my friends when were social distancing during Covid. To play well in Coup you need to take calculated risks, bluffing or calling bluffs at critical moments. You also need to plan ahead. Because some characters counter others in a rock-paper-scissors dynamic, you can aim to create a scenario in the distant future where your character naturally wins, even if doing so weakens you in the short-term.

Game 2: One Night Ultimate Werewolf (ONUW)

You might’ve heard of Werewolf, sometimes played as Mafia. It’s a social deduction game were you have villagers (good guys) and secret werewolves (bad guys). Over a sequence of days and nights the werewolves can kill villagers and the group has to vote players out, trying to eliminate the wolves.

ONUW is simillar in idea but the entire game is condensed into one night, with one round of voting. The village team wins if they vote out a wolf, while the wolves win if they can sneak past undetected. There are special roles that can switch role cards around at night, so it’s possible for a wolf to suddently join the village team, or vice-versa, sometimes without knowing. You win or lose based on your final role. These card-switching actions can be deduced or bluffed. For example, you can claim that you switched someone’s card, incentivizing them to tell you if they were a wolf. (Since now the wolf card would be somewhere else.) But if you lied then they’d still be a wolf, and you can vote them out.

Methodology, Abridged

Most of the code quite was straightforward to handle with coding agents.

There were a few hiccups that required re-thinking the design, leading to extended conversations with the agent.

Running Games

I hosted 70 games of Coup and 250 games of ONUW, testing Anthropic’s Claude Sonnet 5, OpenAI’s GPT-5.6 Luna, DeepSeek V4 Flash, DeepSeek V4 Pro, and Z.ai’s GLM 5.2. Costs limited me to a small set of cheap models as each game could consume hundreds of thousands of tokens. I wanted to run enough games so that after calculating standard errors there would be clear separation between models’ MMRs. If there’s more interest in this I’ll gladly pay for more experiments. Or I could make the code public so others can play with it.

To improve information-gain per game, Coup games were 3-player with rating adjusted by placement. (More players means a longer game, but also gives more information through the final ranking, both scaling linearly with number of players. However, needing to feed the game’s tokens to more agents grows our cost at a rate beyond linear.) ONUW games were 5-player, with separate MMRs for whether you began as a village-aligned role or a werewolf. Each ONUW game featured only 2 different LLMs: I would start each village player with LLM A, and each werewolf with LLM B, allowing for updates of LLM A’s villager rating against LLM B’s werewolf rating. (This isn’t entirely true, as there are cases where a wolf card gets swapped, causing one agent to beat another agent running the same LLM. In these cases the LLM’s rating updates less as it’s winning against itself.) This design was chosen to minimize our standard errors, deriving more precise rating estimates.

Results

Ratings

Despite leading by a wide margin in Coup, Sonnet 5 is a mediocre wolf and the weakest villager in ONUW. GLM-5.2 is a standout in ONUW, being the strongest wolf among all models and nearly the strongest villager (about tied with DeepSeek-v4-pro). GPT-5.6 Luna was the worst Coup player, ONUW wolf, and a mediocre villager.

Note that the DeepSeek-v4 models tested here were before the release of 0731 and 0813 checkpoints for flash and pro, respectively. These improved posttraining checkpoints were released after I had gathered most of my data.

Interesting Obervations

There were some interesting patterns from my perspective as someone who’s spent hours playing each of these games. They’re written assuming some familiarity with their rules and strategy.

Coup

Calling Bluffs

Influence Cards

Sonnet performed the best here by a wide margin of nearly 200 rating points, mainly because it played safely and rationally. It rarely bluffed roles when taking actions (2 out of 95 actions), causing other models to often fail challenges against it. Additionally, it was the only model that won more of its aggressive challenges than it lost. Other models tended to challenge overaggressively, hurting themselves more than their opponents. GPT-5.6 Luna in particular reasoned itself into highly risky challenges that backfired: “1/3 chance to succeed” -> “challenging early in Coup can be strategic” -> challenge anyway.

ONUW

Village

The biggest discrepancy here is Sonnet’s poor village performance despite being strong in other aspects. Sonnet is average in terms of identifying a werewolf, but a group of Sonnet villagers is the most likely to participate in groupthink, aggreeing with the popular concensus 68% of the time even when it’s wrong. Looking at games, Sonnet is gullible as long as the picture is logically sound. Against Sonnet, a wolf can wait until the villagers have mostly claimed, then claim a role that has yet to be claimed, and Sonnet is happy to believe everyone is telling the truth just because it’s not obvious that anyone is lying. GLM-5.2 is the best at forming a consensus against a wolf, so its tendency to follow the crowd is less of a problem. Deepseek-v4-pro is the best at (correctly) dissenting against the public opinion.

One play that’s considered “good” is to look at two center cards as seer instead of another player’s card. This tends to give the village team more information with which to catch a werewolf’s lie. GLM-5.2 and deepseek chose to look at the center the overwhelming majority of the time, while Sonnet 5 and GPT-5.6 Luna preferred to peak at a player card.

Werewolf

I was pretty dissappointed to see that models occasionally made obvious blunders as werewolf, like publicly saying they were a werewolf (unprompted, even without someone else claiming a swap), or by voting their own werewolf partner. When it comes to playing wolf and deceiving the village, GLM-5.2 performs best by creating valid alibis and not slipping up. Though most wolves try to slip under the radar by claiming villager, there was a game were a GLM wolf claimed seer, stating that it saw two villager cards in the middle. This caused the villagers (Sonnet 5) to accuse each other, handing GLM an easy win.

Broader Takeaways

My “vibes” on each model after running these experiments (and trying them out):

Coup

Claude Sonnet’s exceptional performance on Coup could be a reflection of Anthropic’s internal game evals encouraging long-horizon strategic decision-making that eventually distills its way into their cheapest models. Other models, both closed and open, produce coherent reasoning but still perform losing actions in game. The clear weaknesses of open-weights models can be explained by smaller labs having less resources. They likely spend a much greater percentage of their compute and manpower on high-priority tasks like agentic coding, yielding models with less breadth. Given this, I was surprised by GPT-5.6-luna’s poor performance. Perhaps OpenAI does not have as as rich of an evaluation suite when it comes to playing games. Or maybe they do but their distillation process focuses too much on other benchmarks for their smallest model to perform well on game-like decision making. Or it could be that Coup is specifically in the distribution of Anthropic’s training but OOD for the other models tested.

ONUW

I doubt ONUW is in the posttraining distribution of any of these models. All models except for Sonnet exhibited some degree of obvious blunders/misalignment. (I would classify a wolf reporting their fellow wolf as misalignment, as it’s clearly acting against the team’s interests.) In this way, Sonnet’s consistency may be a sign of Anthropic’s focus on alignment. However, Sonnet loses as a villager when it fails to account for the possibility of being lied to in a multiagent system.

← back home