Slay the Spire

17-minute read

Long-horizon tasks are all the rage these days. Many new model releases come with claims about how many hours of autonomous work a model can do and how many steps it can take without losing the thread.

To see how good current models actually are at strategizing over many steps, we let them play Slay the Spire 1, a roguelike deck building game. One run of the game consists of hundreds of consecutive decisions, with randomized events sprinkled in that require the player to adapt on the fly.

We ran 9 frontier models, each at 3 reasoning effort levels, on 5 game scenarios, 10 runs per combination. Skip straight to the results if you are only here for the numbers and pretty graphs. The full report with interactive transcripts and recordings of the runs is available at ai-plays-slay-the-spire.com.

Why Slay the Spire

Slay the Spire is a (frighteningly addictive) roguelike deck-builder released by Mega Crit in 2019. The player picks one of four characters and climbs a spire in three acts, 50 floors in total. Each act consists of 15 floors with multiple rooms each and with a boss on the top floor. Each room is one of: a regular fight, an elite fight, a shop, a rest site, a treasure room, or an unknown room that is usually some special event. On every floor, the player chooses which connected room on the next floor they want to enter. A short fourth act ends at the Corrupt Heart, the game's true final boss. It is only open to players who collected three keys on the way up. Killing the Heart is worth a lot of points, so a player who wants the top score has to plan for those keys from Act 1 on.

The full Act 1 map of the baseline scenario: three starting fights at the bottom, branching paths through fights, events, shops, rest sites and a treasure row, and the Slime Boss at the top.
The Act 1 map of the baseline scenario used in the benchmark. The run starts at one of the three fights at the bottom and ends at the Slime Boss on top.

Fights are turn-based card battles. Enemies announce what they will do in their next turn, so every combat turn is a little tactical puzzle in itself. The player starts with a small deck of basic cards and builds a stronger deck over time by finding, buying, removing, or upgrading cards.

The first fight of the baseline run: the Ironclad faces two slimes, each showing the damage it intends to deal next turn, with five cards in hand and three energy.
The first fight of that run. Each enemy shows what it will do on its next turn; the player has three energy and five cards in hand.

What makes Slay the Spire interesting as a benchmark for long-horizon planning is that it requires careful decisions at three different time scales:

  • Short term: Making good tactical decisions during combat is crucial to preserve health points and stay alive to keep the run going.
  • Medium term: The player can see the full map with all possible paths for the current act. Planning a path that maximizes returns while not being too risky is crucial.
  • Long term: Deck building spans the whole run. Picking the right cards early on can unlock strong synergies later in the run or make future acts and boss fights easier.

Runs are seeded, so the same actions give the same result, but the player does not know the seed and cannot predict what they will draw. The seed determines card draws, battle rewards, and event outcomes. That means the player has to form a plan, see how it works out in practice, and adapt their strategy as the run progresses.

The game also offers a lot of variety without any changes to the benchmark harness. The four characters play differently, different seeds give different maps, and the game includes some built-in rule modifiers that change the game's mechanics.

Last but not least, every run results in a single numerical score, which makes it easy to compare runs. Reaching a floor, killing enemies, beating elites and bosses, beating them without taking damage, and various end-of-run bonuses for deck size, relics, gold, and max HP all add points. The large number of score modifiers means that two runs that reached the same floor vary in score based on how the run was played, which adds a lot of strategic depth to the game.

Building the harness

The model plays the real game (V2.3.4), running headless inside a container. The game itself has no API, but the game's community has built an awesome modding ecosystem. ModTheSpire loads mods into the game, BaseMod provides the hooks that most mods depend on, and MCPTheSpire exposes the game state and the available actions via an MCP server.

We run ModTheSpire and BaseMod as unmodified upstream builds but we added over 50 patches on top of MCPTheSpire to make playing the game as smooth as possible from a model's perspective. A small supervisor process sits between the model and the mod and owns the run. It starts the run with the configured seed, character and ascension (difficulty) level, and it filters out some commands that the mod would otherwise expose for saving, loading, and abandoning runs. Every tool call the model makes goes through the supervisor, which records the command and the resulting game state, so the full rollout is logged. When the run ends, the supervisor writes the outcome and rejects any further tool calls. Act 4 is normally locked until the player has won with three different characters, and the game disables it for runs with custom modifiers. We patched that gate out so every run can go for the Heart.

The model only gets tools to interact with the game, no shell. The tools are what a human player would be able to see and do on screen: read the current screen, show the act map, play a card, end the turn, select rewards, use or discard a potion, proceed, confirm, skip, and cancel. A batch tool lets the model execute a sequence of multiple actions in a single call. Every action returns the new screen state and a short summary of what happened, for example, what actions the enemies executed and how much damage the model took.

container: 4 CPUs, 16 GB RAM Model tool calls state Game tools play_card, choose, end_turn, proceed, execute_actions, … lookup_game_info Knowledge base cards, relics, monsters, events, … MCP Supervisor starts the run, blocks save, load and abandon, records every step Run record rollout, per-floor snapshots, outcome MCP Game JVM Slay the Spire V2.3.4 MCPTheSpire MCP server, 50+ patches ModTheSpire, BaseMod mod loader and hooks, unmodified upstream software OpenGL, 1280 × 720 renders to X display Xorg dummy ffmpeg video
The model only sees the game tools, no shell. Everything but lookups is forwarded to the supervisor, which owns the run and records it. The supervisor talks to the patched MCPTheSpire mod inside the game over MCP. The game renders to a dummy X display, which ffmpeg records as a video.

A run has a wall clock limit of 8 hours and a turn limit of 2,000, where a turn is one model response. No run came close to either: the longest run took 4.5 hours and 961 turns. All tested models have context windows of at least 500K tokens and there is no automatic compaction.

We test for capabilities, not game knowledge

We want the benchmark to measure models' long-horizon capabilities instead of memorized game knowledge, which means that a capable model that should be able to do well in the benchmark when given a description of the game and the scoring formula. As a consequence, the model must be able to look up all the game mechanics during a run, the equivalent of a human player reading up on the game in an online wiki as they play. So we added a lookup tool that provides information about any card, relic, potion, buff, enemy, etc. by name, extracted directly from the game to ensure correctness. Additionally, the instructions explain the game's rules and spell out exactly how the score gets calculated.

Model ergonomics

We want to reduce the friction that models encounter when interacting with the game so that they can focus on the actual task of playing the game. We therefore went through multiple iterations of letting models play the game and then looking at run logs for recurring tool calling errors or soft-locked runs. When several models repeatedly made the same mistakes, we treated that as an environment issue and fixed it by improving the tool interfaces.

The benchmark

We ran the benchmark across 9 models, picked from the top of the Artificial Analysis intelligence index on September 9, 2026. Grok 4.7 came out while the benchmark was running and replaced Grok 4.6. Each model was run at the minimum and maximum reasoning effort available, as well as a medium effort level. The table below shows the models, the providers we used, and what the effort levels map to in the providers' APIs.

Model Provider Min Medium Max
GPT-6 Astra OpenAI low medium max
GPT-5.6 Sol OpenAI low medium max
Fable 5.1 Anthropic low medium max
Muse Spark 1.3 Meta minimal medium max
GLM 5.3 Together AI low high max
Grok 4.7 xAI low medium xhigh
Kimi K3 Together AI low high max
Gemini 3.8 Flash Google Vertex AI low medium high
Qwen 3.8 2.4T A95B Together AI low medium xhigh

We searched over 100M seeds to find scenarios that exercise very different aspects of the game. For each seed we generated the map of the first act and counted the number of monsters, elites, events, shops, rest sites, forks on the map, and the number of possible routes to the first boss. The scenarios we ultimately used for the benchmark are as follows:

  1. Baseline: Play as the Ironclad, seed 8EN5Q, ascension 0. This scenario has the exact population-median Act 1 map on every metric (26 monsters, 4 elites, 12 events, 3 shops, 10 rests, 13 forks, 204 routes), Slime Boss as the Act 1 boss. It represents a typical run of the game.

  2. Cursed: Play as the Silent, same seed as baseline, ascension 0, "Cursed" modifier. The modifier forces the player to obtain curses (cards with negative effects) but at the same time buffs them for every curse they have. Trading off the pros and cons of being cursed adds an additional layer of complexity to this scenario.

  3. Easy trap: Play as the Defect, seed Q776W, ascension 0, Guardian as the Act 1 boss. This scenario lets the player trade difficulty for higher scores. They can choose a path that only requires a single combat before reaching the boss, allowing them to climb floors quickly. However, they can also choose to take a much harder path with up to 5 elite encounters, which can yield a higher score.

  4. Dense map: Play as the Watcher, seed 286QJD, ascension 0, Guardian as the Act 1 boss. The map has 25 forks (almost twice as many as the baseline) and 4,040 possible routes through Act 1. It is one of the most decision-dense scenarios we could find that does not force any elite fights, giving the player a lot of opportunity for strategizing. The Watcher is a strong character if played well, but its stance mechanics take effort to manage.

  5. Forced elites: Play as the Ironclad, seed C5W7N, ascension 1, Hexaghost as the Act 1 boss. This scenario is the hardest of the five on paper because ascension level 1 increases the chances of encountering elites. We took this to the limit by finding a map that forces at least 3 elite encounters in Act 1, with no path around them. Additionally, the Act 1 boss is Hexaghost, which is considered the hardest of the three Act 1 bosses.

The benchmark grid is 9 models x 3 reasoning effort levels x 5 scenarios x 10 runs = 1,350 runs.

If you want to try the scenarios yourself, start a run with the character, seed and ascension level given above. For the Cursed scenario, pick the Cursed Run modifier in Custom mode. To reach the Heart your profile needs Act 4 unlocked, which requires a previous win with each of the Ironclad, the Silent and the Defect. The Cursed scenario cannot reach Act 4 at all without a custom patch.

Results

The results below cover all 1,350 runs. Grok 4.7 has a 500K token context window and 5 of its deepest runs reach the limit in Act 3 or 4. We scored those runs where they stopped, so Grok's numbers are biased low. The whole benchmark cost about $38,000 in inference.

Game score by model and reasoning effort, all five scenarios pooled. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Game score by model and reasoning effort, all five scenarios pooled. Boxes show the median and quartiles, one mark per run, triangles are won runs. Models are sorted by their median score at max effort.
Baseline: Ironclad, ascension 0, seed 8EN5Q. The population-median Act 1 map, Slime Boss as the Act 1 boss. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Baseline: Ironclad, ascension 0, seed 8EN5Q. The population-median Act 1 map, Slime Boss as the Act 1 boss. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Cursed: Silent, ascension 0, the baseline map with the Cursed Run modifier. The player collects curses and starts with relics that reward them. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Cursed: Silent, ascension 0, the baseline map with the Cursed Run modifier. The player collects curses and starts with relics that reward them. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Easy trap: Defect, ascension 0, seed Q776W. A lane with a single fight and no elites against a lane with up to five elites, The Guardian as the Act 1 boss. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Easy trap: Defect, ascension 0, seed Q776W. A lane with a single fight and no elites against a lane with up to five elites, The Guardian as the Act 1 boss. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Dense map: Watcher, ascension 0, seed 286QJD. 25 forks and 4,040 routes through Act 1, The Guardian as the Act 1 boss. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Dense map: Watcher, ascension 0, seed 286QJD. 25 forks and 4,040 routes through Act 1, The Guardian as the Act 1 boss. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Forced elites: Ironclad, ascension 1, seed C5W7N. Every path goes through at least three elites, Hexaghost as the Act 1 boss. Boxes show the median and quartiles, one mark per run, triangles are won runs.
Forced elites: Ironclad, ascension 1, seed C5W7N. Every path goes through at least three elites, Hexaghost as the Act 1 boss. Boxes show the median and quartiles, one mark per run, triangles are won runs.

The score is the game's own run score. Reaching the Act 1 boss is worth about 100 points, a win without the Heart about 700, and a win with the Heart 1,500 or more. Scores are pooled over the five scenarios; the tabs show each scenario on its own.

Long story short: GPT-6 Astra is in a league of its own. It won 79% of its runs, never died before Act 3, and 27 of its 31 deaths were to the Heart, the final boss. The next best model, Gemini 3.8 Flash, won 27% of its runs, and Fable 5.1 26%. Every other model won 7% or less, and the bottom four lose most of their runs in Act 1.

Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort, all five scenarios pooled.
Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort, all five scenarios pooled.
Baseline: Ironclad, ascension 0, seed 8EN5Q. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Baseline: Ironclad, ascension 0, seed 8EN5Q. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Cursed: Silent, ascension 0, the baseline map with the Cursed Run modifier. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Cursed: Silent, ascension 0, the baseline map with the Cursed Run modifier. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Easy trap: Defect, ascension 0, seed Q776W. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Easy trap: Defect, ascension 0, seed Q776W. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Dense map: Watcher, ascension 0, seed 286QJD. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Dense map: Watcher, ascension 0, seed 286QJD. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Forced elites: Ironclad, ascension 1, seed C5W7N. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.
Forced elites: Ironclad, ascension 1, seed C5W7N. Mean game score against mean inference cost per run, log scale. One line per model through min, medium and max effort, marker shade darkening with effort.

How models play

Weak models die early. The strong models only lose to the Heart while the weak ones mostly already die in Act 1. Qwen died in Act 1 in 117 of its 150 runs, Muse Spark in 92, Kimi in 79 and GLM in 67. These models die to ordinary fights and the Act 1 elites as much as to the bosses. They also die rich, with 30 to 40% of the gold they earned unspent. They rest (heal) at campfires instead of upgrading cards despite being healthy, and Muse Spark died holding an unused potion in 64% of its runs. Fable almost never dies in Act 1, but loses more runs in Act 2 than any other model. Gemini dies everywhere and has by far the widest spread of any model: the same scenario at the same effort can score 100 or 1,700.

Nobody can play the Watcher. The dense map is the worst scenario for every model, but the map is not the reason: its routes have fewer elites than the baseline and models used 66 different paths in Act 1. The models dies because they don't understand how to use the Watcher's stances. Her core mechanic is Wrath, a stance that doubles both the damage she deals and the damage she takes. The trick is to enter it for the attack and leaving it again before the enemy acts. The models unfortunately failed to pull off the second part of the trick. 40% of all Watcher deaths happened with Wrath active. In most of those turns the model had entered Wrath that same turn and had no card in hand to leave it. Even Astra ends 19% of its attack turns in Wrath, which is a big part of why it only wins 27% of its Watcher runs, despite being the strongest model overall.

Does more reasoning help?

Increasing the reasoning effort mostly helps from minimal effort to medium. Gemini is a great example for this. At minimal effort, it plays like, well, a bot. At medium effort it is the second strongest of the bunch.

Pushing the effort from medium to max is where some interesting things happen. Every model except Fable gains on their median score, but the gains are modest and the mean is within noise. Fable goes the other way and loses 230 points on the median, a drop that holds up across scenarios. It thinks about four times longer per move at max and uses that time to play more cautiously outside of combat and more recklessly inside it. It turns down trades that cost max HP but strengthen the deck, and in boss and elite fights it ends more turns without enough block, losing 40% more HP per enemy attack than at medium with the same cards in hand. On the Easy trap scenario, half of its max runs died to the Act 2 boss, where none of its medium runs did.

What it costs

Gemini at medium effort is the clear value pick at 833 points for $12. Above that, nothing but Astra is worth buying. Astra at min effort scores 1,377 for $64. Astra at max costs 67% more for 9% more score. Fable sits at 660 to 790 points for $60 to $86, below Gemini medium at five to seven times the price.

Taking Astra to the limit

First off, it seems evident that Astra has either been trained directly on Slay the Spire or had a lot of related content in its training data. It asks the fewest questions by far, calling the lookup tool about once per run, with 70% of its runs never calling it at all (against 6 lookups per run for Fable and 13 for Muse Spark). When it does look something up, it is mostly rare cards. It also barely reasons: about 35 output tokens per request at every effort level, where Fable produces 200 to 900. It plays like a machine, taking the same path through each map in almost every run, making the fewest tool errors, and losing the least HP of any model in the first fight (which is identical for every run of a scenario). That is a real capability, but it is not the one this benchmark was built to measure. For Astra, the game stops being a long-horizon planning task and becomes a recall task.

However, we also went a bit easy on the models, seeing that many of them already struggled at low ascension levels. If you want to gain a Spire veteran's respect, you gotta show that you can consistently beat the game at the maximum difficulty level: ascension 20.

GPT-6 Astra at max effort, ten runs per scenario at Ascension 20 next to its Ascension 0/1 runs. Ascension 20 stacks all twenty of the game's handicaps and adds 5% to most score sources per level, so the raw scores are inflated; compare wins and spread, not the numbers.
GPT-6 Astra at max effort, ten runs per scenario at Ascension 20 next to its Ascension 0/1 runs. Ascension 20 stacks all twenty of the game's handicaps and adds 5% to most score sources per level, so the raw scores are inflated. Compare wins and spread, not the numbers.

After putting Astra to the test, its win rate drops from 86% to 20%, and the ways in which it dies go from the Heart alone to a dozen different fights. It still reaches at least Act 3 in 64% of its runs, which is still extremely strong. According to the experts on Reddit, this win rate is not absolute elite tier, but far beyond any casual gamer. Despite the increased difficulty, Astra still barely thinks and plays almost entirely from memory.

While it is fascinating to see Astra play the game at such a high level, it also means that Slay the Spire stops being a good measure of its planning capabilities, since it doesn't plan all that much. The obvious next step would be to change the rules of the game so that it can't play from memory. But we gotta leave something for future research.

Full report and videos

Every run of the benchmark, including the Ascension 20 runs, is available on ai-plays-slay-the-spire.com with interactive transcripts and videos.