✦ Simulator

Two people, two choices, and no way to talk it over.

Type any payoff matrix and this solves it: best replies, dominant strategies, every Nash equilibrium including the mixed ones, and which outcomes waste value. Start from a famous game or enter your own numbers.

The payoff matrix
Every number is editable. The first payoff in each cell belongs to the row player, the second to the column player. A cell where both are flagged as a best reply is a Nash equilibrium.
What the solver finds
Recomputed on every keystroke, straight from the numbers above.
Why each player mixes the way they do
Each line is what one strategy pays as the opponent's behavior shifts. Where the lines cross, both strategies pay the same, and that crossing point sets the opponent's equilibrium mix.
Playing it again
A prisoner's dilemma played once has no escape. Played repeatedly against the same person, with no known final round, cooperation can become the rational choice.
How the model works

Best replies

Within each column, the row player's largest payoff is marked. Within each row, the column player's largest is marked. Ties mark both, which is how weakly dominant strategies reveal themselves.

Pure Nash equilibria

Any cell carrying both marks. Neither player can improve by moving alone, so the pair of choices is self-enforcing even without an agreement.

Mixed equilibria

Solved from the indifference conditions. The row player's own payoffs determine the column player's mix, and vice versa. q* = (d − b) / (a − b − c + d)

Pareto efficiency

An outcome fails if another exists that is at least as good for both players and strictly better for one. Those cells are drawn with a dashed border.

Dominance

A strategy strictly dominates when it beats the alternative against every possible opponent choice, and weakly dominates when it never loses and sometimes wins.

The repeated game

Under grim trigger, cooperating forever pays R / (1 − δ) while cheating once pays T + δP / (1 − δ). Comparing them gives the threshold δ ≥ (T − R) / (T − P).

What you are actually watching

Game theory has a reputation for being about outsmarting people. Most of it is the opposite: working out what happens when nobody can outsmart anybody, because everyone has already thought it through.

How to solve one of these by hand

The method the solver uses is the one worth learning, and it takes about twenty seconds on paper. Go column by column and underline the row player's best payoff. Then go row by row and underline the column player's best payoff. Any cell with two underlines is a Nash equilibrium.

That is all a Nash equilibrium is: a cell where each player is already doing the best they can against what the other is doing. It says nothing about whether the outcome is good. On the Prisoner's Dilemma preset the equilibrium pays both players 1 when they could have had 3, and it is still an equilibrium, because getting to the better cell needs both of them to move at the same time and neither can trust that.

The dilemma has nothing to do with prisoners

Load the Cartel quota preset. The numbers are different and the story is about two firms agreeing to restrict output, but the structure is identical: each firm does better by secretly overproducing whatever the other does, so both overproduce and the high price they were protecting disappears.

The same shape turns up in arms races, in overfishing a shared stock, in doping in sport, and in every situation where restraint benefits the group while cheating benefits the individual. Recognising the structure matters more than the story attached to it, and it is why the solution is almost never persuasion. It is binding contracts, external enforcement, or repetition.

Coordination games and anti-coordination games

Compare the Stag Hunt with Chicken. Both have two pure equilibria, and they are nothing alike. The Stag Hunt is a coordination game, sometimes called an assurance game. Chicken is an anti-coordination game, the same structure as the hawk-dove game in biology. Battle of the Sexes is a coordination game too, with the extra twist that the players disagree about which equilibrium they would prefer.

In the Stag Hunt the equilibria sit on the diagonal: both hunt stag, or both hunt hare. Everyone agrees that the stag outcome is better. The obstacle is purely doubt about the other person, which makes it a problem of trust and assurance. A promise, a convention, or a visible signal can solve it outright.

In Chicken the equilibria sit off the diagonal: one swerves and the other does not. Both players prefer to be the one who holds firm, so they disagree about which equilibrium to land on. Talking helps far less here, and what decides it is credible commitment, which is why the folk advice is to visibly throw away your steering wheel.

Mixing is stranger than it looks

Load Matching Pennies. There is no pair of definite choices that both players would stand by, so the only equilibrium is a mixed strategy: each side randomizes.

Now look at the two charts and notice what sets the probabilities. Your mix is calculated from your opponent's payoffs, not your own, because the condition being solved is that your opponent is left exactly indifferent between their options. If they were not indifferent they would have a favourite, and a favourite is a pattern you could be exploited for.

So mixing is not about maximizing your own payoff. It is about removing information. That is why goalkeepers and penalty takers, or auditors and people considering fraud, end up behaving unpredictably on purpose.

Why repetition changes everything

Drag the discount factor on the Prisoner's Dilemma. Below 0.50 cooperation collapses. At or above it, cooperation holds, and nothing about the payoffs has changed.

The threshold compares one round of gain from cheating against giving up cooperation permanently. Cheat and you gain 5 instead of 3, worth 2 today. Get punished forever and you receive 1 instead of 3 in every future round. Whether that trade is worth taking depends entirely on how heavily the future is weighted, which is what the discount factor measures.

This is the formal version of a familiar idea: people behave better with those they expect to deal with again. It also explains the sharp edge, since a game with a known final round unravels backwards. In the last round there is no future to lose, so both defect, and once that is anticipated the second-to-last round has no future worth protecting either.

What this model leaves out

A great deal, and the omissions are worth naming. Both players move at the same moment, so nothing here covers games where one side moves first and the other responds, which is where threats and commitments actually live. There are only two players and two options each. Everyone knows the full payoff matrix, so there is no private information and no bluffing in the poker sense.

The payoffs also carry more weight than they look like they do. They are meant to capture everything a player cares about, including guilt, reputation, and fairness. When real people cooperate in laboratory dilemmas more often than the matrix predicts, the usual explanation is that the matrix was written down wrong rather than that the players were irrational.

Things to try

Each one takes a few keystrokes and lands a point the matrix alone will not.

1Find out how big a fine has to be

On the Prisoner's Dilemma, cut both payoffs of 5 down to 2, as though cheating carried a penalty. The dominant strategy vanishes, but you now get two equilibria: the dilemma has turned into a stag hunt, which is a problem of trust rather than of incentives. Now also drop both mutual-confession payoffs from 1 to −1. Only then does staying silent become dominant and the good outcome unique.

2Turn a stag hunt into a dilemma

On the Stag Hunt, raise just the two payoffs of 3 that come from taking the hare while the other chases the stag, up to 6. Leave mutual hare at 3. Hunting hare becomes dominant for both, and the result is worse for everyone than the stag they gave up.

3Find the exact tipping point

On the Prisoner's Dilemma, drag the discount factor slowly through 0.50 and watch the verdict flip. Then raise the temptation payoff and watch the threshold move.

4Make a zero-sum game

Set the payoffs so every cell totals the same number. The solver will say so, and you will find there is nothing left to cooperate over.

5Build a game with no equilibrium

Start from Matching Pennies and change one number at a time until a pure equilibrium appears. It takes less than you would expect.

6Check your own homework

Type in a matrix from a problem set and compare the answer. The best-reply marks show the working, not just the result.

Common questions about game theory

What is a Nash equilibrium?

A pair of strategies where neither player can do better by changing only their own choice. It is a statement about stability rather than about goodness: an equilibrium can be bad for everyone involved and still be an equilibrium, because escaping it would require both players to move at the same time. That gap between stable and good is what most of the interesting problems in economics are about.

How do you find a Nash equilibrium by hand?

Use the underlining method. For each column, mark the row player's highest payoff. For each row, mark the column player's highest payoff. Any cell where both payoffs carry a mark is a Nash equilibrium, because each player is already giving their best reply to what the other is doing. This solver marks them for you, so you can check your working rather than just your answer.

Why is the prisoner's dilemma outcome bad if both players are rational?

Because each player has a dominant strategy: confessing pays better whatever the other does. Rationality points both of them toward it, and the result is worse for both than staying silent would have been. The failure is structural rather than a mistake, which is why the practical fixes are binding contracts, outside enforcement, or repeated dealings, rather than encouraging people to think harder.

What does a mixed strategy actually mean?

Deliberately randomizing between your options with fixed probabilities. The counterintuitive part is where the probabilities come from: each player mixes so that the opponent is left exactly indifferent between their own choices. You are not tuning the mix to maximize your own payoff, you are removing any pattern the opponent could exploit. Penalty shootouts and tax audits both work this way.

Why does repeating a game change the outcome?

Because punishment becomes possible. Under a grim trigger strategy, where any cheating ends cooperation permanently, cooperating is worth it when the discount factor is at least the gain from cheating divided by the gain from cooperating rather than feuding. Patient players facing an indefinite horizon can sustain an outcome that a single round cannot. A known final round breaks it, since the last round has no future to protect and the logic unravels backwards from there.

Do real people actually play these equilibria?

Often, but not always, and the exceptions are informative. People cooperate in one-shot dilemmas more than the matrix predicts, reject unfair offers that leave them better off, and coordinate on focal points the payoffs alone do not single out. The usual reading is that the written payoffs were incomplete rather than that the players failed at arithmetic, since fairness and reputation are real payoffs even when they are hard to put in a box.

About this model: two players, two strategies each, simultaneous moves, and complete information about every payoff. Real strategic problems usually involve sequential moves, more options, private information, and players who care about things the matrix does not record. Use this to learn the machinery and to check working, not to settle an argument about how somebody will behave.