What you are actually watching
Game theory has a reputation for being about outsmarting people. Most of it is the opposite:
working out what happens when nobody can outsmart anybody, because everyone has already
thought it through.
How to solve one of these by hand
The method the solver uses is the one worth learning, and it takes about twenty seconds on
paper. Go column by column and underline the row player's best payoff. Then
go row by row and underline the column player's best payoff. Any cell with
two underlines is a Nash equilibrium.
That is all a Nash equilibrium is: a cell where each player is already doing the best they
can against what the other is doing. It says nothing about whether the outcome is good. On
the Prisoner's Dilemma preset the equilibrium pays both players 1 when they could have had 3,
and it is still an equilibrium, because getting to the better cell needs both of them to move
at the same time and neither can trust that.
The dilemma has nothing to do with prisoners
Load the Cartel quota preset. The numbers are different and the story is about two firms
agreeing to restrict output, but the structure is identical: each firm does better by
secretly overproducing whatever the other does, so both overproduce and the high price they
were protecting disappears.
The same shape turns up in arms races, in overfishing a shared stock, in doping in sport,
and in every situation where restraint benefits the group while cheating benefits the
individual. Recognising the structure matters more than the story attached to it, and it is
why the solution is almost never persuasion. It is binding contracts, external enforcement,
or repetition.
Coordination games and anti-coordination games
Compare the Stag Hunt with Chicken. Both have two pure equilibria, and they are nothing
alike. The Stag Hunt is a coordination game, sometimes called an assurance
game. Chicken is an anti-coordination game, the same structure as the
hawk-dove game in biology. Battle of the Sexes is a coordination game too, with the extra
twist that the players disagree about which equilibrium they would prefer.
In the Stag Hunt the equilibria sit on the diagonal: both hunt stag, or both
hunt hare. Everyone agrees that the stag outcome is better. The obstacle is purely doubt
about the other person, which makes it a problem of trust and assurance. A
promise, a convention, or a visible signal can solve it outright.
In Chicken the equilibria sit off the diagonal: one swerves and the other
does not. Both players prefer to be the one who holds firm, so they disagree about which
equilibrium to land on. Talking helps far less here, and what decides it is
credible commitment, which is why the folk advice is to visibly throw away
your steering wheel.
Mixing is stranger than it looks
Load Matching Pennies. There is no pair of definite choices that both players would stand
by, so the only equilibrium is a mixed strategy: each side randomizes.
Now look at the two charts and notice what sets the probabilities. Your mix is calculated
from your opponent's payoffs, not your own, because the condition being solved is
that your opponent is left exactly indifferent between their options. If
they were not indifferent they would have a favourite, and a favourite is a pattern you
could be exploited for.
So mixing is not about maximizing your own payoff. It is about removing information. That
is why goalkeepers and penalty takers, or auditors and people considering fraud, end up
behaving unpredictably on purpose.
Why repetition changes everything
Drag the discount factor on the Prisoner's Dilemma. Below 0.50 cooperation collapses. At or
above it, cooperation holds, and nothing about the payoffs has changed.
The threshold compares one round of gain from cheating against giving up cooperation
permanently. Cheat and you gain 5 instead of 3, worth 2 today. Get punished forever and you
receive 1 instead of 3 in every future round. Whether that trade is worth taking depends
entirely on how heavily the future is weighted, which is what the
discount factor measures.
This is the formal version of a familiar idea: people behave better with those they expect
to deal with again. It also explains the sharp edge, since a game with a known final round
unravels backwards. In the last round there is no future to lose, so both defect, and once
that is anticipated the second-to-last round has no future worth protecting either.
What this model leaves out
A great deal, and the omissions are worth naming. Both players move at the same moment, so
nothing here covers games where one side moves first and the other responds, which is where
threats and commitments actually live. There are only two players and two options each.
Everyone knows the full payoff matrix, so there is no private information and no bluffing in
the poker sense.
The payoffs also carry more weight than they look like they do. They are meant to capture
everything a player cares about, including guilt, reputation, and fairness. When real people
cooperate in laboratory dilemmas more often than the matrix predicts, the usual explanation
is that the matrix was written down wrong rather than that the players were irrational.
Things to try
Each one takes a few keystrokes and lands a point the matrix alone will not.
1Find out how big a fine has to be
On the Prisoner's Dilemma, cut both payoffs of 5 down to 2, as though cheating carried
a penalty. The dominant strategy vanishes, but you now get two equilibria: the
dilemma has turned into a stag hunt, which is a problem of trust rather than of incentives.
Now also drop both mutual-confession payoffs from 1 to −1. Only then does staying silent
become dominant and the good outcome unique.
2Turn a stag hunt into a dilemma
On the Stag Hunt, raise just the two payoffs of 3 that come from taking the hare while
the other chases the stag, up to 6. Leave mutual hare at 3. Hunting hare becomes dominant
for both, and the result is worse for everyone than the stag they gave up.
3Find the exact tipping point
On the Prisoner's Dilemma, drag the discount factor slowly through 0.50 and watch
the verdict flip. Then raise the temptation payoff and watch the threshold move.
4Make a zero-sum game
Set the payoffs so every cell totals the same number. The solver will say so, and you
will find there is nothing left to cooperate over.
5Build a game with no equilibrium
Start from Matching Pennies and change one number at a time until a pure equilibrium
appears. It takes less than you would expect.
6Check your own homework
Type in a matrix from a problem set and compare the answer. The best-reply marks show
the working, not just the result.
Common questions about game theory
What is a Nash equilibrium?
A pair of strategies where neither player can do better by changing only their own choice. It is a statement about stability rather than about goodness: an equilibrium can be bad for everyone involved and still be an equilibrium, because escaping it would require both players to move at the same time. That gap between stable and good is what most of the interesting problems in economics are about.
How do you find a Nash equilibrium by hand?
Use the underlining method. For each column, mark the row player's highest payoff. For each row, mark the column player's highest payoff. Any cell where both payoffs carry a mark is a Nash equilibrium, because each player is already giving their best reply to what the other is doing. This solver marks them for you, so you can check your working rather than just your answer.
Why is the prisoner's dilemma outcome bad if both players are rational?
Because each player has a dominant strategy: confessing pays better whatever the other does. Rationality points both of them toward it, and the result is worse for both than staying silent would have been. The failure is structural rather than a mistake, which is why the practical fixes are binding contracts, outside enforcement, or repeated dealings, rather than encouraging people to think harder.
What does a mixed strategy actually mean?
Deliberately randomizing between your options with fixed probabilities. The counterintuitive part is where the probabilities come from: each player mixes so that the opponent is left exactly indifferent between their own choices. You are not tuning the mix to maximize your own payoff, you are removing any pattern the opponent could exploit. Penalty shootouts and tax audits both work this way.
Why does repeating a game change the outcome?
Because punishment becomes possible. Under a grim trigger strategy, where any cheating ends cooperation permanently, cooperating is worth it when the discount factor is at least the gain from cheating divided by the gain from cooperating rather than feuding. Patient players facing an indefinite horizon can sustain an outcome that a single round cannot. A known final round breaks it, since the last round has no future to protect and the logic unravels backwards from there.
Do real people actually play these equilibria?
Often, but not always, and the exceptions are informative. People cooperate in one-shot dilemmas more than the matrix predicts, reject unfair offers that leave them better off, and coordinate on focal points the payoffs alone do not single out. The usual reading is that the written payoffs were incomplete rather than that the players failed at arithmetic, since fairness and reputation are real payoffs even when they are hard to put in a box.