Behavioral Economics
EC 404 ·Psychology and Economics

Expected Utility, concluded · The Evidence Against It begins

Two applications that finish the standard model; what the numbers a utility function returns do and do not mean; how to say that one person is more risk averse than another; and then the first cracks, where the theory stops describing anybody.

Slides PDF · 516 KB

Open the deck in a new tab ↗ If the viewer doesn't load, the link above always works.

Everything I said is now written down

Last class we opened choice under uncertainty: expected value, expected utility, and enough of a calculus toolkit to work with them.

A note before anything else. Not everything in a deck can be talked through in the room, and you should know that the gap is covered. Everything I say has now been turned into lecture notes, more or less — which is what you are reading, and why I record. So when in doubt, open the slides and do some scrolly-scrolly. Somebody asked about the CRRA utility function, and the answer was sitting in the slides the whole time. Credit where it is due. There is also a calculus refresher in there, which I will not walk through: slopes, the chain rule, a few worked examples. If calculus makes you nervous, that is the thing to read.

Today we finish the standard model with two applications, insurance and financial assets. Then we start pulling it apart.

Problem Set Lesley is due a week from today — Tuesday, September 15. Do not wait until the last minute, though you are where you are.

Drill and re-drill

Our key definitions from last time were these notions of risk aversion, and I want to drill them into you. We will drill and re-drill. Memorize this; it is not hard.

A person is risk averse if, for any lottery, she prefers to have the expected value of that lottery instead of the lottery itself.

That is it. There are analogous definitions for risk neutrality and risk loving.

We illustrated the danger of reaching for intuition instead. Take a lottery whose expected value is $5, and suppose I offer a person $4.50. Will she take the $4.50 rather than the lottery? The intuitive account of risk aversion says: yes, probably, she does not like risk.

Now — are the words does not like risk anywhere on the screen? No. So we apply the definition instead, and the definition says we need more information. We would need to know how risk averse she is.

Then we tied those definitions to a model of choice, expected utility theory. And I flagged that the version I first wrote down was a little sloppy, because it left out wealth. The correct statement carries the changes in wealth — the incomes, the pluses and minuses — into the wealth you started with, and weights each by its probability:

EU(x)=∑i=1npi u(w+xi)EU(\mathbf{x}) = \sum_{i=1}^{n} p_i \, u(w + x_i)

A function with a dial on it

There are many functions we could use for uu. Log. Square root. Plenty of others. One is worth singling out, and it says revisited on the slide although we are visiting it properly for the first time.

The CRRA utility function is

u(x)=x1−ρ1−ρfor ρ∈[0,∞)u(x) = \frac{x^{1-\rho}}{1-\rho} \qquad \text{for } \rho \in [0, \infty)

The letters stand for constant relative risk aversion. Why that name is the right name is something we will earn at the end of this stretch; you will also explore it on your own.

The reason I introduce this particular function is that it has a parameter in it. The Greek letter is rho. If you do not care for a lowercase rho, write whatever letter you like — it does not matter. Try my favourite, squiggle. Or eta, beta, gamma, delta, epsilon. At which point I am simply naming sororities and fraternities.

That parameter captures variation across people. The admissible values run from zero up to positive infinity, so it must be a positive real number. And remember that infinity always gets a soft bracket.

Now watch what the dial does. Put ρ=0\rho = 0 in. Just put it in. This is a person — call her Sarah — and Sarah has ρ=0\rho = 0. So Sarah’s utility for total wealth xx is

u(x)=x1−01−0=x11=xu(x) = \frac{x^{1-0}}{1-0} = \frac{x^{1}}{1} = x

Sarah’s utility function is linear. And we established last time that a linear utility function is synonymous with risk neutrality.

All other values — every positive ρ\rho — imply some degree of risk aversion. This function does not accommodate risk lovingness. If you start feeding it negative numbers you will get zany things.1

Two lotteries and a poor person

We can work with this. Suppose a person has wealth of $100 — they are very poor — and is considering two lotteries.

Lottery x\mathbf{x} is ( $40, 1/2 ; $0, 1/2 ). Lottery y\mathbf{y} is ( $120, 1/3 ; −$30, 2/3 ).

First, consult your intuition, though you are not going to eyeball this one, so do not try. I am probably way better at mathematics than you are, and I cannot eyeball it either. We have to do some plugging in.

But before any plugging in, we have to take sentences — descriptions of lotteries — and turn them into mathematics. That is the recurring move. We ask when the expected utility of x\mathbf{x} equals the expected utility of y\mathbf{y}, and we leave ρ\rho untethered, then solve for particular cases.

Which does the person pick if risk neutral? There you are putting in ρ=0\rho = 0, or equivalently a linear utility function, or equivalently again, you are calculating expected value. Risk neutrality means the person acts as an expected value maximizer. If that does not ring a bell, review the first lecture.

And which does the person pick as an expected utility maximizer with ρ=1/2\rho = 1/2?

You can do this on your own, and it is a fun exercise, so I will not spoil it. I will give you one hint, which is on the slide: in this particular case you can reach the answer without grinding out both expected utilities. What we are doing here is capturing a variety of risk attitudes with one simple functional form.

Concave, linear, convex

To bring risk attitudes into expected utility at all, we have to assume something about the shape of the function. We cannot leave it generic — a uu with a parenthesis after it and nothing said about it.

I presented the result last time. Under expected utility theory, a person is risk averse if and only if her utility function is concave. If it is linear, the person is risk neutral. If it is convex — cup up — the person is risk loving.

This is my own reminder to say: the calculus refresher is in the deck if you need it. We are about to use some.

A thousand dollars, and where to put it

It is an economics class, so here is a radical notion: let us watch somebody buy something.

Suppose you have $1,000, you are going to invest it for one year, and then consume it. By consume it, I mean spend it, so it goes through your utility function. This is $1,000 that Nana gave you and that you will use next year. You have only $1,000 because you are a poor student — a theme to which we will return many times.

There are two assets. One is a so-called risk-free asset yielding a certain return of zero percent. What is that asset? Cash. Keep the brain on; stay limber. The other is a risky asset yielding a return of 50% with probability one half and −40% with probability one half.

If I were merely asking whether the person wants this asset or that one, the answer would turn on how risk averse she is, and we could nearly eyeball it. Suppose she had to put everything into one or the other. The risky asset has positive expected value — the average of +50% and −40% is still positive — so a risk-neutral agent puts it all in the risky asset.

But the very name of the thing suggests that somebody who succumbs to risk aversion might want to put in less than everything. And she might want to put in some amount in between, continuously. So we cannot brute-force this. We cannot ask is it this one or that one. We have to choose an amount and maximize.

So we create a choice variable. Call it yy: the amount you invest in the risky asset, which leaves 1000−y1000 - y in the risk-free one.

I am going to say this repeatedly, because many of you will get bogged down by it: I will reuse letters and I will switch letters. They are letters. The context makes it obvious, and I define them right where I use them. Here it is a yy. Sometimes it will be an alpha. I am a zany guy; I do many things.

Now turn the situation into a lottery by describing the contingent states. What can happen? Only two things. In the good world you get a 50% return on the amount you invested; in the bad world you take the loss.

The amount you keep as cash comes back to you whole — you get 100% of it — and the amount you invested gets multiplied by 1.5:

(1000−y)(1.00)+y(1.50)=1000+0.5y(1000 - y)(1.00) + y(1.50) = 1000 + 0.5y

If things go badly, the same logic applies, with the invested portion multiplied by 1−0.41 - 0.4:

(1000−y)(1.00)+y(0.60)=1000−0.4y(1000 - y)(1.00) + y(0.60) = 1000 - 0.4y

To make this a lottery we only have to attach probabilities, and here that is trivial, because it is fifty-fifty:

( $1000 + 0.5y, 1/2 ; $1000 − 0.4y, 1/2 )

Notice that yy is still sitting inside it. Of course it is: yy is the thing we are choosing.

Now turn that sentence — note the semicolon — into mathematics by taking the expected utility. In a real problem there would be a wind-up: Jane is an expected utility maximizer, something of that sort. Take the probability of each outcome, multiply by the utility of that outcome, and note that the wealth is already folded in, because you started with $1,000:

EU(x(y))=12 u(1000+0.5y)  +  12 u(1000−0.4y)EU(\mathbf{x}(y)) = \tfrac{1}{2}\,u(1000 + 0.5y) \;+\; \tfrac{1}{2}\,u(1000 - 0.4y)

And now I have to maximize something with respect to a continuous variable. You can throw up your hands in grief, or do calculus, or both — they are not mutually exclusive.

Of course, you need an actual function. Let us take a log. Take the derivative with respect to yy, set it to zero, move things around, and in this case it comes out at y∗=250y^{*} = 250.

Try it yourself. People often ask what good exercises look like. That is one.

The squiggle means precisely tied

Now the insurance example from last class.

You face a lottery, and you face a choice between that lottery and another. What we did last time was set up both lotteries and impose an indifference condition — that is what the little squiggly equals sign, ∼\sim, means. It says the lottery on the left is indifferent to the lottery on the right.

The word indifference is a linguistic trap. In ordinary speech it carries a connotation of meh. Here it means the person is precisely tied. There is a version of meh that fits: if you said to a person who is truly tied, “I will just pick the left one,” they would say, “I do not care.” And the right one? “I do not care.” Because they are tied.

Turning that sentence into mathematics means taking the expected utility of each side and setting them equal, which gave us

u(1000−p∗)=0.9 u(1000)+0.1 u(750)u(1000 - p^{*}) = 0.9\,u(1000) + 0.1\,u(750)

Insuring a fraction

That is where we stopped. But in the real world — and I promise there will be small steps toward real things, even if I cannot make them properly real — you do not face a single take-it-or-leave-it policy. You insure against a proportion of the loss, and you pay in proportion.

The insurance company does not mind. It will happily sell you 50% insurance at 50% of the premium. So: to insure the whole $250 loss costs pp; to insure a proportion α\alpha of it costs αp\alpha p.

What is the optimal α\alpha?

Here is where people get tripped up. There are two letters in play, one English and one Greek, and the question arises: which one am I maximizing over? Just remember to be a person. pp is the price. You do not get to optimize against the price, because your optimal price is zero — no kidding — or negative, if the company is willing to pay you. So that is not the choice variable. The choice variable is the proportion of the loss you cover.

How do we do it in practice? Calculus, same as before, and the first step is the same as in the financial case: set up the lottery as a function of the choice parameter. Then take expected utility as a function of it. Then do calculus.

You had $1,000, and the chance of the loss was 10%. So 90% of the time nothing happens to you medically — but something still happens to you financially, because you paid the premium. In the language of the financial example, that is the good state: I do not get hit by a car, and I do not lose $250. I am only out the fraction of the premium I bought.

In the bad state I lose the $250 but get some fraction of it back. So:

( $1000 − pα, 0.9 ; $750 + (250 − p)α, 0.1 )

You can write that second outcome however you like. Expanded, it is 750+250α−pα750 + 250\alpha - p\alpha. The pαp\alpha is always there, in both states, because it is the fraction of the premium you pay no matter what.

The next step is easy. Put the probability out front, put the garbage inside a utility function — I have not told you which function, so it just goes inside a uu — and now it is mathematics:

EU(x(α))=0.9 u(1000−pα)  +  0.1 u(750+(250−p)α)EU(\mathbf{x}(\alpha)) = 0.9\,u(1000 - p\alpha) \;+\; 0.1\,u\big(750 + (250 - p)\alpha\big)

This is the object our participant, our agent, our person is going to maximize with respect to α\alpha — the proportion he or she chooses, zero and one included.

Escrow, and other words built to confuse you

A helpful definition can come in here, and this is a bit of bridging between the ideas in my head and the world outside.

Insurance comes wrapped in a great deal of jargon. My theory about that — and this is my actual theory, I am not being cute — is that the jargon exists to confuse you. It is there to create demand that would otherwise be lower if the product were less opaque.

You will eventually buy a home. Every month, on top of principal and interest, I pay something called escrow. I believe that term exists to confuse people. Escrow is this: they take your annual property taxes, divide by twelve, and hold the money in a side account so that they can pay the taxes for you and you never have to write a cheque to your taxing authority. That is the whole thing. But then there is an escrow agent, and you pay the escrow agent, and this is all scams. My working theory of a good deal of what happens in insurance markets is opacity in the service of demand.

So part of what we are doing is pulling back the curtain and asking what these terms actually mean.

Actuarially fair, unfair, overfair

Insurance is actuarially fair if the premium — what you pay no matter what — equals the expected payout.

This is not insurance that exists on the market. It is hypothetical, and it is hypothetical for an obvious reason: an insurance company selling it makes no money.

And yet insurance is an interesting thing. Anybody here snowboard? Or ski? One person. I suppose it is a very flat state; in Oregon it is everybody. If you are a snowboarder, or a skier, or a helicopter snowboarder, or a practitioner of whatever your own favourite reckless activity happens to be, then everyone else in this room is subsidizing your health insurance to some degree — because statistically the rest of them are pretty healthy. The healthy are subsidizing those who do risky things. And it is quite plausible that you personally hold insurance that is actuarially unfair to you.

So there are the corollary definitions. Insurance is actuarially unfair if the premium is larger than the expected payout, and actuarially overfair if the premium is smaller.

Actuarially overfair insurance cannot happen on the whole. Mathematically it cannot. But it very much can in the small, and the reason is that the insurance company does not know about you. It knows about the statistical you. If you are a far more reckless driver than others of your age, you are getting actuarially overfair insurance, and the safe drivers are paying for it.

Now suppose insurance were actuarially fair in our running example. The company has to pay out $250, and that happens with probability 10%, so the price of actuarially fair insurance is

p=(0.10)(250)=$25p = (0.10)(250) = \$25

In general, the actuarially fair price is the highest amount a risk-neutral person is willing to pay — it is the price at which they sit right on the line, indifferent between buying and not.

A risk-averse person is willing to pay more. Why? You could say: because they are risk averse, they do not like risk. You are not entirely wrong, but you are not using the definition I need drilled into your head. Here is the reason. When insurance is actuarially fair, the company is offering you the expected value of a gamble. Take full insurance and you have no risk at all; you are out a little money, and the amount you are out is precisely the expected value of the gamble you would otherwise have played with your car, or your health, or your life.

Now remind yourself of the definition. A person is risk averse if she would rather have the expected value of a lottery for sure than play that lottery out. It is on the slide. So a risk-averse person would certainly take insurance at $25. She might pay $26. She might pay $27. How much she will pay depends on how risk averse she is.

Why the risk averse fully insure

Substitute p = \25$ into what we built:

EU(x(α))=0.9 u(1000−25α)  +  0.1 u(750+225α)EU(\mathbf{x}(\alpha)) = 0.9\,u(1000 - 25\alpha) \;+\; 0.1\,u(750 + 225\alpha)

In general we would now need calculus. But let us use our brains a bit first.

Here is a portable and profitable habit, and it is the kind of exercise to do while sitting on a park bench. Never mind this class for a moment — the class is not the important thing. The important thing is building skills you can carry out of here and use profitably elsewhere. And one of those is: think about the extremes. Think about edge cases. Edge cases are usually far easier than interior ones.

So instead of solving for the best α\alpha, ask what happens if we buy as much insurance as they will sell us, and what happens if we buy none.

Buy everything, α=1\alpha = 1:

EU(x(1))=0.9 u(1000−25)+0.1 u(750+225)=0.9 u(975)+0.1 u(975)=u(975)EU(\mathbf{x}(1)) = 0.9\,u(1000 - 25) + 0.1\,u(750 + 225) = 0.9\,u(975) + 0.1\,u(975) = u(975)

because nine tenths plus one tenth is one. Buy nothing, α=0\alpha = 0:

EU(x(0))=0.9 u(1000)+0.1 u(750)EU(\mathbf{x}(0)) = 0.9\,u(1000) + 0.1\,u(750)

Notice anything? We have already said something about this pair. That number, 975, is the expected value of the uninsured lottery:

0.9×1000+0.1×750=9750.9 \times 1000 + 0.1 \times 750 = 975

Somebody asked what uu is here, and the answer is worth stating plainly: uu is whatever utility function the person happens to have. We are deliberately not writing it down. We are not restricting it yet. We take only the perspective that if she is risk averse, then that function is concave — which we know by definition, because that is what the result said.

And so, from the definition and nothing else: if insurance is actuarially fair and you are risk averse, you will fully insure. That is an always-true conclusion.

Twenty-five dollars, and then two hundred fifty thousand

The conclusion does not depend on the magnitudes. A risk-averse consumer reaches the same answer whether the loss is our silly $250, or a $25,000 loss on an actual car, or $250,000. Nothing about the small numbers is misleading; I write small numbers because they fit on slides.

And here is where I want you to start consulting your intuitions, because the class has that sneaky word in front of economics on every slide, and so far we have done none of it. Let us try.

Does the question feel the same when the loss is $250,000 as when the loss is $25? The mathematics is identical — divide everything by ten thousand and see. If you do not believe me, put it on the board and start dividing. It will not matter.

But we all lose $25. Who cares? I am not buying insurance against $25 losses. For $250,000, maybe — probably. Even at an actuarially fair price, though, am I really going to hand over $2.50 to protect $25? I am going to roll the dice.

That verbal description of behaviour does not correspond to expected utility. It is something else.

The same answer, without naming a function

By the way — and this is an italicized definition, a helpful fact from high-school mathematics that you probably do not remember:

A function ff is concave if, for any xx and yy and any α∈[0,1]\alpha \in [0,1],

f((1−α)x+αy)  ≥  (1−α)f(x)+αf(y)f\big((1-\alpha)x + \alpha y\big) \;\geq\; (1-\alpha)f(x) + \alpha f(y)

It is saying, visually, that the straight line always lies below the curve. That is all concavity is, written in slightly opaque language.

We know that risk aversion is synonymous with concavity. So apply concavity to what we have, which just means moving the mixing inside the function:

u(0.9(1000−25α)+0.1(750+225α))  ≥  0.9 u(1000−25α)+0.1 u(750+225α)u\big(0.9(1000 - 25\alpha) + 0.1(750 + 225\alpha)\big) \;\geq\; 0.9\,u(1000 - 25\alpha) + 0.1\,u(750 + 225\alpha)

and the left-hand side collapses:

u(975)  ≥  0.9 u(1000−25α)+0.1 u(750+225α)u(975) \;\geq\; 0.9\,u(1000 - 25\alpha) + 0.1\,u(750 + 225\alpha)

Same conclusion by a direct route. Paying $25 for certain — holding $975 for certain — beats playing out the lottery.

Notice what I did not have to use. I did not need a log. I did not need any particular function. I needed the definition of concavity. So it is true for every concave function.

The levels do not mean anything

Some helpful tricks now, and the first is geeky. It is completely, 100% mathematically true, and its geekiness is useful in other economics classes. If you are an econ major, write it down and learn what it means. If you are not, or you are otherwise indifferent, it is a fun fact that is genuinely helpful. It is not on any exam.

It has a very geeky wind-up: the utility function is unique up to a positive affine transformation.

We have to break that down, starting with what a positive affine transformation is. It looks like y=mx+by = mx + b — your high-school slope-intercept form. Take any function, call it uu, and call the transformed one u~\tilde{u}. Then

u~(x)=a⋅u(x)+bfor some a>0 and any b\tilde{u}(x) = a \cdot u(x) + b \qquad \text{for some } a > 0 \text{ and any } b

You multiply by any positive number and add any number at all. The additive one can be anything — positive, negative, 0.0000001. It does not matter.

And two utility functions reflect the same preferences if, for any pair of lotteries, EU(x)>EU(y)EU(\mathbf{x}) > EU(\mathbf{y}) precisely when EU~(x)>EU~(y)E\tilde{U}(\mathbf{x}) > E\tilde{U}(\mathbf{y}). Swapping the functions preserves the ordering of expected utilities. Since expected utilities are what dictate choice, the two functions mean the same thing.

The theorem is that uu and u~\tilde{u} reflect the same preferences if and only if one is a positive affine transformation of the other.

So what? Who cares? This is dumb, I am sleepy. All viable comments. Let me make you halfway care.

Consider another class — a less interesting class, at some hypothetical other university, since all of ours are excellent — where you are made to plug and chug. Somebody hands you the utility function 10x10\sqrt{x} and asks what the person chooses. I am telling you that the answer is the same as if you had used x\sqrt{x}. Multiply by one tenth. It gives the same answer in every class, always.

The deeper purpose shows up in the second example. Suppose I tell you that

10x−20,000,00010\sqrt{x} - 20{,}000{,}000

represents the same preferences as x\sqrt{x}. Wait a second — are not all my numbers going to be negative now? I am subtracting twenty million.

Yes. And: can a utility function be negative? Yes. Can an expected utility be negative? Yes.

The numbers do not mean anything. The order of the numbers means everything. If the utility of lottery x\mathbf{x} is −1,000 and the utility of lottery y\mathbf{y} is −900, the person prefers y\mathbf{y}, because −900 is bigger than −1,000. There is no sense in which those are bad lotteries. I can turn any lottery into a negative-numbered lottery by subtracting twenty million.

The numbers do not correspond to human sensations. They correspond to preferences — to what a person picks when choosing among lotteries. I promised in the first real lecture that this section is about a person choosing over lotteries, and that is what I meant. It is not about how they feel about lotteries. There is a natural relationship between the two, of course. But the levels do not matter, only the comparisons. And you can use this trick any time. It will always work.

If the last few minutes felt like gobbledygook, let it wash over you like a warm Caribbean wave. You are floating in an ocean of happiness, and it is fine. It will not be on an exam. It might help you in life, and exposing yourself to different ideas is, I conjecture, a portable skill as well.

Two coins at once

Sometimes you face more than one lottery at the same time, and expected utility operates on the reduced lottery — you have to combine them first.

Suppose I flip one coin and then another. The first, x\mathbf{x}, pays $100 with probability one half and −$90 with probability one half. The second, y\mathbf{y}, pays $95 and −$85, again fifty-fifty.

I used coins as the example, but not everything is independent. If these are independent coin flips, you enumerate: heads-heads, heads-tails, tails-heads, tails-tails. Heads-heads is the good outcome twice, $100 and $95, so $195, and its likelihood is one half times one half, a quarter. Heads-tails is good then bad: $15, again a quarter. Tails-heads is −$90 plus $95, so $5. And tails-tails is −$175.

But they could have any correlation structure at all. Under perfect positive correlation I only need to flip the first coin, because the second is then determined: heads on the first guarantees heads on the second. So you get $195 or −$175, and nothing in between. Under perfect negative correlation you get the two middle outcomes instead.

A two-by-two grid of the four combined outcomes of the two coin flips, and a table showing the probability each outcome carries under independence, perfect positive correlation, and perfect negative correlation. the four combined outcomes y = +95 y = −85 x = +100 x = −90 +195 +15 +5 −175 probability each carries +195 +15 +5 −175 independent 1/4 1/4 1/4 1/4 perfectly positive 1/2 — — 1/2 perfectly negative — 1/2 1/2 —
Fig. 1. Each coin still pays its good outcome half the time in every row here. What changes is the correlation, and it decides which of the four cells you can reach: all of them when the flips are independent, only the corners when the coins move together, only the middles when they move against each other.

If you would like to get into real financial trouble in your life, ignore correlations. Pretend they do not exist. That is a pro tip for losing all of your money — or for getting electoral predictions very wrong. The outcomes in Michigan in a few months are correlated with the outcomes in Texas, and more correlated still with the outcomes in Ohio.

So when you face several lotteries at once, reduce them first. This will come up on Thursday.

How risk averse, in a number

Finally, returning to a dangling promise about the CRRA function. There is a natural way to describe how risk averse a person is. It is a little opaque through the mathematics, though the goal is easy to state.

The goal is something like this: if Allison is more risk averse than Barbara, then Allison is less likely to accept a particular gamble. That is a true-ish statement, and we want to sharpen it. So how do we formalize more risk averse? We reach for this rather nasty expression.

The coefficient of risk aversion at wealth ww is

r(w)=−u′′(w)u′(w)r(w) = \frac{-u''(w)}{u'(w)}

Notice first that it is assessed at a given level of wealth. Why? Because a person’s degree of risk aversion may intuitively depend on how much wealth they have. In every other way I am exactly like Elon Musk, except our wealth. That is not true, as far as I can tell — I do not know the man. But intuitively, a person with very high wealth may have a different degree of risk aversion.

We capture it with a ratio of the second derivative to the first. The idea: the second derivative captures how cupped down the function is at that point — how cuppy it is, if I make a cup with my hand. The first derivative says how fast it is rising overall. So the ratio asks: relative to how much the function is going up, how curved is it? That is what the expression captures, intuitively.

And to compare Allison with Barbara, we evaluate each at her own wealth. If Allison’s coefficient exceeds Barbara’s, then Allison is less likely to accept a generic risky gamble. That is the whole utility of the approach.

I am not going to make you take two derivatives. One derivative you will do a few times; two is impossible. But note the two features. First, it depends on wealth — wealth is the argument. Second, it depends on the utility function.

Absent further assumptions I cannot tell you how it depends on wealth. I gave you the intuitive version: as wealth goes up, risk aversion might go down. So economists, and humans, often assume utility functions for which that is true. But it is an assumption, and it comes from the functions we choose.

By the way, the affine-transformation trick does nothing to this. Why? Because the derivatives make all that garbage go away.

Why CRRA is called that

The CRRA function has a particularly nice feature here. I am going to swap in a ww, because we are talking about wealth, and elsewhere you will see a zz or some other stand-in — the argument might itself be a lottery, in which case it is ww plus something and ww minus something, and I cannot reuse ww or xx if those letters are busy elsewhere in the problem. Expect squiggles and Greek letters.

We have seen the function. At ρ=0\rho = 0 it is risk neutrality; at ρ=1\rho = 1 it is log. And for every value of ρ\rho, the coefficient of risk aversion is

r(w)=ρwr(w) = \frac{\rho}{w}

Put that in a little box in your notes — though you could also rediscover it yourself, since it is just the ratio of the negative second derivative to the first. I know you are already annoyed by the amount of calculus, so I have taken it for you.

Look at what that expression does. The coefficient is increasing in ρ\rho: turn the dial up, and the number goes up. So ρ\rho is doing the job I promised it was doing all along, capturing the degree of risk aversion. And the coefficient is decreasing in wealth, because we are dividing by wealth — which preserves the notion that Jeff Bezos is less risk averse than I am.

And because it is a ratio, scaling everything up by a common factor produces the same relative sensations. Which is where the name comes from. Constant relative risk aversion: the feelings about a gain-$11, lose-$10 gamble from a wealth of $1,000 are the same as the feelings about a gain-$110, lose-$100 gamble from a wealth of $10,000, and the same again about gain $1,100, lose $1,000 from $100,000. We will explore that feature as we go.

How are we doing? Brains still somewhat present? I have got some scared faces out there. I do not want to look at anybody and call them out for their scared face, but you know who you are.

Normative and positive

A brief aside. Stop taking notes. Just chill. Or take notes if you like. You do you.

This one relates to the political conversation around economics, and you may have heard the distinction between normative and positive theories. Normative: what people should do. Positive: what people do do.

Many people regard expected utility as a good normative theory — as what one ought to do. More so, in fact, than the effective altruists who believe in expected value maximization; it is far less controversial to say you ought to behave according to expected utility. Here is one reason why.

Take the following properties and suppose we want people’s choices to have them. We, as analysts, are saying: this is how we would like humans to behave.

Preferences are complete and transitive. Complete means there are no holes: you cannot construct some coin flip and have the person say I do not know. Transitive means that if you prefer lottery AA to BB and BB to CC, then you prefer AA to CC. Reasonable, uncontroversial.

Preferences are continuous. Suppose you like lottery x\mathbf{x} a little more than lottery y\mathbf{y}. Now I add a 0.001 chance of an orange — a literal orange — to both. You shrug. It does not suddenly reverse your preference between them. Again, uncontroversial.

The independence axiom. Related to continuity, and in fact the same thing the way I just said it. If you prefer x\mathbf{x} to y\mathbf{y}, then take some third lottery z\mathbf{z} and mix it equally into both; the preference should not flip.

Somebody asked about the symbol, ≿\succsim, which allows for a tie — it means prefers or is indifferent to. You may ignore the tie if you like; drop the little squiggle underneath and the conclusion is the same. And the exercise is: I prefer x\mathbf{x} to y\mathbf{y}, so now I offer a 50% chance of x\mathbf{x} and 50% chance of z\mathbf{z} against a 50% chance of y\mathbf{y} and 50% chance of z\mathbf{z}. My preference has to stay put. And yes — it is the same as the orange. The only difference is that continuity requires it to hold as the z\mathbf{z} gets small. It is a technical assumption; it is not important.

Do these feel like reasonable things a person should do? If you believe that, then a person behaves according to expected utility theory. Those axioms are equivalent to it.

If you want to prove it, go to graduate school. It is hard. It is math-hard, not deep — just annoying. John von Neumann proved it originally: a polymath of the early-to-mid twentieth century, probably the smartest person we are aware of ever having lived. Einstein said he was way smarter than him. Look him up; he is a fun figure.

Ultimately, questions about what one should do are left either to a group of people — who can decide collectively, typically by voting — or to you yourself.

The end of the standard model

So there is a real reason people think we should describe humans this way, and I agree with it. It would be nice if humans behaved like this.

They do not. And this is the end of the standard model.

The next stretch is where we get our feet wet in actual behavioral economics. The good news is that the topics are somewhat front-loaded in intensity: we needed to understand the inner workings of the standard model first, and that was vitally important, because only now can we see the edges where it is brittle.

The Allais paradox

Imagine I ask you this. In question one, do you want a million dollars for sure — option A? Or option B: a million dollars with 89% chance, five million with 10% chance, and nothing with 1% chance?

Hands went up for A, then for B.

Question two. Option C: a million dollars with 11% chance and nothing with 89%. Option D: five million with 10% chance and nothing with 90%.

Hands for C, then for D.

If you chose A and then D, you have succumbed to the paradox. Not many did in here — though I suspect if I multiplied every number by ten I would catch more of you. A million dollars is not cool any more. I have done this whole bit.

So how is it a paradox? How does it violate expected utility? Step one, as always: put it on the wall and write it down.

The expected utility of A is 1×u(1M)1 \times u(1\text{M}) — I will write M for a million, because writing zeros is tedious. Suppose the person picked A over B, so that is greater than the expected utility of B:

u(1M)  >  0.89 u(1M)+0.10 u(5M)+0.01 u(0)u(1\text{M}) \;>\; 0.89\,u(1\text{M}) + 0.10\,u(5\text{M}) + 0.01\,u(0)

A brief note: you can put a + w+\,w inside every one of those, and that is the more accurate way of writing it. I am not doing it because I will run out of board space, and because I know it will not affect the conclusion. You will see why in a moment.

Before writing anything else down, do some high-school algebra. There is one u(1M)u(1\text{M}) on the left and 0.89 u(1M)0.89\,u(1\text{M}) on the right, so subtract:

0.11 u(1M)  >  0.10 u(5M)+0.01 u(0)0.11\,u(1\text{M}) \;>\; 0.10\,u(5\text{M}) + 0.01\,u(0)

Everybody see what I did? Just subtract. Now the second question, where the person picked D over C. Why am I even looking at the board? I know these from memory.

0.90 u(0)+0.10 u(5M)  >  0.11 u(1M)+0.89 u(0)0.90\,u(0) + 0.10\,u(5\text{M}) \;>\; 0.11\,u(1\text{M}) + 0.89\,u(0)

Same trick. There is 0.89 u(0)0.89\,u(0) on both sides, so move it across:

0.01 u(0)+0.10 u(5M)  >  0.11 u(1M)0.01\,u(0) + 0.10\,u(5\text{M}) \;>\; 0.11\,u(1\text{M})

Now stare at the two boxes. Does everybody see the contradiction? If not, check your glasses. The first says one quantity is bigger; the second says the same quantity is smaller.

These choices are a mixture. They are the orange trick from earlier — they are the independence axiom, and that is where the violation comes from.

It can also be revealed by simply putting the omitted numbers back in. A million for sure can be written as a million with probability 0.89 and a million with probability 0.11. A stupid way to write it, but a legal one. Analogously, the 90% chance of nothing in option D can be broken into an 0.89 chance of nothing and a 0.01 chance of nothing.

The four Allais options drawn as probability bars. Within each question the leftmost 0.89 block is identical across the two options; the remaining 0.11 of probability is identical across the two questions. Question 1 $1M · 0.89 $1M A $1M · 0.89 $5M B Question 2 $0 · 0.89 $1M C $0 · 0.89 $5M D 0.89 of the probability the other 0.11
Fig. 2. The same four options, with the omitted numbers written back in. Within a question, both options hand you the identical block on the left — a million in question one, nothing in question two — so it cannot bear on which one you take. That leaves the 0.11 on the right, which poses the same decision both times. Pick A in the first question and you have already picked C in the second.

And notice: these are literally the same question, twice.

The Ellsberg paradox

I have one more paradox to describe. Or two. Paradoxi? Paradoxes. Paradactyls. That is how it is.

The next is courtesy of Daniel Ellsberg, a national hero — the man who published the Pentagon Papers — who also dabbled in behavioral economics. So there is hope for me yet.

The randomizing device of choice in economics is an urn: an object you cannot see through, containing coloured balls, used to generate random outcomes. Grandma may also be in there. Who knows.

So we have 90 balls, plus grandma. Thirty are red. The other sixty are black and yellow — black and yellow, black and yellow — in unknown proportions. I am going to draw one ball at random.

Question one. Option A: you win $100 if the ball is red. Option B: you win $100 if the ball is black. Hands for A, hands for B.

Question two. Option C: you win $100 if the ball is either red or yellow. Option D: you win $100 if it is either black or yellow. Hands for C, hands for D.

Got you all that time.

The pattern of A over B together with D over C violates expected utility. In this case it means you are using nonsense probabilities. We can do the mathematics, but let us use our eyeballs instead, which is a fun alternative.

The urn's three colours with their counts, and which colours pay under each of the four options; choosing A over B implies fewer than thirty black balls, while choosing D over C implies more than thirty. red black yellow 30 ? ? 60 between them, split unknown A pays on red B pays on black C pays on red or yellow D pays on black or yellow A over B ⟹ black < 30 D over C ⟹ black > 30
Fig. 3. Both questions turn on the same unknown, the number of black balls. Preferring A to B backs thirty red against the black ones, which puts black under thirty. Preferring D to C adds yellow to both sides, where it cancels, leaving black against red again — only now you have backed black. The first choice needs fewer than thirty black balls in the urn, the second needs more than thirty, and nobody has touched the urn in between.

There are thirty red. If I choose A over B, that means I think there are fewer than thirty black. I have to. But if there are fewer than thirty black, then there are more than thirty yellow, and I should definitely pick red-or-yellow over black-or-yellow — because that sum is thirty plus a number bigger than thirty.

What is being revealed here is often called ambiguity aversion. Ambiguity is a circumstance of this kind, easy to describe in words and rather hard to put in mathematics: there are sixty balls in unknown proportions. It is ambiguous. And whatever mental model you hold, you cannot think that the number of balls is changing between the two questions.

For those who succumbed, you may tell yourself sweet lies about what you were thinking. You were not. But it is fine. This is very common and very easy to catch people on. You just did it.

Kahneman and Tversky, to be continued

One final one — or the start of one, which we will pick up on Thursday. This is courtesy of Danny Kahneman and Amos Tversky.

First, some details. The evidence comes from hypothetical choices, which one normally takes with a grain of salt, except that this has been replicated seven million times. It may be the most replicated result in economics. They asked students and faculty to respond to a series of binary choices, no more than a dozen problems per questionnaire, varying the order of questions and the position of the options.

Critically, their notation drops zero-dollar outcomes. So “(4000, 0.8)” means $4,000 with probability 0.8 and nothing with probability 0.2.

These numbers come from the original paper. Here is their problem three. Option A: an 80% chance of $4,000, and correspondingly a 20% chance of nothing. Option B: $3,000 with certainty. Hands for A, hands for B.

And problem four. Option C: a 20% chance of $4,000. Option D: a 25% chance of $3,000, and otherwise nothing. Hands for C, hands for D.

Those who picked B and then C — you know who you are — exhibited the pattern.

On Thursday I will explain why. But people always ask what they can do in this class. You can do the things I tell you to do; that is a good start. Here is a better one. You tell me. It is a simple question:

Why is that pattern a violation of expected utility theory?

Before Thursday

Office hours begin momentarily. Keep the timing in mind for the problem set: there are office hours today, and there are office hours on Thursday. There are no office hours on Monday, when you will want them. Plan accordingly.

Footnotes

  1. One gap worth closing, since I waved at it in class. At ρ=1\rho = 1 the formula reads x0/0x^{0}/0, which is not a number, and yet we cheerfully say that CRRA “is” log utility there. In class I invoked L’Hôpital’s rule — dredging up the old stuff in your brain — and moved on. The careful version is nicer, because it leans on something else we do today. Subtract a constant first: x1−ρ−11−ρ\frac{x^{1-\rho} - 1}{1-\rho} differs from x1−ρ1−ρ\frac{x^{1-\rho}}{1-\rho} by the constant 11−ρ\frac{1}{1-\rho}, which makes it a positive affine transformation — and by the result later in this lecture, it therefore represents the same preferences. Now let ρ→1\rho \to 1 in that version and L’Hôpital hands you ln⁡x\ln x. If you do not remember L’Hôpital’s rule, it will not come up again. If you do, you may feel cool for a moment. ↩