There is no average driver

There’s a tendency to think of averages as being “fair”. If we describe people in terms of averages, then we are somehow more likely to do them justice (or less likely to be unjust). That’s not the case.

There’s a tendency to think of averages as being “fair”. If we describe people in terms of averages, then we are somehow more likely to do them justice (or less likely to be unjust). That’s not the case. An average feels like the fair, rigorous, data-driven summary — the responsible thing to reach for — and that feeling is the whole problem. Averages are sneaky things. They decide which crowd you belong to, and then they flatten that crowd into one number that describes no one actually in it. Then they dress themselves in a costume called Objectivity and hope no one asks to see ID.

Averages are on my mind after my jury duty wrapped up this week. It was a personal injury case. Car accident. The resonant moment for me was listening to an expert witness who explained that, in his assessment, he calculated the defendant’s rate of acceleration from a stopped position into the intersection using the average rate of acceleration for most people in such circumstances.

The number the witness was using was correct, but only for the group it was drawn from. The problem is that the average describes a crowd, and the defendant wasn’t a crowd. He was one person, and as one person, he belonged to many crowds all at once. It’s true that he was a driver — he’s part of that crowd — but he was also an 80-year-old. And he’s the driver of a particular car, with particular cognitive and physical capabilities (that are not, themselves fixed!). The expert ignored the pile of possible groups, picked the broadest possible one, and called the result a measurement. But what he really did was make a judgment about which crowd we should accept the defendant as being a part of. The defense was either unaware of that or was hoping the jury wouldn’t think twice about it.

It was difficult to sit through the testimony. I tried to Jedi-mind-trick my way into the plaintiff’s attorney’s head to get him to answer the questions I wanted answered: Why choose all drivers as the referent? If you’re going to use an average, wouldn’t it be more accurate to use the average rate of acceleration of a more specific reference class — for example, the average acceleration among 80-year-old men?

The Force failed me. The witness decamped, and I sank back in my chair. Still, we can imagine how it might have played out if I had been able to pull an Obi-Wan.

The defense might have objected to a biased line of questioning. Reasoning about the defendant as “an 80-year-old” is stereotyping. And he’d be right. It’s not fair or just to apply that lens to the case. But notice the bind that creates. The broad group — most drivers — is valid, stable, and nearly useless for judging this one man. At the same time, the narrower group — men of his age — is far more relevant to what actually happened, but that added relevance is what introduces prejudice. So we’d have to narrow again: drivers of that particular car, people with a similar cognitive profile, individuals with similar motor function, the visibility at that intersection on that day … and every step that makes the class more relevant also makes it more his, until the only group precise enough to be truly fair is a group of one, which is no longer an average. There is no neutral place to stop. Each stopping point is a decision about who this man is comparable to, and the data never makes that decision for you.

But suppose we could agree on the right class. The average would still fail us, because no one is the average. Take the group we finally settled on — drivers of his age. Some react faster than the typical 30-year-old; some shouldn’t be behind the wheel at all. The mean sits in the middle of that range and describes almost none of them. Sam Savage calls this the flaw of averages. He argues that plans built on average conditions are, on average, wrong. So the average asks us to concede that the crowd this person has been assigned to is somehow the correct one (Goldilocks style — not too broad, not too narrow, just right) and then that that crowd’s midpoint can and should stand in for the person in front of you. In this case, the person on trial.

The law already recognizes this. The rules of evidence exclude this sort of reasoning — things like character, propensity, and demographic generalization. Even when any of those things would genuinely predict something, the court dismisses them because judging a person by the group they belong to isn’t and shouldn’t be the same thing as judging their actions. Our legal system has deliberately decided that relevance is not enough. It says that fairness takes precedence. Museums have made no such decision.

Take Membership as an example. A museum has limited resources for its renewal campaign, so the membership department uses a model to determine how to allocate those resources. The model scores each lapsing member on the likelihood of renewal using a segmentation approach. For example, when they joined or how many times they visited. Jane, who is approaching the end of her first year of membership, is identified as “low propensity” — first-year members renew at far lower rates than long-tenured ones, so the model scores nearly every one of them low before it knows anything else about her — and she gets an automated email instead of a personal phone call. Jane doesn’t renew. The model was right.

But the model never met Jane. It sorted her into a crowd and let that crowd stand in for her decision before she had made it. Jane isn’t the average of her cohort; she’s one person it happened to contain. And because the model flagged her group as unlikely to renew, the museum invested less in reaching her. Not investing in a group is, of course, a good way to lower how many of them come back. The prediction was self-fulfilling, and it then took credit for being right. The model decided who Jane was, so that the museum wouldn’t have to.

Every time we segment an audience by fixed characteristics, we choose which crowd a person belongs to and let that choice speak for them. Sometimes that’s fine. As George Box said, “All models are wrong, some are useful”. The useful ones tend to be the ones built by listening to real individuals, though. When a model is built from what people tell us, they’ve played an active role in defining their own group, and a good model lets them move between groups as they change. That kind of classification is one that the person helped author.

The trouble is that most models don’t work that way. They run on values that were assigned from outside rather than co-constructed by those the group represents. That’s the line the courtroom draws: a person may be judged by what they did, because they authored the action, but not by the group someone else filed them under.

Before you let an average or a segment decide something about a person, ask whether the people in this group have any hand in forming it, and whether they move out of it as they change. A museum can build its segments alongside the people those categories represent, rather than imposing them from the outside.

Too often, we reach for models because they project rigor. But in trying to soothe our uncertainty, we act on evidence that we’d deem ethically unsound in other contexts.

Have a good weekend,

Kyle