Influence
Part 2  Reading the Body
Chapter 56 of 360

Behavior Detection at the Airport: The SPOT Program Failure

The Transportation Security Administration's SPOT program — Screening of Passengers by Observation Techniques — put behavior detection officers in American airports to identify travelers displaying signs of stress, fear or deception, on the theory that mal-intent produces observable behavioral indicators. Officers scored passengers against a checklist of behaviors and referred high scorers for additional screening. Roughly three thousand officers were deployed at its peak. Around nine hundred million dollars had been spent by the time the Government Accountability Office reported on it in November 2013, with total program costs since inception exceeding a billion.

The GAO's conclusion was that the underlying scientific premise was unvalidated, that TSA's own evidence did not demonstrate effectiveness, and that Congress should consider limiting future funding.

The report is a useful document because it did not simply say the program failed. It said why it could not have succeeded, and the reasoning has two independent parts.

The first is the cue-validity problem this part of the guide has been building toward. GAO reviewed the meta-analytic literature, including Bond and DePaulo, and found that human ability to detect deception from behavioral cues is at or very near chance. A screening system built on a 54 percent instrument cannot produce reliable output regardless of how well it is administered. Internal analysis also found that referral rates varied enormously between airports and between individual officers, which is what you expect when officers are applying subjective criteria to noise.

The second is the base rate problem, and it is the more general lesson.

Consider a screen that is genuinely good — say, 90 percent sensitive and 90 percent specific, far beyond anything this program could achieve. Apply it to a population where the prevalence of the thing you are looking for is extremely low, as it is for terrorists among airline passengers. Out of every million passengers, the screen flags roughly a hundred thousand innocent people, alongside whatever tiny number of genuine threats exist. The overwhelming majority of every alarm is false, not because the test is bad but because the arithmetic of rare events is unforgiving. This is Bayes' theorem, and it applies to every low-prevalence screening program ever built, including several in Part 4 of this guide.

The consequences were predictable and documented. Complaints and internal accounts alleged that officers, working from ambiguous criteria under pressure to produce referrals, referred passengers on the basis of ethnicity and national origin — which is precisely what happens when a scoring system has no valid signal in it. Something has to drive the decisions, and in the absence of signal, it will be whatever priors the officer brought to work.

There is a broader point here about institutions rather than individuals. Nobody in this story was stupid. Behavior detection is intuitively compelling, it is easy to explain to a legislature, it produces visible activity, and it is nearly impossible to falsify from inside — an officer who refers a hundred innocent people and no terrorists has no way to learn they are performing at chance, because there were no terrorists to miss.

Which is the same feedback vacuum as chapter 52, scaled up to a billion dollars.

The case

The US Transportation Security Administration’s SPOT (Screening of Passengers by Observation Techniques) program and the Government Accountability Office’s 2013 report GAO-14-159, which found the underlying science unvalidated and recommended limiting future funding after roughly a billion dollars of spending.

The mechanism

SPOT rested on the premise that trained officers can identify mal-intent from stress and fear indicators; GAO’s review of the meta-analytic literature concluded human ability to detect deception from behavioral cues is at or near chance, and TSA’s own data did not demonstrate effectiveness. The deeper flaw is base rates — with an extremely low prevalence of actual threats, even a modestly inaccurate screen generates overwhelming false positives and discriminatory outcomes. This is the definitive institutional case study in over-trusting nonverbal reading.

What this chapter covers

  1. Origin of behavior-detection screening programs
  2. Chance-level cue validity and base-rate failure
  3. GAO’s 2013 evaluation of TSA SPOT
  4. Aviation security, border control, and profiling
  5. Stress indicators mistaken for hostile intent

See Defense: base rates before behavioral judgment