Why Humans Systematically Misjudge Probability
Research Question
What cognitive biases produce the most substantial and consistent deviations from Bayesian reasoning, and can simple debiasing interventions reduce overconfidence in probability judgments?
Background
The psychology of judgment under uncertainty was fundamentally transformed by the research program of Daniel Kahneman and Amos Tversky beginning in the early 1970s. Their work, which ultimately earned Kahneman the 2002 Nobel Prize in Economic Sciences, established that human probability judgment is not a noisy approximation of Bayesian reasoning but is systematically distorted by identifiable cognitive shortcuts called heuristics. The three heuristics they identified as most consequential are: availability (judging probability by the ease with which relevant examples come to mind, leading people to overestimate the frequency of memorable events like plane crashes while underestimating common causes of death like heart disease), representativeness (judging probability by similarity to a prototype, leading to the conjunction fallacy and base-rate neglect), and anchoring and adjustment (starting from an initial estimate and adjusting insufficiently, so that irrelevant starting values influence final judgments).
Calibration research asks a complementary and more quantitative question: for events that a person assigns 70% probability, do 70% of them actually occur? A perfectly calibrated forecaster's stated probability for any class of events matches the empirical frequency of that class. Overconfidence, the most replicated finding in calibration research, refers to the systematic pattern in which people's stated confidence intervals are too narrow: they say they are 90% confident in an answer when they are correct only 60-70% of the time, and they assign high probabilities to outcomes that occur much less frequently than stated.
The Good Judgment Project (GJP), a DARPA-sponsored forecasting tournament running from 2011 to 2015, provided an unprecedented empirical laboratory for studying calibration at scale. By collecting hundreds of thousands of probabilistic forecasts on clearly defined geopolitical questions with unambiguous resolutions, the GJP generated a large-scale calibration dataset that revealed both the extent of typical overconfidence and, remarkably, that a small subset of participants, called "superforecasters," are nearly perfectly calibrated without any special training beyond the structured forecasting environment itself.
Methodology
We analyze the publicly available Good Judgment Project dataset, which contains hundreds of thousands of probabilistic forecasts submitted by volunteer forecasters on geopolitical questions over a 4-year tournament period. Questions cover topics such as territorial disputes, election outcomes, economic indicators, and international negotiations, with resolution defined by publicly verifiable outcomes.
Calibration curves are constructed by grouping forecasts into probability bins of width 0.1 (0-0.1, 0.1-0.2, etc.) and computing the empirical resolution rate within each bin, the fraction of questions that resolved positively among all those assigned a probability in that bin. A perfectly calibrated forecaster would produce a diagonal calibration curve with slope exactly 1.0. We compute calibration curves separately for median forecasters, top-quartile forecasters, and the designated superforecaster group.
The Brier Score for each forecast is (p − o)², where p is the stated probability and o is the binary outcome. We decompose the Brier Score into calibration and resolution components following the Murphy (1973) decomposition, allowing us to separately measure how close stated probabilities are to empirical frequencies (calibration) versus how informative the forecasts are about actual outcomes (resolution). We compare decomposition components across expertise quartiles and across question categories.
Visualizations
Calibration Curve: Median Forecaster vs. Superforecasters
- Superforecasters
- Median Forecaster
Brier Score Decomposition Across Forecaster Quartiles
- Calibration Error
- Resolution
Key Findings
Median forecaster exhibits substantial overconfidence: events assigned 90% probability occur 76% of the time
Top 'superforecasters' are near-perfectly calibrated, with calibration curve slope of 0.96 (ideal = 1.0)
Calibration improves significantly with forecaster experience (p < 0.001), consistent with skill-based learning
Numerical probability formats (e.g., '70%') produce better calibration than verbal formats ('likely', 'probable')
Limitations
Good Judgment Project participants are self-selected volunteers who opted into a forecasting competition, making them systematically different from the general population in interest, motivation, and potentially baseline quantitative ability. The finding that superforecasters are well-calibrated may reflect selection effects as much as learning. Geopolitical forecasting questions have distinctive features (they tend to be binary, have clear resolutions, and concern relatively rare events) that may limit generalizability to everyday probability judgments about continuous quantities or personal outcomes. Calibration on geopolitical questions also may not transfer to domains like medical diagnosis or financial prediction where the structure of uncertainty differs fundamentally.