You tell a friend you are pretty sure it will rain. They say they are pretty sure too. You might mean nine times out of ten. They might mean six. Neither of you can tell from the words. And when tomorrow comes, nobody can check who was closer. The whole exchange leaves no trace.
Sources
Fischhoff, B., Slovic, P., & Lichtenstein, S. (1977). Knowing with certainty: The appropriateness of extreme confidence. Journal of Experimental Psychology: Human Perception and Performance, 3(4), 552-564. Full PDF text-extracted and read this pass. Verified in the source text: the quote block is the paper's opening sentence, verbatim; 73% correct at 100:1, 81% at 1,000:1, 87% at 10,000:1, 90% at 1,000,000:1 or greater with 9:1 the appropriate odds; 'Of 6,996 odds judgments, 3,560 (51%) were greater than 50:1'; 'Almost one fourth of the responses were greater than 1,000:1'; subjects 'reasonably well calibrated' at 1:1 through 3:1 and 'little or no increase in accuracy' from 3:1 to 100:1; Experiment 3's 20-minute lecture, 74% correct at 50:1, odds of about 3:1 appropriate, extreme odds on 'approximately one third of the items', 'instruction tempered subjects' extreme overconfidence, but only to a limited extent'; Experiment 5's '13 participating subjects missed 46 of the 387 answers (11.9%)' at 50:1 or greater, out of 19 who were asked; Table 4's median of 4 cases of extreme overconfidence in Experiment 4 and 'removing those two subjects had no effect on our conclusions'; paid volunteers recruited by University of Oregon student-newspaper ad; five experiments. https://nuovoeutile.it/wp-content/uploads/2014/10/Knowing-with-certainty.pdf
Murphy, A. H., & Winkler, R. L. (1977). Can weather forecasters formulate reliable probability forecasts of precipitation and temperatures? National Weather Digest, 2(2), 2-9. Full PDF text-extracted and read this pass. Verified: 'Since 1965, National Weather Service (NWS) forecasters have expressed their forecasts of precipitation occurrence in probabilistic terms'; 'several million PoP forecasts have been formulated and disseminated to the general public during the last ten years'; Chicago WSFO July 1972 through June 1976, 'a total of 17,514 forecasts'; 'forecasts of 30% were issued on 1574 occasions during the period, on 449 (28.5%) of which measurable precipitation actually occurred' — the unit is OCCASIONS, and 'PoP forecasts are usually issued three or four times a day'; PoP defined as 'measurable precipitation (i.e., > 0.01 inches) will occur during a specified period at a particular point in the forecast area (generally the official raingage)'; 'the reliability of these subjective precipitation probability forecasts was excellent'; 'a slight tendency did exist for the forecasters to overforecast ... for most probability values'; the Brier score described as 'simply the mean square error of the forecasts', the single-gap version the module teaches. https://nwafiles.nwas.org/digest/papers/1977/Vol02No2/1977v002no02-MurphyWinkler.pdf
Mellers, B., Ungar, L., Baron, J., Ramos, J., Gurcay, B., Fincher, K., Scott, S. E., Moore, D., Atanasov, P., Swift, S. A., Murray, T., Stone, E., & Tetlock, P. E. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5). Full PDF text-extracted and read this pass, including Table 1. Verified: IARPA-funded, tournament Sept 2011-April 2013, 199 geopolitical outcomes (85 questions closed in Year 1, 114 in Year 2); Year 1 began with 2,246 participants, Year 2 with 1,648; entry required 'a bachelor's degree or higher and completion of a battery of psychological and political tests that took an average of 2 hr', participants 76% US citizens, 83% male, average age 36; training 'took approximately 45 min, and the modules could be reexamined throughout the tournament'; probability training 'guided forecasters to consider reference classes; average multiple estimates from existing models, polls, and expert panels; extrapolate over time ...; and avoid judgmental traps such as overconfidence, the confirmation bias, and base-rate neglect'; Table 1 Year 1 individual forecasters no training 0.44/0.40/0.31 and probability training 0.40/0.36/0.29, Year 2 individual no training 0.46/0.39/0.26 and superforecasters 0.25/0.19/0.07; superforecasters were the top 2% of Year 1 placed in 'five teams of 12 members each'; superforecasters 'made an average of 7.8 predictions per question' against 1.4 for independent forecasters; 'Brier scores can be decomposed into three additive parts—variability, calibration, and resolution', variability being 'a function of the base rate for events ... question difficulty rather than skill'; 'superforecasters' accuracy was in large part due to their greater resolution' and they were also better calibrated. CRITICAL for scale: 'This score measures individual accuracy; 0 is the best score, and 2 is the worst score', computed as (.9-1)^2 + (.1-0)^2 — the summed, doubled version. https://sydneyscott.nfshost.com/pubs/Psychological_Strategies_for_Winning_a_G.pdf
Chang, W., Chen, E., Mellers, B., & Tetlock, P. (2016). Developing expert political judgment: The impact of training and practice on judgmental accuracy in geopolitical forecasting tournaments. Judgment and Decision Making, 11(5), 509-526. Open-access article read this pass. Verified verbatim: 'Although the training lasted less than one hour, it consistently improved accuracy (Brier scores) by 6 to 11% over the control condition' — this is the paper's own summary sentence and is what the module quotes; per-year figures were 10% and 11% (Year 1), 12% (Year 2), 6% (Year 3), 7% (Year 4). Also verified: 'training improved the calibration and resolution of forecasters by reducing overconfidence'; four tournament years, Sept 2011-May 2015; the CHAMPS KNOW checklist (comparison classes, hunt for information, adjust/update, models, post-mortems, effort level). https://dlab.sauder.ubc.ca/sjdm/journal/16/16511/jdm16511.html
Meyer, A. N. D., Payne, V. L., Meeks, D. W., Rao, R., & Singh, H. (2013). Physicians' diagnostic accuracy, confidence, and resource requests: A vignette study. JAMA Internal Medicine, 173(21), 1952-1958. Article page read this pass. Verified: 118 physicians, 4 vignettes, 2 easier (difficulty 3.2 and 3.7 of 7) and 2 more difficult (6.0 and 5.2); 'Physicians correctly diagnosed 55.3% of easier and 5.8% of more difficult cases'; confidence 7.2 vs 6.4 on a 0-10 scale, a difference the authors call small and 'likely clinically insignificant'; 'Higher confidence was related to decreased requests for additional diagnostic tests (P = .01)'; 'Diagnostic accuracy was not significantly related to the request of any type of resource'; stated limitation, 'Our methods of case delivery limit real-world validity of the study'. https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/1731967
Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502-517. Full PDF text-extracted and read this pass. Verified: the three-way split into overestimation, overplacement and overprecision; 'on difficult tasks, people overestimate their actual performances but also mistakenly believe that they are worse than others; on easy tasks, people underestimate their actual performances but mistakenly believe' they are better; '90% confidence intervals contain the correct answer less than 50% of the time (Alpert & Raiffa, 1982; Klayman ... 1999; Soll & Klayman, 2004)', elicited with questions like 'How long is the Nile River?'; 'Overprecision appears to be more persistent than either of the other 2 types of overconfidence'. https://healy.econ.ohio-state.edu/papers/Moore_Healy-TroubleWithOverconfidence.pdf
Mosteller, F., & Youtz, C. (1990). Quantifying probabilistic expressions. Statistical Science, 5(1), 2-34. Full PDF text-extracted and read this pass, including Table 2. Verified: 'For 20 different studies, Table 1 tabulates numerical averages of opinions on quantitative meanings of 52 qualitative probabilistic expressions'; 'One exception was possible, because it had distinctly different meanings for different people' and 'possible has a bimodal distribution ... possible is unsatisfactory as a qualitative expression'; the science-writer survey was a mail questionnaire with about a 37% response rate; Table 2 quartiles for the writers' own estimates — Possible 7.5 / 38.5 / 50.2 (IQR 42.7, the widest), Even chance 49.7 / 50.0 / 50.2 (IQR 0.5), Certain 98.7 / 99.6 / 99.8, Always 99.6 / 99.7 / 99.8. http://web.stanford.edu/~clark/1990s/Mosteller,%20F.%20_%20Youtz,%20C.%20_Quantifying%20probabilistic%20expressions_%201990.pdf
NOT READ, credited in prose only: Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1-3. Every mirror returned HTTP 403 on this pass too. The attribution rests on Murphy & Winkler (1977) and Mellers et al. (2014), both read in full, both of which name it. Note the two papers use the score at DIFFERENT SCALES — Murphy & Winkler as mean squared error (worst = 1), Mellers et al. summed over both outcomes (worst = 2). The module now teaches the first and states the second explicitly.
Shahmaran catalogue, checked by SQL against the live database this pass, to place the two cross-references and rule out re-teaching: 'Decision frameworks' already teaches resulting (with the Annie Duke, Thinking in Bets (2018) quote block); 'Cognitive biases' teaches hindsight bias by name in two sections. No published module contains calibration in the forecasting sense, the Brier score, or superforecasters, and no module title covers bets, probability, forecasting, confidence or certainty.