AtlasReasonCalibration trainer

Games · free, no sign-up

Calibration trainer

Give 90% ranges for quantities computed fresh each round, then see how many held the truth, a calibration curve, and your hits round by round.

In short

Being calibrated means that the things you are 90% sure of turn out true about 90% of the time. The trainer asks for 90% ranges on quantities it computes fresh every round, shows how many of your ranges held the true value, and keeps a running total with a calibration curve. Most people start with ranges that are far too narrow.

The trainer runs in your browser, and your browser is not running the script for this page. The method and the sources below still apply.

Nothing you type on this page leaves your browser tab: no answer, guess or score is sent anywhere or stored.

How the trainer works

Each round is five questions. For each you give a low and a high number such that you are 90% sure the truth lies between them. The questions are computed rather than looked up: counts of drawn dots, powers, square roots, compound interest, committee counts, primes below a number. So every truth is computed, not remembered or looked up, and a new round never repeats an old one.

After each round you see the truths, your hits in that round, and a running total across rounds. The running total is what matters: five questions bounce around, and twenty start to say something.

The curve and the badge

The calibration curve reads each of your ranges as the middle 90% of a bell curve on a log scale, then asks how many truths would have fallen inside at levels from 50% to 99%. If you are calibrated, the dots follow the diagonal; if they sit below it, your ranges were too narrow. That bell-curve reading is an assumption about what a range means, stated so you can weigh it.

The trainer shows a Calibrated badge once you have given at least 20 ranges in one visit and between 80% and 95% of them held the true value. It is a statement about those ranges, not about you.

Why practise this

Interval estimates are one of the most consistently overconfident judgements in the research literature: in studies such as Soll and Klayman's, ranges meant to hold the truth 80% or 90% of the time held it far less often. Training helps. In a large forecasting tournament, Mellers and colleagues found that a short training module on probabilistic reasoning improved forecasters' accuracy.

Sources

Text, questions and code on this page are original to MyTestAtlas. The cited studies describe the classic procedure; none of their items is reproduced.