Trust layer
How the tests calculate a result
Read how raw accuracy, weighted attainable ranges, source classes, local browser processing, and explicit evidence limits shape every result.
Every test shows its source, item count, result shape, and material limit before the first question. Questionnaire dimensions are weighted sums normalized only to each test’s attainable range. That number is not a population percentile.
Reasoning tasks
Original tasks report a MyTestAtlas score: each item counts once for the easy tier, twice for medium and three times for hard, and the weighted total is normalized to that test’s own attainable range (0–100). The long-form routes add accuracy by difficulty tier and list every item with your answer, the correct answer, and its rule family. The tiers are our estimate of difficulty, not a measured item statistic. The score compares you with the item set, not with other people.
The IQ-style estimate
Only the 36-item free IQ test prints a range on the familiar 100/15 scale, and it is an estimate from a published formula, not a measured IQ. Each item is modelled with a three-parameter logistic curve: a guessing floor of one over the number of options (0.25 for four options), a discrimination of 1, and a difficulty we set by tier — easy −1.0, medium 0.0, hard +1.2 logits. Your ability estimate is the point on that curve where the model expects exactly your number of correct answers; its standard error comes from the summed Fisher information at that point. Two anchor constants turn ability into the IQ scale: 100 sits where the model expects 55% of the items correct, and 1.15 logits equal 15 points. These two constants are our decisions, not measurements; a different pair would move the same result by several points, and the range printed is a 68% band with a minimum half-width of 7 points. The range is shown only on the first completed attempt within 26 minutes, is clamped to 80–135 because 36 items cannot resolve results beyond that, and a point estimate is never printed. This test has no representative norm sample, so the range must not be represented as an IQ score to anyone. The IQ-style range is an estimate from this published formula with author-set anchor constants, not a normed score.
The full-scale form
The 72-item full-scale test uses the same three-parameter model and a longer bank. Its five core tiers and the two routed blocks carry the difficulties, option counts and discriminations below; the guessing floor c is one over the option count, which is why the option count rises with difficulty — at the top of the scale guessing is what limits the test, and eight options carry about 27% more information per item than four.
| Tier | Items | b (logits) | Options | c | a | IQ at b |
|---|---|---|---|---|---|---|
| Foundation | 10 | -1.6 | 4 | 0.250 | 1.0 | 79 |
| Standard | 16 | -0.4 | 4 | 0.250 | 1.0 | 95 |
| Hard | 18 | 0.9 | 5 | 0.200 | 1.1 | 112 |
| Very hard | 16 | 2.3 | 6 | 0.167 | 1.2 | 130 |
| Extreme | 12 | 3.6 | 8 | 0.125 | 1.2 | 147 |
| Routed block 1 | 8 | 3.0 | 8 | 0.125 | 1.2 | 139 |
| Routed block 2 | 12 | 4.3 | 8 | 0.125 | 1.2 | 156 |
The two anchor constants are the same in kind as on the quick form and the same size of decision: 100 sits where the model expects 45% of the 72 core items correct, and 1.15 logits are treated as 15 points. The band is a 68% band with a minimum half-width of 5 points, and the wider 95% band is printed beside it on every result. Your ability is found over every item you were actually given — 72, or 92 if you were routed — while the scale itself is always anchored on the 72-item core, so a routed and an unrouted result mean the same thing. Below about 60 the test says it cannot resolve the result; above 160 it prints “above 160; this test stops here”, because four standard deviations is where the items run out and where a comparison sample would have to come from a population we do not have. A taker who is offered the extreme block and finishes without it is not printed above 145, because the core alone cannot separate a result above 145 from one at it.
What happens if we are wrong about the constants
The band’s width is computed from the information the items carry. Its position rests on two numbers we chose. Here is the same ability printed under three values of the second one. At 100 the choice changes nothing; at 155 it is worth about seven points in each direction, more than the band we print. An error of 0.3 logits in the first constant would move every printed number by about four points in the same direction, for everyone.
| Ability that we print as | if 1.00 logits were 15 points | at 1.15 (published) | if 1.30 were |
|---|---|---|---|
| 100 | 100 | 100 | 100 |
| 115 | 117 | 115 | 113 |
| 130 | 135 | 130 | 127 |
| 145 | 152 | 145 | 140 |
| 155 | 163 | 155 | 149 |
| 160 | 169 | 160 | 153 |
We publish this rather than hide it, and we will replace both constants with values estimated from real response data once enough people have finished this version, printing the old constant, the new one and the date so that anyone comparing two results across the change can see why they differ.
Timed tasks
The attention and memory tasks measure response times and spans with the browser clock on your device and report them as your own figures with the device named. Devices, screens, input methods and browsers differ, so those figures are never compared with other people, never converted into a norm or percentile, and never sent anywhere. The reasoning and questionnaire routes do not score timing at all.
Questionnaires
Positive items add their selected response value; reverse-keyed items invert that value before the dimension total is normalized between the minimum and maximum possible for this questionnaire. The adult attention route and every route under /adhd/ use the same transparent method and have no clinical cut-off or total score.
Results
Use a result as a hypothesis: choose one pattern, observe it in real situations, then decide whether the description was useful.