Fabelwaren

How the thinking test measures and where the limits are

What the test measures

Rounds of 11 puzzles in eight kinds: pattern matrices, picture matrices (Sandia), number series, odd-one-out tasks, a memory matrix (visual working memory), figure weights (quantitative reasoning without numbers), block counting (spatial visualisation) and three language puzzles in the respective language. The focus is on what research calls “fluid intelligence”: spotting rules in new material. The language-free part is the main value and runs identically in every language; the language section gets a band of its own.

It does NOT measure: knowledge, vocabulary, creativity, practical wisdom, social intelligence. Thinking is much broader than any online test. Memory and attention are only touched on (6 puzzles each) and do NOT yield a robust score on their own.

Why a range instead of a number

Every test score carries measurement error: daily form, tiredness, practice, lucky guesses. Serious assessment therefore works with confidence intervals. An online test selling you “IQ 128” as an exact number hides this uncertainty. We show it.

What the calibration phase means

This test is new and being calibrated. For now your placement rests on a provisional, uncalibrated model: it shows your hit rate corrected for lucky guessing with a statistical uncertainty range (Wilson interval), not a comparison against a sample. We say so right on the result. Every anonymous run grows the data base; once it carries, we switch to empirical percentile ranks and re-anchor all results for free. Online samples remain self-selected. We will state that too.

Live values from the runs so far

Loading values …

Rounds without a fixed end — and why you can only stop between them

The test runs in rounds of 11 puzzles, as many as you like (at most 30). All rounds follow the same template: the same puzzle kinds, the same difficulty mix, each starting easier and ending harder. That is why you can stop at the end of any round: what you worked through is a fair miniature of the test. Mid-round stopping is deliberately blocked — you would skip the hard puzzles at the round's end. From three rounds on your run is comparable with everyone else's.

The uncertainty range narrows with every round: the error of a proportion falls with the square root of the number of puzzles. The round card therefore states the width now and the width after one more round — you decide when it is sharp enough for you. The “meaningfulness” bar is scaled to eight rounds.

Fairness & limits

  • No speed score: the generous per-puzzle limits only discourage outside help. Speed never enters your result. This keeps the test fair on phone and computer.
  • Unsupervised means estimate: nobody checks who takes the test and how. For robust scores there is professional, supervised assessment.
  • Not a clinical instrument, no diagnosis, not for hiring or admissions. And no judgement about a person.
  • Puzzles are colour-vision-friendly (shape and pattern carry the information, never colour).

Where the puzzle kinds come from

Our own generator produces every puzzle; no item comes from another test. The puzzle kinds have models we name: figure weights follows the “Figure Weights” type, which shows the highest loading on the general factor in WAIS-IV research; block counting follows the public-domain Army General Classification Test (1940) and the Army Beta test (1917); matrices follow the Carpenter–Just–Shell rule family. Since 2 Sept 2026 (Form C) the quick look is gone — in our own data it barely loaded on what the test is meant to measure — and puzzle 24 of the old form is discarded because it was solved below chance level.

Language puzzles are never translated but written separately for each language and blind-checked by a foreign model (options shuffled, no answer hint; flagged items replaced). All eight languages have a verified list. The picture matrices come from the Sandia Matrices (Matzen et al. 2010, BSD licence): 30 of 840 published items, used with their normed solutions — we name the source because the licence requires it and because it is true.

The number on request

The result is a range first, and that stays. Whoever explicitly wants a number gets one on request — honestly placed, in two stages: first we show the position among all full runs of this test so far (“better than about X%”), as a point with an uncertainty band on a curve. Once about 300 full runs have accumulated, a provisional value on the IQ scale (mean 100, spread 15) is added — derived from the position within our own participant pool, never without a band, never as a classification.

Two limits are stated explicitly: our participant pool is not a representative population sample — people who volunteer for reasoning puzzles sit above the population average, so the comparison tends to be strict. And only the language-free part of runs with at least three rounds is compared — shorter runs are too noisy, and the language section has a separate puzzle list per language.

The business model, stated openly

Free result, detailed breakdown for a one-time 4.99 €. No subscription, no ads, no data sharing. Purchased breakdowns update for free with every norm improvement.