About this site

How Our Test Works: Questions, Scoring and Limits

On this page
  1. What does our test measure?
  2. What does our test not measure?
  3. How are the questions made?
  4. How is each question checked?
  5. How is your score calculated?
  6. Why do we show a range?
  7. How precise is the test right now?
  8. What are our norms, and how will they change?
  9. Who takes online tests, and why does it matter?
  10. What numbers do we publish about test takers?
  11. What data do we keep?
  12. What has changed?
  13. How can you report a faulty question?

Our IQ test is a short online reasoning test. It uses generated questions, a published scoring model and a visible range around every result. This page explains how the questions are made and checked. It also shows how the score is computed, how precise it is today, and what it cannot do. Every number below is computed from the live code or data, not typed by hand.

What does our test measure?

Our test measures reasoning on five kinds of puzzle. They are matrices, number and letter series, mental rotation, logic and simple math. It needs no special knowledge, and every question can be solved from what is on the screen.

One test has 12 matrix questions, 6 series, 4 rotation, 4 logic and 4 math questions. Matrix puzzles get the largest share. Research on matrix tests found two main differences in higher scorers. They find abstract rules more easily, and they keep more goals in mind at once.

What does our test not measure?

Our test is not a clinical assessment and not a diagnosis. A full clinical test also measures vocabulary, memory and processing speed. Ours does not. It cannot qualify anyone for Mensa or any other society, because those accept only supervised tests.

It also cannot measure the ends of the scale precisely. We print no number below 70 or above 145. Results beyond those points show as “below 70” or “145 or higher”. A range that reaches past one end shows only its other end, for example “up to 84” or “128 or higher”.

How are the questions made?

The questions are made by software for each attempt, not taken from a fixed bank. Each kind of puzzle has a generator. It picks rules, builds the question, and builds wrong options from common mistakes. A number called a seed makes each question reproducible. So we can rebuild any question later to check it.

This approach is not new. Research has shown that software can generate large numbers of matrix puzzles from rule settings. In one study, the generated items covered and extended the difficulty range of the original test.

Generated questions have two benefits. Every attempt is different, so answers cannot be memorized or shared. And the practice set on our IQ test questions page uses separate seeds. Those never appear in a real test.

Difficulty rises with the number and kind of rules in a question. That is our design choice as a starting point. The real difficulty of each rule is then measured from real answers, as described below.

How is each question checked?

Each question is checked by a separate piece of software, a validator, before anyone sees it. The validator works out the answer again from scratch. It rejects the question unless exactly one option is correct.

For matrices and series, it also tries every simple rule in its library. It checks across rows and down columns. If a second rule could explain a different answer, the question is thrown away. For logic, it checks every possible arrangement of the groups or every possible order of the people. For math, it recomputes the answer from the stored numbers.

Every change to the generators is tested on thousands of new questions, with zero failures allowed. We also review samples by eye for clarity, overlap and contrast. When we find a problem, we fix the generator.

How is your score calculated?

Your score is calculated from which questions you answered correctly and how hard each one is. We use an item response model, a standard way to link answers to an underlying level.

How our score is madeFour steps: your answers, then the ability level that best fits them, then an IQ of 100 plus 15 times that level with a 90 percent range, then the percentile.1Your answerswhichquestions yougot right2Ability levelthe levelthat bestfits thoseanswers3IQ and range100 + 15 ×level, with a90% range4Percentilethe share ofpeople whoscore lower
How our score is madeFour steps: your answers, then the ability level that best fits them, then an IQ of 100 plus 15 times that level with a 90 percent range, then the percentile.1Your answerswhich questionsyou got right2Ability levelthe level thatbest fits thoseanswers3IQ and range100 + 15 × level,with a 90% range4Percentilethe share ofpeople who scorelower

For each question, the model gives the chance of a correct answer at each level of ability. It allows for guessing: with five options, a pure guess is right one time in five.

P(correct) = c + (1 - c) / (1 + exp(-a × (θ - b)))

Here θ (theta) is your level and b is the question’s difficulty. The letter a is how sharply the question separates people. And c is the guessing chance, 1 divided by the number of options.

Difficulty comes from the question’s features, such as how many rules it has. This follows the idea of the linear logistic test model. In it, a question’s difficulty is the sum of the difficulties of its parts.

Your level is then estimated by a method called EAP, short for expected a posteriori. It combines your answers with a starting assumption that most people are near the middle. Its uncertainty closely matches the standard error of measurement. Finally, the level becomes an IQ:

IQ = 100 + 15 × θ

The same math is used on the result screen, on our score pages and in the IQ calculator. What each score range means is shown on the IQ scale.

Why do we show a range?

We show a range because every score has measurement error, and hiding it would be dishonest. The standard error of measurement is how much a person’s score would vary over many sittings.

A result with its 90% rangeA number line from 70 to 145. The estimate 115 is marked with a dot, and the shaded band from 103 to 127 shows the 90% range around it.7085100115130145Estimate: 11590% range: 103 to 127
A result with its 90% rangeA number line from 70 to 145. The estimate 115 is marked with a dot, and the shaded band from 103 to 127 shows the 90% range around it.7085100115130145Estimate: 11590% range: 103 to 127

Our range is a 90% interval: the estimate plus or minus 1.645 standard errors. On our test today, a result of 100 usually comes with a range of about 88 to 112. A result of 115 usually comes with about 103 to 127, and 130 with about 117 to 143.

These ranges are wider than on a full clinical test, which uses many more questions. That is the honest cost of a short test.

A result screen from our test: the estimated IQ with its label and percentile, its range shaded on a bell curve, and a bar for each kind of question.
Every result comes with its range, shaded on the curve under the number.

How precise is the test right now?

The test is most precise in the middle of the scale and least precise at the ends. We measure this by simulation. We give the test to simulated people whose ability we set in advance. Then we compare each result with that set ability.

How well the current scoring recovers known levels (400 simulated people per row)
Simulated ability (IQ)Average result shownDifference from the simulated ability (points)Typical error (points)90% range contains the simulated ability
70.078.78.710.678%
77.583.05.58.687%
85.088.03.07.591%
92.594.31.86.695%
100.099.2-0.86.694%
107.5105.1-2.47.094%
115.0112.0-3.07.292%
122.5117.8-4.78.089%
130.0122.8-7.29.481%

The table shows three things. First, between about 85 and 120, results land close to the simulated ability. There the range contains it about nine times in ten, as it should. Second, toward the ends, results are pulled toward 100. A positive difference means the result came out higher than the simulated ability; a negative one, lower. In our method the estimate leans on the starting assumption when answers give little information. A short test gives little information at the ends. Third, at the ends, the range contains the simulated ability less often than it claims.

In practice this means a perfect answer sheet shows about 140 today, not higher. We say this on every result. These numbers will be replaced once the test has been calibrated on real test takers.

What are our norms, and how will they change?

Our norms are provisional. Today the scale comes from a statistical model, not from a sample of real people. The norm version is “provisional”. Every result shows its norm version.

Current norm sample size: none yet (provisional model). Last calibration: not yet calibrated.

Calibration will use real answers from adults on their first attempt. They must finish the test in a normal time. Feature weights can be re-estimated once 1,000 such attempts exist. A new norm replaces “provisional” only after 3,000 adult first attempts. It must also pass checks for fit and stability.

Old results keep their norm version. A shared result never changes after the fact.

Who takes online tests, and why does it matter?

People who take online tests choose to take them. So they are not a random sample of the population. Research on web surveys shows that such self-selection can make results unreliable as a picture of everyone.

Unsupervised online tests also score a little higher than supervised ones on average. A review of 49 studies found a gap of about 3 IQ points. And retaking tests raises scores a little. That is why we ask whether you have taken the test before.

We do not adjust our norms to correct for who takes the test. Instead, we disclose it here. Any future adjustment will be documented on this page with its method and its effect.

What numbers do we publish about test takers?

We publish counts only from our real database. We also wait until they are large enough to mean something. Before 1,000 completed attempts, we show no counts at all.

Completed attempts so far: not shown yet (below the publishing threshold). Adult first attempts: not shown yet.

What data do we keep?

We keep only what the scoring and calibration need, and nothing that identifies you. We store your answers, the questions’ seeds and features, and the time per question. We also store your score and range, the day (not the time), the kind of device and the country code. If you choose an age group, we store that band; it does not change your score. If you answer the optional question after your result, we store whether you have taken the test before.

Our database never stores IP addresses, device fingerprints, names or email addresses. Our hosting provider, Cloudflare, does process your IP address to deliver the site and protect it, as every web host must. No account is needed. If you say you are under 13, nothing from your test is sent to us or stored. Raw answers are deleted after 24 months; only anonymous totals are kept. Our privacy policy has the full details.

What has changed?

Every change to the questions, the scoring or the norms is listed here, newest first.

DateWhat changed
September 26, 2026Matrix generator re-draws any pattern that a second rule could also explain, so every matrix question has exactly one answer.
September 26, 2026Question quality review of generated items by eye: clearer explanations, lowest-term ratios, stronger contrast in figures, and a real "cannot be worked out" option in ordering puzzles.
September 26, 2026Test launched in preview with provisional norms: item difficulty from feature weights, not yet calibrated on people.

How can you report a faulty question?

If a question looks wrong or unclear, please tell us through the contact page. Describe the question and the answer you expected. Every question can be rebuilt from its seed. So we can check it exactly and fix the generator for everyone.

For how we research and correct our articles, see our editorial policy. Who runs the site is explained on our about page. To see what the numbers mean for you, read are IQ tests accurate. Or try the free test yourself.