How We Build and Score Our Tests

Every test on this site is free, instant and unsupervised — which makes it all the more important to be precise about what a result does and does not tell you. This page documents the whole method: what is measured, how the number is produced, and where it stops being meaningful.

Written & maintained by Christoph Ballmer · Updated August 18, 2026

1. What each test measures

Two different kinds of test live on this site, built to different specifications.

The critical thinking test

This one is structured around five reasoning skills, following the long-established structure used by critical-thinking appraisals — most recognisably the Watson–Glaser Critical Thinking Appraisal published by Pearson TalentLens, whose subtests carry these same five names. The underlying idea of critical thinking as a set of separable, teachable skills comes from the 1990 APA Delphi Report (Facione), the expert-consensus statement most later assessment work builds on. Our questions are written in-house; only the skill structure is shared.

SkillWhat the questions actually test
InferenceHow well you judge the likely truth of a conclusion drawn from stated facts.
Recognition of AssumptionsWhether you spot what an argument quietly takes for granted.
DeductionWhether you can tell what necessarily follows from what merely seems to.
InterpretationWhether a conclusion is actually warranted by the passage in front of you.
Evaluation of ArgumentsHow reliably you separate genuinely strong arguments from weak ones.

For the broader question of how to reason well, rather than how to score it, we lean on the Paul–Elder framework from the Foundation for Critical Thinking, which separates the elements of thought from the intellectual standards used to judge them. The University of Louisville maintains a clear public summary of it.

The aptitude and reasoning tests

Each of these isolates one reasoning channel — verbal, numerical, deductive, abstract, spatial or mechanical — and mirrors the question format used in graduate and job assessments. They are kept separate rather than merged because the skills genuinely come apart: strong verbal reasoning tells you very little about someone's spatial reasoning.

The Wonderlic and CCAT practice tests are the exception — they mix several channels into one round, because the real exams do. Both use original questions written to match the published format and difficulty. They are not retired or leaked items from the real tests, which are copyrighted by their publishers.

2. How the score is produced

There is no hidden weighting and no proprietary algorithm. The whole calculation is four steps:

  1. Draw. Questions are shuffled and a fixed number is drawn from a larger bank, so a retake is not simply the same test again.
  2. Mark. One point per correct answer, every question weighted equally. No penalty for a wrong answer, no partial credit.
  3. Convert. Your score is the share of drawn questions you answered correctly — correct divided by drawn, rounded to a whole percent.
  4. Break down. The same calculation runs again within each skill, which is what produces the per-skill bars and the "your weakest area" pointer.

Here is the exact size of every bank and how many questions each test draws from it:

TestQuestions in bankDrawn per attemptSkills scored
Critical Thinking Test 40 20 5
Logical Reasoning Test 30 20 5
Verbal Reasoning Test 30 20 5
Numerical Reasoning Test 26 20 6
Deductive Reasoning Test 24 20 4
Mechanical Reasoning Test 12 12 4
Abstract Reasoning Test 10 8 5
Spatial Reasoning Test 10 8 3
Wonderlic Practice Test 40 20 3
CCAT Practice Test 38 20 3

Note the small banks. On the abstract, spatial and mechanical tests you will see most or all of the bank in a single attempt, so a retake is close to a repeat. Treat repeat scores on those three as practice rather than as an improving measurement. Expanding these banks is on our list.

What the bands mean

Alongside the percentage you get a short verdict. These bands are our own editorial descriptions, written to be useful feedback — they are not statistically derived cut scores, because we have no norm sample (see the limits below). The thresholds are fixed and identical on every test:

ScoreBand shown
85–100%Excellent — top-tier reasoning
70–84%Strong
55–69%Solid, with clear room to grow
40–54%Developing — worth deliberate practice
0–39%Early stage — practise each question type

What is stored

Scoring runs entirely in your browser and your answers are never transmitted anywhere. The only thing kept is your best percentage per test, saved in that browser's local storage on that one device, so you can see whether you beat it. Clearing your browser data erases it. There is no account, and no way for us to see your result. The privacy policy has the full picture.

3. These are practice tools, not diagnostics

This is the distinction that matters most, so it is worth stating flatly. Everything here is built for training and self-reflection. None of it is a diagnostic or clinical instrument, an IQ test, a certification, or a selection tool for hiring.

A genuinely diagnostic instrument needs things we do not have and do not claim: a representative norm sample to compare you against, published reliability coefficients, validity evidence against an external criterion, supervised administration, and controlled item exposure. Commercial appraisals such as Watson–Glaser carry that apparatus. Ours do not. What they offer instead is immediate, explained feedback on real reasoning questions — which is what actually builds the skill.

If you are preparing for a specific employer's assessment, use these to rehearse the format and find your weak spot, then treat that employer's own published guidance as authoritative.

4. What a result cannot tell you

The honest limits of any score you get here:

5. How the questions are written and checked

Found an error?

If a question, answer key or explanation looks wrong, we want to know — a wrong key in a practice test is worse than no practice test at all. Corrections reach us through the contact route on the about page; confirmed fixes are made directly and the page's updated date changes with them.