How We Build and Score Our Tests
Every test on this site is free, instant and unsupervised — which makes it all the more important to be precise about what a result does and does not tell you. This page documents the whole method: what is measured, how the number is produced, and where it stops being meaningful.
1. What each test measures
Two different kinds of test live on this site, built to different specifications.
The critical thinking test
This one is structured around five reasoning skills, following the long-established structure used by critical-thinking appraisals — most recognisably the Watson–Glaser Critical Thinking Appraisal published by Pearson TalentLens, whose subtests carry these same five names. The underlying idea of critical thinking as a set of separable, teachable skills comes from the 1990 APA Delphi Report (Facione), the expert-consensus statement most later assessment work builds on. Our questions are written in-house; only the skill structure is shared.
| Skill | What the questions actually test |
|---|---|
| Inference | How well you judge the likely truth of a conclusion drawn from stated facts. |
| Recognition of Assumptions | Whether you spot what an argument quietly takes for granted. |
| Deduction | Whether you can tell what necessarily follows from what merely seems to. |
| Interpretation | Whether a conclusion is actually warranted by the passage in front of you. |
| Evaluation of Arguments | How reliably you separate genuinely strong arguments from weak ones. |
For the broader question of how to reason well, rather than how to score it, we lean on the Paul–Elder framework from the Foundation for Critical Thinking, which separates the elements of thought from the intellectual standards used to judge them. The University of Louisville maintains a clear public summary of it.
The aptitude and reasoning tests
Each of these isolates one reasoning channel — verbal, numerical, deductive, abstract, spatial or mechanical — and mirrors the question format used in graduate and job assessments. They are kept separate rather than merged because the skills genuinely come apart: strong verbal reasoning tells you very little about someone's spatial reasoning.
The Wonderlic and CCAT practice tests are the exception — they mix several channels into one round, because the real exams do. Both use original questions written to match the published format and difficulty. They are not retired or leaked items from the real tests, which are copyrighted by their publishers.
2. How the score is produced
There is no hidden weighting and no proprietary algorithm. The whole calculation is four steps:
- Draw. Questions are shuffled and a fixed number is drawn from a larger bank, so a retake is not simply the same test again.
- Mark. One point per correct answer, every question weighted equally. No penalty for a wrong answer, no partial credit.
- Convert. Your score is the share of drawn questions you answered correctly — correct divided by drawn, rounded to a whole percent.
- Break down. The same calculation runs again within each skill, which is what produces the per-skill bars and the "your weakest area" pointer.
Here is the exact size of every bank and how many questions each test draws from it:
| Test | Questions in bank | Drawn per attempt | Skills scored |
|---|---|---|---|
| Critical Thinking Test | 40 | 20 | 5 |
| Logical Reasoning Test | 30 | 20 | 5 |
| Verbal Reasoning Test | 30 | 20 | 5 |
| Numerical Reasoning Test | 26 | 20 | 6 |
| Deductive Reasoning Test | 24 | 20 | 4 |
| Mechanical Reasoning Test | 12 | 12 | 4 |
| Abstract Reasoning Test | 10 | 8 | 5 |
| Spatial Reasoning Test | 10 | 8 | 3 |
| Wonderlic Practice Test | 40 | 20 | 3 |
| CCAT Practice Test | 38 | 20 | 3 |
Note the small banks. On the abstract, spatial and mechanical tests you will see most or all of the bank in a single attempt, so a retake is close to a repeat. Treat repeat scores on those three as practice rather than as an improving measurement. Expanding these banks is on our list.
What the bands mean
Alongside the percentage you get a short verdict. These bands are our own editorial descriptions, written to be useful feedback — they are not statistically derived cut scores, because we have no norm sample (see the limits below). The thresholds are fixed and identical on every test:
| Score | Band shown |
|---|---|
| 85–100% | Excellent — top-tier reasoning |
| 70–84% | Strong |
| 55–69% | Solid, with clear room to grow |
| 40–54% | Developing — worth deliberate practice |
| 0–39% | Early stage — practise each question type |
What is stored
Scoring runs entirely in your browser and your answers are never transmitted anywhere. The only thing kept is your best percentage per test, saved in that browser's local storage on that one device, so you can see whether you beat it. Clearing your browser data erases it. There is no account, and no way for us to see your result. The privacy policy has the full picture.
3. These are practice tools, not diagnostics
This is the distinction that matters most, so it is worth stating flatly. Everything here is built for training and self-reflection. None of it is a diagnostic or clinical instrument, an IQ test, a certification, or a selection tool for hiring.
A genuinely diagnostic instrument needs things we do not have and do not claim: a representative norm sample to compare you against, published reliability coefficients, validity evidence against an external criterion, supervised administration, and controlled item exposure. Commercial appraisals such as Watson–Glaser carry that apparatus. Ours do not. What they offer instead is immediate, explained feedback on real reasoning questions — which is what actually builds the skill.
If you are preparing for a specific employer's assessment, use these to rehearse the format and find your weak spot, then treat that employer's own published guidance as authoritative.
4. What a result cannot tell you
The honest limits of any score you get here:
- No norm group. Your percentage says how many of these questions you got right. It places you in no percentile against any population, because we have not normed the tests on one.
- No published reliability or validity study. We have not run one. Any claim that a score predicts job performance or academic success would be unsupported, so we do not make it.
- Small banks, repeated items. With banks of 10 to 40 questions, repeat attempts recycle items. A rising score across many retakes partly reflects familiarity, not only improvement.
- Unsupervised and untimed. Nothing stops you pausing, looking something up, or retaking. Real assessments are timed and invigilated, which changes performance substantially.
- A narrow slice of thinking. Multiple-choice questions can measure whether you pick the valid conclusion. They cannot measure whether you ask the right question in the first place, sit with ambiguity, or change your mind on new evidence — arguably the more important parts of thinking well.
- One sitting is noisy. Attention, fatigue and luck of the draw move a single score by a meaningful margin. A trend across several attempts on different days says far more than any one result.
5. How the questions are written and checked
- Every question is written in-house. Nothing is scraped, spun, or lifted from another test.
- Every question ships with a worked explanation of why the correct answer is correct — and, where a wrong option is tempting for a specific reason, why it fails.
- Answer keys are verified by running each bank against its own key, and correct answers are spread across option positions so a test cannot be gamed by always picking the same slot.
- Visual questions (abstract, spatial, mechanical) are generated from shape and diagram helpers, so the figure and the stated rule cannot drift apart.
- Pages carry a visible "updated" date and are revised whenever material is corrected or expanded.
Found an error?
If a question, answer key or explanation looks wrong, we want to know — a wrong key in a practice test is worse than no practice test at all. Corrections reach us through the contact route on the about page; confirmed fixes are made directly and the page's updated date changes with them.