How The IQ Gym works — and what it honestly can't do
The IQ Gym is built on a simple, well-evidenced claim: practice, familiarity, and strategy reliably improve performance on IQ tests. It deliberately avoids a second claim that the evidence does not support: that training here permanently raises your underlying general intelligence. This page explains both halves, and how our scores are computed.
What the science supports
Meta-analyses of test-retest research covering more than 150,000 participants find that simply retaking a cognitive ability test raises scores by roughly a third of a standard deviation — about five IQ points — with further gains on a third attempt before the effect plateaus. Gains are larger when practice is accompanied by coaching, and they are largest of all on figural content like matrix reasoning: exactly the material most IQ tests lead with.
Solving strategy is trainable, too. Research on matrix reasoning distinguishes 'constructive matching' — predicting the answer before scanning the options — from 'response elimination', and finds that the predictive strategy is what stronger performers do. Every The IQ Gym lesson teaches these strategies explicitly: rows-first scanning, one attribute at a time, decision ladders for number series, bridge sentences for analogies, set diagrams for syllogisms.
What we will not claim
Studies that trained working memory hoping to raise fluid intelligence — the famous dual n-back literature — did not hold up: training improves the trained task and close variants, with little credible transfer to broader ability. Regulators have also acted against brain-training marketing that promised real-world cognitive gains without evidence. So our promise is exact: The IQ Gym improves your test performance — your familiarity, pattern knowledge, strategy, speed, and confidence. We do not promise a new brain.
What the test samples — and what it doesn't
Modern intelligence research organizes abilities in the Cattell–Horn–Carroll (CHC) model: a general factor at the top, broad abilities beneath it — fluid reasoning, comprehension-knowledge, visual-spatial processing, working memory, processing speed — and narrow abilities under those. Every The IQ Gym family is tagged with the CHC ability it draws on: matrix puzzles and series load on fluid reasoning (induction), vocabulary items on comprehension-knowledge, rotation and folding on visual-spatial processing. Our four scored domains sample the broad abilities that mainstream IQ tests weight most heavily.
Sampling is the honest word. A short online battery measures a useful slice of the CHC hierarchy, not the whole of it — a full clinical evaluation samples more abilities, in person, with an examiner. And a domain profile is a study guide, not a diagnosis: scoring lower in one domain here means that domain deserves training time, nothing more.
One honest caveat about the categories themselves: our "Verbal" domain groups two different abilities — vocabulary knowledge (synonyms, antonyms, analogies) and verbal reasoning (letter series, deduction, syllogisms). They correlate, but a sharp gap between them is meaningful, so read a Verbal result as a blend and lean on the lesson-level breakdown when the two diverge. Numerical and Spatial likewise mix pure reasoning with a little learned knowledge. We tag every family with its precise CHC ability and are working toward reporting these as separate strands.
How your score is estimated
Every question in our bank carries a difficulty rating that updates as real people answer it — the same family of methods (Rasch/Elo and item-response theory) used by serious adaptive testing programs. A fixed set of anchor questions is frozen to pin the difficulty scale in place over time, new questions earn live calibration data before they may affect anyone's score, and ultra-fast responses are screened out so guessing sprees don't distort the statistics. When you finish an assessment we estimate your ability per domain from which questions you answered correctly and how hard they were, combine the domains, and convert to the familiar IQ scale: mean 100, standard deviation 15.
We always report a band, never a bare number, because measurement without uncertainty is fiction. Even gold-standard clinical tests carry a margin of several points; a 12-minute online sample carries more. Your band is an honest 90% confidence range — the coverage the assessment standards use for an individual score — and it narrows as you complete longer assessments.
Testing professionals evaluate a score the way the Standards for Educational and Psychological Testing require: state the intended interpretation, then back every link in the argument. Ours, in one sentence: your The IQ Gym band estimates your current performance level on IQ-test-style reasoning tasks, calibrated against our user base — backed by verified single-correct-answer items, model-based scoring, and disclosed provisional norms. It is not certified evidence of clinical IQ, and we never sell it as such.
Why online scores are estimates
Professional IQ tests are administered one-on-one by trained examiners under controlled conditions, and normed on carefully constructed population samples. An online test can control none of that: people take it on phones at midnight, sometimes twice. Our reference distribution is built from our item design and recalibrated from real response data on a stated schedule — and its provisional nature is disclosed here rather than hidden. No online score, ours included, qualifies you for high-IQ societies or constitutes a clinical assessment.
The training method
- Teach the generating rules of each question family — the small set of patterns items are built from.
- Worked examples from easy to hard, then untimed drills with instant worked explanations.
- Adaptive difficulty tuned to the learning research: roughly 85% success while a skill is new, tightening to about 72% as you strengthen — hard enough that errors surface, easy enough that corrections land.
- Miss an item and you get one unaided retry before the explanation — generating the correction yourself is what makes it stick.
- A one-tap confidence rating after each answer trains calibration: lucky guesses get scheduled for review even when they landed, and confident errors get the correction that memory research shows sticks hardest.
- Spaced review: your misses — and your slow or lucky corrects — return on an expanding schedule until the pattern sticks.
- Timed sets and full mock exams under exam conditions, because pacing is a skill of its own.
- A trajectory chart of banded mock scores, so improvement is measured, never imagined.
One more honesty note: because our assessments never reuse an item you have seen, your trend line reflects genuine skill growth on fresh material — not memorization of answers.