Short answer

A raw score is the result before conversion, usually the number of points earned or items answered correctly. A scaled score is a derived value on a score scale created for a particular test. The conversion uses that test's scoring rules and reference data, often taking account of age or another defined comparison group. In practice, a raw score tells you what happened on that form; a scaled score helps you interpret how that performance compares within the test's intended framework. Neither number can be interpreted responsibly by itself. You need the instrument, form, administration conditions, norm group, scale definition, and measurement precision. A raw score of 30 is not automatically comparable with 30 on another test, and a scaled score of 10 on one instrument is not automatically equivalent to 10 on another. For a low-stakes introduction, take the free 50-question reasoning test, review its worked answers, and treat the result as educational performance feedback rather than a professional intelligence assessment.

The practical difference in one minute

Imagine two score reports arrive after two different tests. The first says 32 correct out of 40. The second says 32 correct out of 60. The shared number is a raw score, but the performances are not the same proportion and the items may differ greatly in difficulty. Even a percentage would not solve every comparison problem, because it does not show how difficult the items were or how the result compares with the intended population.

A scaled score adds a defined frame of reference. It is produced from the raw result, or sometimes from a model-based ability estimate, using the instrument's own scale. That scale may be designed so its reference group has a particular average and spread. Some intelligence subtests use scaled scores centered at 10 with a standard deviation of 3; an overall intelligence score is often reported on a standard-score scale centered at 100 with a standard deviation of 15. Those are conventions, not universal laws. The test manual determines what a number means.

So the short answer is: raw score counts performance, while scaled score interprets it on a specified scale. The conversion does not change the answers you gave. It changes the units used to describe them.

What a scaled score adds

Scaling is the process of placing results into score units intended to make interpretation easier. The Standards for Educational and Psychological Testing describe raw scores as often difficult to interpret without additional information and explain that scale scores can show relative standing or improve comparability across alternate forms of the same test. That last phrase matters: a scale is not a universal exchange rate between all tests.

A scaled score can put very different raw-score ranges into a more manageable reporting system. For example, a subtest might report values around 10, while a composite might report values around 100. The numerical size of the scores is not a measure of how much intelligence the test contains. A 12 on one subtest is not “more” than a 105 on an overall index simply because 12 is a smaller number. They are different scales serving different reporting purposes.

Some scales are norm referenced, meaning the result is interpreted against a defined group. Others support criterion-referenced interpretations, where score ranges describe performance against a task standard or domain. The same test can sometimes support more than one interpretation, but each intended use needs evidence. A scaled label alone does not authorize a conclusion about school placement, employment, diagnosis, or future achievement.

A worked example without a proprietary test item

Consider an original practice section with 12 newly written, right-or-wrong reasoning questions. A reader answers 9 correctly. The raw score is 9 out of 12. That statement is direct and reproducible. If the section is only being used for practice, it may be all that is needed to discuss the result: review the three missed operations, inspect the explanations, and try fresh examples that use the same reasoning method.

Now suppose an actual test developer administers the same section to a defined reference group and creates a conversion table. The reader's raw score of 9 could map to a scaled score on that instrument. The map might not be a simple percentage calculation. It could reflect the observed distribution of raw scores, age-specific comparisons, or a model used by the test. A raw score of 9 would not safely map to the same scaled value on another test, even if both sections contained 12 questions.

This example shows why a website cannot responsibly convert an isolated number from an unknown quiz into an IQ score. To do that, it would need a documented instrument, an appropriate reference sample, a defensible scoring method, and evidence for the intended interpretation. The example also shows why a scaled score is not a reward added to the raw score. It is a new description generated under defined rules.

Scaled score, standard score, percentile, and IQ are different labels

Score reports often place several kinds of numbers together. A scaled score is a score on a specified reporting scale. A standard score is usually a transformed score with a stated reference average and spread. A percentile rank describes the position of a result within a defined reference population. It is not the percentage of questions answered correctly. An IQ score is a particular kind of standard-score report used by some intelligence instruments, not a generic synonym for every online quiz result.

The conventions can be easy to confuse. A subtest scaled score with a reference mean of 10 and standard deviation of 3 cannot be read as an IQ score with a mean of 100 and standard deviation of 15. A percentile also does not mean that the person has that percentage of intelligence or that they answered that percentage of items correctly. It reports relative standing under the relevant norm definition.

The current sample report published by Pearson illustrates the separation: it lists subtest scaled scores, composite standard scores, percentile ranks, and confidence intervals in different fields. That is a useful visual reminder that these numbers are related through a test's scoring system, but they are not interchangeable. Never borrow a conversion from a different edition, language, age group, or instrument.

Why the norm group changes the meaning

A norm is information about a reference population used for comparison. The important question is not whether a test has a large-looking norm table. It is whether the reference group is suitable for the interpretation being made. The testing standards state that norm-referenced interpretation depends partly on the appropriateness, technical quality, and continuing usefulness of the reference group.

Age, language, education, culture, administration conditions, and the date of data collection can affect the relevance of a comparison. A score report may therefore use age-specific norms, and some instruments may offer additional reference frames. That does not mean the test is “correcting” a person's identity. It means the publisher has selected a comparison rule for a stated purpose. The rule should be disclosed and supported.

A national or professional organization can give general guidance about norm quality, but it cannot turn an unknown browser quiz into a professionally normed instrument. For this reason, do not compare your raw score with a friend's raw score from a different test, and do not treat a scaled score as portable merely because the label sounds familiar.

Precision still matters after conversion

Scaling can make a result easier to read, but it cannot remove uncertainty. Every test score is an estimate of performance under particular conditions. Fatigue, distraction, language familiarity, timing, guessing, item sampling, and ordinary measurement error can affect the observed result. A report that prints many digits does not necessarily measure more finely.

The testing standards make this practical point directly: too few scale points can discard information, while too many can encourage people to interpret differences smaller than the amount of measurement error. Professional reports may therefore show a confidence interval around a composite score. A confidence interval is a range produced by a stated method to communicate score precision; it is not permission to choose the most flattering endpoint.

This changes how to read a difference. If one scaled result is 10 and another is 11, the difference may be worth noticing descriptively, but it is not automatically a meaningful change in ability. Ask whether the report provides a standard error, interval, reliability evidence, and a comparison rule for that exact score. Without those details, use cautious language such as “higher on this administration,” not “my intelligence increased by one point.”

What to do with your own result

Start by naming the result you actually have. If it is a raw score, record the number of items, scoring rule, time condition, domain, and whether the form was repeated. If it is scaled, find the instrument name, edition, scale definition, norm group, and any reported interval. If the report gives a percentile, check the population it describes. Do not fill missing information with a familiar online conversion chart.

For low-stakes curiosity, the free Test IQ Free assessment is a fixed-form educational reasoning quiz with 50 original questions across pattern, quantitative, verbal, spatial, and logical reasoning. Its appropriate use is to attempt the work, inspect the raw and domain performance supplied by the site, and learn from the worked answers. It is not a professional IQ assessment, and its result should not be used to diagnose a condition, establish giftedness, or make a high-stakes decision.

If a score will affect education, employment, disability support, or a clinical question, bring the report and the decision to a qualified professional. Ask what instrument and norm group fit the question, how uncertainty will be reported, and what other evidence will be considered. A paid practice report can add structure, error-pattern review, and fresh exercises after the free quiz, but practice guidance is not a guarantee of a higher general-intelligence score.

The conversation that prevents a bad conclusion

The most useful next question is not “What number did I get?” It is “What does this number represent, and what decision do I want it to inform?” A raw score can support a transparent review of performance on one form. A scaled score can support a comparison when its scale and reference data are known. A percentile can summarize standing in a defined population. None of them, alone, is a statement about a person's worth, identity, diagnosis, or fixed potential.

When someone shows you an intelligence-test result, ask: “Is this a raw score or a derived score, which instrument and edition produced it, who is the comparison group, and what uncertainty does the report show?” If those answers are unavailable, keep the result at the level of informal performance feedback. That is a smaller claim, but it is the claim the evidence can support.

For an open, low-stakes practice starting point, take the free 50-question test and read every worked answer. The point is not to attach a permanent label to a number. It is to see which reasoning operations you handled under the stated conditions, identify what to practise next, and keep the interpretation proportional to the measurement.

Sources

  1. Standards for Educational and Psychological Testing, Chapter 5: Scores, Scales, Norms, Score Linking, and Cut Scores

    Supports the definitions of raw and scale scores, norm-referenced interpretation, score linking, norms, and the role of measurement error.

  2. APA Dictionary of Psychology: Raw Score

    Supports the definition of a raw score as a score before conversion into another unit or form.

  3. NIEHS Report on Evaluating Features and Application of Neurodevelopmental Tests in Epidemiological Studies

    Supports how normative data convert raw scores into scaled, standard, and percentile scores and why sample fit matters.

  4. Pearson WAIS-5 Sample Score Report with Demographically Referenced Scores

    Illustrates, without reproducing protected items, the separate reporting of raw-to-scaled process scores, composite scores, percentiles, and confidence intervals.

Try the difference yourself.

Work through 50 adaptive reasoning questions, then see your estimated IQ band. Unlock a detailed personalized report after purchase.

Take the free test