AP-PSYCH-2.8

U2.8 Intelligence and Achievement

Master AP Psych 2.8: compare Spearman's g, Sternberg, Gardner, and emotional intelligence, plus test reliability, validity, standardization, and heritability.

What you'll do in this lesson

A voice-first session with the Crimsora tutor on U2.8 Intelligence and Achievement, then targeted practice and FRQs — with the tutor adapting to where you get stuck.

What this lesson covers

How do we measure something as slippery as intelligence — and can a single number really capture it? Topic 2.8 tackles the biggest debates in psychology: whether intelligence is one general ability or many, how tests are built to be fair and meaningful, and what genetics and environment contribute to differences between people. This lesson gives you the vocabulary and reasoning the exam expects. You'll learn to distinguish the major theories, evaluate what makes a test trustworthy, and think carefully about heritability and group differences without overstating what the data show. These distinctions show up in multiple-choice items and FRQ prompts alike, so precise definitions matter.

Theories of Intelligence

Psychologists disagree about whether intelligence is one thing or many. Charles Spearman argued for a general intelligence factor, called gg, based on the observation that people who score well on one type of mental task tend to score well on others. He used factor analysis to identify this common thread running through diverse abilities.

Other theorists broke intelligence apart. Robert Sternberg proposed the triarchic theory, dividing intelligence into three types: analytical (problem-solving, the kind schools test), creative (adapting to novel situations, generating new ideas), and practical ("street smarts," handling everyday tasks). Howard Gardner went further with multiple intelligences, proposing relatively independent domains such as linguistic, logical-mathematical, spatial, musical, bodily-kinesthetic, interpersonal, intrapersonal, and naturalist. Gardner pointed to savant syndrome and brain damage that spares some abilities while destroying others as evidence.

Emotional intelligence (Salovey, Mayer, popularized by Goleman) is the ability to perceive, understand, manage, and use emotions. It predicts social success but is controversial as a form of "intelligence."
TheoryKey idea
Spearman ggOne general factor underlies all abilities
SternbergThree types: analytical, creative, practical
GardnerMany independent intelligences
Emotional intelligencePerceiving and managing emotions
A common misconception is treating Gardner and Sternberg as identical — both reject a single gg, but Sternberg names three domains while Gardner names many.

History of Intelligence Testing

Modern testing began with Alfred Binet and Théodore Simon in France, who created the first test to identify schoolchildren needing extra help. They introduced the idea of mental age — the chronological age that typically corresponds to a given level of performance. A child performing like an average 8-year-old has a mental age of 8.

Lewis Terman at Stanford adapted Binet's work for American use, producing the Stanford-Binet test. This context introduced the intelligence quotient, originally calculated as IQ=mental agechronological age×100IQ = \frac{\text{mental age}}{\text{chronological age}} \times 100. Modern tests no longer use this formula; instead they compare scores to age peers.

David Wechsler developed the widely used Wechsler Adult Intelligence Scale (WAIS) and versions for children, providing an overall score plus subscores for verbal comprehension, perceptual reasoning, working memory, and processing speed.

Be careful to distinguish achievement tests, which measure what you have already learned (like a history final), from aptitude tests, which attempt to predict future performance or capacity to learn. The exam frequently tests this difference. A misconception is that IQ is a fixed, precise measure of worth — it is a statistical comparison to a norm group, not a fixed biological quantity.

Test Properties: Standardization, Reliability, Validity

For a test to be scientifically useful, it must meet three properties. Standardization means administering the test to a representative sample (the standardization sample) to establish norms, so any individual score can be compared meaningfully. Scores typically form a normal distribution — a bell curve where most people cluster near the average and fewer fall at the extremes.

Reliability means the test yields consistent results. Types include test-retest reliability (same score on repeat administrations), split-half reliability (two halves of the test agree), and inter-rater reliability (different scorers agree).

Validity means the test measures what it claims to measure. Content validity covers whether items sample the relevant domain; predictive (criterion) validity covers whether scores forecast the outcome they should — for example, whether an aptitude test predicts college grades.
PropertyQuestion it answers
StandardizationCompared to whom? Are there norms?
ReliabilityAre results consistent?
ValidityDoes it measure the right thing?
A key insight the exam tests: a test can be highly reliable yet not valid. A bathroom scale that always reads 10 pounds heavy is perfectly consistent (reliable) but wrong (invalid). Reliability is necessary but not sufficient for validity. Also note the Flynn effect — the observed rise in average IQ scores over the 20th century, requiring periodic re-standardization.

Heritability and Group Differences

Heritability is the proportion of variation in a trait, within a population, that can be attributed to genetic differences. It is one of the most misunderstood concepts on the exam. Heritability is about variation among individuals in a group, not about how much of one person's intelligence comes from genes. A heritability estimate of 0.50.5 does not mean half of your intelligence is genetic.

Twin and adoption studies suggest intelligence is substantially heritable — identical twins raised apart still show correlated scores. But environment matters enormously: enriched or impoverished environments, nutrition, education, and stress all shape measured intelligence.

Critically, heritability within groups says nothing about the causes of differences between groups. Two groups raised in systematically different environments can differ for entirely environmental reasons even if the trait is heritable within each. Score gaps between groups have historically been influenced by unequal access to education, test bias, and stereotype threat — the self-fulfilling anxiety that arises when a person fears confirming a negative stereotype about their group, which can lower performance.

A test can show test bias if it disadvantages a group due to cultural content unrelated to the ability being measured. The exam expects you to recognize that group differences in scores are not evidence of genetic differences between groups.

Key terms

General intelligence (gg).
Spearman's proposed single underlying factor that accounts for correlated performance across different mental tasks, identified through factor analysis.
Triarchic theory.
Sternberg's model dividing intelligence into analytical, creative, and practical abilities.
Multiple intelligences.
Gardner's theory that intelligence consists of several independent domains such as linguistic, spatial, musical, and interpersonal.
Standardization.
Establishing test norms by administering it to a representative sample so individual scores can be compared.
Reliability.
The consistency of a test's results across time, forms, or scorers.
Validity.
The extent to which a test measures or predicts what it is intended to measure.
Heritability.
The proportion of variation in a trait within a population attributable to genetic differences; not a statement about a single individual.
Stereotype threat.
Performance decline caused by anxiety about confirming a negative stereotype about one's group.

Worked example

A school psychologist gives a new reasoning test to 500 students. The same students retake it a month later and get nearly identical scores. However, the test scores do not correlate with students' later grades or with other established reasoning measures. Evaluate the test using the concepts of reliability and validity.
First, identify what the consistent scores tell us. Because students earned nearly identical scores one month apart, the test shows strong test-retest reliability — it produces consistent, repeatable results.

Next, evaluate validity. Validity asks whether the test measures what it claims to measure. Here the scores fail to correlate with grades or with other reasoning measures. That means the test lacks predictive (criterion) validity and likely lacks construct validity — it is not measuring reasoning ability as intended.

Finally, connect the two. This case illustrates the exam's favorite principle: a test can be reliable without being valid. Consistency alone does not guarantee that a test measures the right thing. The bathroom-scale analogy applies — a scale reading 10 pounds heavy every time is reliable but invalid. The psychologist should not use this test to make decisions about students, because its consistent scores do not reflect the trait it claims to assess.

Practice questions

A researcher claims her intelligence test is highly reliable but critics argue it is not valid. Which finding would best support the critics' claim?
  1. Students receive nearly the same score when they retake the test
  2. The two halves of the test produce closely matching scores
  3. Test scores fail to predict any real-world academic or cognitive outcomes
  4. Different examiners scoring the test arrive at the same results

Answer: Test scores fail to predict any real-world academic or cognitive outcomes

Validity concerns whether the test measures what it claims. If scores do not predict relevant outcomes, the test lacks predictive validity. The other options describe forms of reliability (test-retest, split-half, inter-rater), which is consistency — not the same as validity. This item tests the crucial distinction that reliability does not guarantee validity.
Compare Spearman's concept of general intelligence with Gardner's theory of multiple intelligences, and explain one type of evidence Gardner used to support his view.

Answer: Spearman argued that a single general factor, gg, underlies performance across all mental tasks because scores on different abilities tend to correlate. Gardner rejected a single factor, proposing several relatively independent intelligences (such as linguistic, spatial, and musical). Gardner supported his view with evidence like savant syndrome and cases of brain damage, where one ability can be preserved or destroyed independently of others.

A full answer must contrast the core claims — one factor versus many — and cite Gardner's actual evidence. Savant syndrome, in which a person has an exceptional isolated skill despite limited general ability, suggests intelligences can operate independently, undercutting the idea of a single unifying gg.
A standardized IQ test administered decades ago is re-normed today, and researchers notice average raw performance has risen over time. What phenomenon does this illustrate, and why does it require re-standardization?

Answer: This illustrates the Flynn effect, the observed rise in average intelligence test performance over the 20th century. Re-standardization is required because norms are based on how a representative sample performs; if overall performance rises, old norms would make today's average person appear above average, distorting comparisons. Updating the standardization sample keeps the mean anchored appropriately.

The Flynn effect shows that test norms are relative to a reference population and can drift, so tests must be periodically re-standardized to remain meaningful. This reinforces that IQ is a comparative statistic, not a fixed absolute measure.

FAQ

What is the difference between reliability and validity?
Reliability is consistency — getting the same result repeatedly. Validity is accuracy — measuring what you actually intend to measure. A test can be reliable without being valid, like a scale that consistently reads 10 pounds too high. But a valid test must be reliable, so reliability is necessary but not sufficient for validity.
Does high heritability mean intelligence is mostly genetic and can't change?
No. Heritability describes how much of the variation among people in a population is linked to genes, not how fixed a trait is or how much of one person's intelligence is genetic. Traits with high heritability can still be strongly shaped by environment, and environmental improvements can raise scores.
What's the difference between an aptitude test and an achievement test?
An achievement test measures what you have already learned, like a final exam. An aptitude test attempts to predict your future performance or capacity to learn. The distinction can blur in practice, but on the exam, focus on past learning versus future prediction.
Why do group differences in test scores not prove genetic differences between groups?
Heritability within a group says nothing about causes of differences between groups. Groups raised in systematically different environments — unequal schooling, nutrition, or facing stereotype threat and test bias — can differ for purely environmental reasons, even when a trait is heritable within each group.

Learn this with a teacher, not a page

The Crimsora tutor teaches U2.8 Intelligence and Achievement live — explaining on a whiteboard, asking you questions, and adapting to where you get stuck.