Decide when a sample supports generalization

Lesson progressPractice problems 0/3
Difficulty
Beginner
Estimated time
25 minutes
Techniques
Random-samplingGeneralizationSelection-biasTarget-population

What you’ll learn

  1. Recognize questions that ask who a result can speak for.
  2. Tell the group a researcher wants to describe from the group that was actually sampled.
  3. Explain why a random sample can represent the population it was drawn from.
  4. Spot selection bias in convenience, volunteer and self-selected samples.
  5. Pick the largest population the sampling method supports.

Prerequisites

You’re ready. No earlier Aniko lesson is required.

Why this matters on the SAT

Match the claim to the sample

A survey can be accurate for the people it measured and still not speak for the larger group a conclusion names. On the SAT, the answer often turns on one detail: where the sample came from.

Solution to the example

The agency did pick people at random, but only from its list of monthly bus-pass holders. So the result can reasonably reach everyone on that list, but not every adult in the county. That rules out A.

A sample percentage is an estimate, so about fits better than exactly. That rules out B. Choice C is correct.

SAT example

A transit agency wants to estimate the proportion of all adults in a county who would use a new express bus route. The agency randomly selects 240240 people from a list of adults who currently hold monthly bus passes. Of those surveyed, 68%68\% say they would use the route.

Which conclusion is best supported?

  1. A

    About 68%68\% of all adults in the county would use the route.

  2. B

    Exactly 68%68\% of monthly bus-pass holders would use the route.

  3. C

    About 68%68\% of the county’s monthly bus-pass holders would use the route.

  4. D

    Using a monthly bus pass causes an adult to support the route.

Spot a population question

These questions describe a survey, then ask something like:

  • Which population can the result be extended or generalized to?
  • Is the sample representative of a certain group?
  • Is the survey biased?
  • Which sampling method would give reliable information about a population?
  • What keeps the claim from reaching a larger group?

Each one comes down to two groups:

  • The target population is the whole group the researcher wants to describe.
  • The sample is the smaller group that was actually picked and measured.

Before you read the answer choices, fill in this sentence:

The researcher wants to describe ______, but the data came from ______.

If the blanks name different groups, look closely at how the sample was picked.

Check your understanding:

A college wants to describe all enrolled students, but it surveys students leaving the fitness center. Fill in the two blanks.

Let the source population set the limit

A random sample uses chance to pick people or objects from a population. In a simple random sample, every member of that population has the same chance of being picked.

Chance doesn’t play favorites, so a random sample helps you get a representative sample: one that doesn’t systematically leave out or overcount any important part of the population it came from.

Here’s the central rule:

Random sample from a population⇒generalize to that population\boxed{\text{Random sample from a population} \Rightarrow \text{generalize to that population}}

The last three words do the work. Picking at random doesn’t make the population any bigger.

Say a museum randomly picks 150150 names from its complete list of annual members. The results can represent all annual members of that museum. They don’t automatically represent:

  • all museum visitors
  • all residents of the city
  • members of other museums

Anyone who isn’t a member of this museum was never on the list, so chance never had a shot at picking them. The list chance actually picks from is called the sampling frame, and it sets the limit. So the key question is: who could have been picked?

Common mistake:

It’s tempting to see the word “random” and pick the answer with the biggest population. But random selection only works inside the list it picked from. Name that list, then reject any conclusion that reaches past it.

Check your understanding:

A company randomly selects 9090 workers from a complete list of employees at its west warehouse. What’s the largest population the survey directly supports?

Spot selection bias

Selection bias happens when the way people get into a sample systematically favors some parts of the target population over others. The sample can then lean toward an answer the full population wouldn’t give.

Convenience sample

A convenience sample uses whoever is easy to reach, such as:

  • the first customers through one entrance
  • people at one event or place
  • one class, because it’s nearby
  • visitors to a business tied to the survey topic

Easy to reach doesn’t mean representative. A place can draw people who share an interest, a schedule, an age range or a habit.

Volunteer or self-selected sample

In a voluntary-response sample, people decide for themselves whether to answer. That’s why it’s also called a self-selected sample.

People with strong opinions, or a special interest in the topic, may be more likely to answer. And posting a survey link where everyone can see it doesn’t make the respondents random. Everyone could answer, but chance didn’t choose who did.

An organized list isn’t automatically random

Picking the first 100100 names in alphabetical order, or the people at the top of a sign-up sheet, follows a rule, but chance plays no part. The order can be tied to something that matters. The top of a sign-up sheet, for example, holds the people who signed up first, and they may be the most eager. On top of that, not every member had the same chance of being picked.

Common mistake:

It’s natural to think a bigger sample fixes bias, but it doesn’t. More answers from the same tilted source only repeat the same tilt. To fix the design, change who can be picked, and let chance pick from the whole target population.

Try it yourself:

A city posts a voluntary survey about noise rules in an online group for local musicians. Before you open the check, decide who might be overrepresented and which population the city can’t describe yet.

Check your understanding:

Why doesn’t that survey represent all city residents?

Describe only the group the evidence reaches

When a sample wasn’t randomly picked from the target population, don’t swing to the other extreme and decide the data mean nothing. They still describe the people who answered.

Say 73%73\% of 200200 visitors surveyed at a food festival prefer outdoor seating. Then:

  • it’s true that 73%73\% of those 200200 visitors gave that answer
  • the survey can tell you something useful about those people
  • but it doesn’t justify claiming that about 73%73\% of all city residents feel the same way

So there are three cases:

  1. If chance picked from a list of the whole target population, like every household in town, the result can represent that whole population.
  2. If chance picked from a narrower list, like one museum’s members, the result speaks only for the people on that list.
  3. If the sample was convenient or self-selected, like the festival visitors, the result describes the people who answered and goes no further.

How close a sample percentage is likely to land to the population’s, and how sample size affects that, is a different question, the one Estimate populations and margin of error answers.

Check your understanding:

A survey randomly samples apartment renters from a county housing registry. Can the result represent all adults in the county? If not, who can it represent?

Example: A random sample from a narrower list

Worked example

A recreation department wants to estimate the proportion of all households in the town that support adding weekend sports clinics. The department uses the complete registration list for the town’s youth recreation programs and randomly selects 180180 households from that list. Most sampled households support the clinics.

Which conclusion is best supported?

  1. A

    The survey supports an estimate for all households in the town because the households were selected at random.

  2. B

    The survey supports an estimate for households registered for the town’s youth recreation programs, but not for all town households.

  3. C

    The survey cannot support any conclusion because fewer than all registered households were surveyed.

  4. D

    The survey proves that registering for a youth program causes a household to support weekend clinics.

Step 1

Who does the department want to describe?

The department wants to describe all households in the town. That’s the target population.

Step 2

Who could have been picked?

Only households on the youth-program registration list could land in the sample. A town household with no one registered never had a chance of being picked.

So the sampling frame is narrower than the target population.

Step 3

What does the random pick give you?

Picking at random helps the 180180 households represent the registration list they came from.

It can’t fill in the households that were never on that list.

Step 4

Pick the conclusion that stays inside the list

Choice B names the population the design supports: households registered for the town’s youth recreation programs.

Choice A reaches past the list. Choice C is too cautious, because a random sample can represent a larger list without surveying everyone on it. Choice D claims cause, which a survey like this doesn’t test.

Check your understanding:

What would the department need to change to generalize to all town households?

Keep random sampling and random assignment separate

These two sound alike, but they answer different questions:

  • Random sampling decides who gets picked from a population. It sets who the result can speak for.
  • Random assignment decides which treatment each participant gets. It matters when a question asks about cause and effect.

For the questions here, you only need to ask one thing:

Was the sample randomly picked from the population the conclusion names?

When a question asks whether a treatment caused an outcome, that’s the job of random assignment, and it comes next in Decide when a study supports causation.

Choose the population by hand

You won’t need Desmos for these. A calculator can’t tell you which list a sample came from, whether it’s representative or who the result can describe. Numbers in the question don’t change that, because the decision is about groups, not arithmetic.

Instead, work through the question in three moves:

  1. Underline the group the conclusion or research goal names, like “all households in the town.”
  2. Circle the people who could actually have been picked, like “households on the registration list.”
  3. Check whether chance did the picking. That tells you which of the three cases above you’re in.

Then read the choices for overreach. Words like all, every and exactly, or a group bigger than the one you circled, are the usual giveaways.

Practice problems

For each one, ask who could have been picked, and how.

Name the narrower population

Practice problem

An aquarium randomly selects 160160 people from its complete list of annual pass holders. Each selected person is asked whether the aquarium should stay open later on Fridays. Of those surveyed, 61%61\% say yes.

Which inference is most appropriate?

Answer choices
Calculator loads as you approach
Find the list the random sample came from.

Reject a topic-linked sample

Practice problem

A county wants to estimate the proportion of all county households that plan to install solar panels within the next 55 years. Surveyors question 300300 visitors leaving a weekend home-energy expo, and 74%74\% say they plan to install panels.

Which statement best evaluates the county’s use of the result?

Answer choices
Calculator loads as you approach
Ask who an event like this is likely to draw.

Reject a broad voluntary poll

Practice problem

A city wants to estimate the proportion of all voting-age residents who support adding late-night bus service. The city posts an open online poll and mails every household a postcard with the survey link. Of the 840840 residents who choose to respond, 69%69\% support the service.

Which conclusion is best supported?

Answer choices
Calculator loads as you approach
Look closely at how people ended up answering.

Finish the lesson

3 practice examples left

Finish the remaining questions correctly to complete this lesson.

Quick recap

  • The target population is the group the researcher wants to describe. The sample is the group actually picked and measured.
  • A random sample can speak for the population it was drawn from. Ask: who could have been picked?
  • Random selection from a narrower list supports only that list.
  • Convenience samples can overrepresent people tied to a place, schedule, event or topic.
  • Volunteer and self-selected samples can overrepresent people with strong opinions or a special interest.
  • A bigger sample doesn’t fix a biased way of picking people.
  • Without random sampling from the target population, describe the people who answered and stop there.
  • Random sampling decides who a result speaks for. Random assignment is about cause, a separate question.
  • Work these by hand. A calculator can’t tell you who was sampled or spot selection bias.

Next lesson

Decide when a study supports causation

Use random assignment and comparison groups to decide whether a cause-and-effect conclusion is justified.

Start next lesson

Practice

Practice this lesson

476 SAT questions use what this lesson teaches. Practice a few in a study session at the difficulty you choose.

Start practice