UK-based online statistics and data analysis support for USA, UK, and international clients. No exams, no impersonation, no fabricated data.
AP Statistics Exam

Introducing Statistics: Population, Samples, Data, and Investigative Questions

AP Statistics • Unit 1 • Topic 1.1 • Visual course chapter Introducing Statistics What can we learn from data? Begin with the language of a...

Statistics guide Ethical learning support SPSS/R/Python/Excel friendly


AP Statistics • Unit 1 • Topic 1.1 • Visual course chapter

Introducing Statistics

What can we learn from data? Begin with the language of a statistical study: the population we care about, the sample we observe, the individuals that supply information, the variables recorded, and the question that gives every number a purpose.

Population & sampleData & datasetsParameter & statisticInvestigative questions

Previous chapter: What Is the AP Statistics Exam?Course path: AP Statistics → Unit 1 → Topic 1.1
Advertisement · Google AdSense top placement reserved here

Introducing Statistics begins with a simple idea that is easy to say but powerful to apply: data do not explain themselves. A spreadsheet can contain thousands of values, yet those values become statistical evidence only when we know who or what was observed, which larger group matters, what was measured, how the observations vary, and which question the study is intended to answer.

Imagine that a school principal is shown the number 18.7. Is it encouraging, worrying, or meaningless? It might be the average number of minutes students wait for a bus, the percentage of seniors absent on a particular day, the mean score on a twenty-point quiz, or the number of books borrowed per student during a semester. The same numerical symbol can support completely different interpretations. Statistical language supplies the missing context.

This chapter treats a statistical study as a connected story. The story starts with an investigative question. The question points to a population and to one or more characteristics that must be recorded. Because observing the entire population may be impossible, expensive, slow, or unnecessary, researchers often examine a sample. Each item or person observed is an individual or observational unit. The recorded pieces of information become data, and the organized collection becomes a data set. A numerical description of the population is a parameter; a numerical description calculated from the sample is a statistic.

The official AP Statistics framework places Topic 1.1 at the beginning because every later graph, probability, confidence interval, test, and regression model depends on this vocabulary. Introducing Statistics is therefore not a list of definitions to memorize. It is training in how to read, design, and explain a study without losing the real-world meaning of the evidence.

NPopulation size
nSample size
1Question gives purpose
Possible patterns in data
Advertisement · Google AdSense in-content placement 1 reserved here
01
The big idea

What statistics means

Statistics is the science and practice of learning from data while accounting for variation, uncertainty, context, and the way the data were produced. The word also has a second, narrower meaning: a statistic is a numerical value calculated from sample data. These two meanings are related but not identical. A course in statistics teaches a way of reasoning; a statistic is one output within that reasoning.

Introducing Statistics should change how a reader reacts to claims. Instead of asking only, “What number was reported?” a statistically minded reader asks, “What was the question? Who or what did the data describe? How were observations selected? What exactly was measured? How much did the observations differ? Does the reported number describe the sample or claim something about a larger population?” These questions prevent a polished calculation from being mistaken for strong evidence.

Statistics is often divided into two broad activities. Descriptive statistics organize and summarize the data that were actually observed. Tables, graphs, percentages, averages, and measures of spread can all describe a sample or a complete population. Statistical inference uses sample data to learn about a larger population, while acknowledging that another sample could have produced a different result. Topic 1.1 introduces the language required for both activities, even though formal inference is developed later in the course.

A number becomes statistical evidence only after it is connected to a question, a group, a measurement, and a source of variation.Central idea of Introducing Statistics

Statistics is more than arithmetic

Arithmetic can tell us that the mean of 4, 6, and 8 is 6. Statistics asks whether those three values came from the people or objects relevant to the question, whether the measurement was meaningful, whether 6 is an appropriate summary, how the values varied around 6, and whether the result can be generalized. A correct calculation can coexist with a poor study. Conversely, a thoughtfully designed study may produce a result with unavoidable uncertainty. Statistics makes that uncertainty visible rather than pretending it does not exist.

For example, suppose a student surveys five close friends and finds that all five prefer online homework. The sample proportion is 1.00, or 100 percent. The arithmetic is correct. However, if the intended population is all 1,200 students in the school, the selection process is unlikely to represent the population. The result is a statistic describing those five friends. It is not automatically a trustworthy estimate of the population parameter.

Statistics is a complete investigation cycle

1

Question

State what you want to learn and about whom or what.

2

Data

Decide what information is needed and obtain it responsibly.

3

Analysis

Represent, summarize, compare, or model the observations.

4

Conclusion

Answer the question in context and state the limits.

In practice, these stages interact. A preliminary look at data may reveal that a variable was recorded ambiguously. A planned analysis may show that the question needs a clearer population. A conclusion may lead to a better follow-up question. Introducing Statistics teaches the vocabulary needed to move around this cycle without confusing one component with another.

Introducing Statistics checkpointThe subject of statistics is not merely “numbers.” It is evidence-based reasoning with data. Words, categories, images converted into features, dates, measurements, ratings, and yes/no outcomes can all become data when they are recorded in a structured way to answer a question.
Advertisement · Google AdSense in-content placement 2 reserved here
02
Read the whole story

The anatomy of a statistical study

A statistical study can be read as a set of connected components. If one component is missing, the meaning of the study becomes uncertain. Introducing Statistics uses the following map throughout the chapter.

Study map

From a real-world question to statistical evidence

Investigative questionWhat do we want to learn, and why does the answer require data?
PopulationThe complete group of items or individuals about which the question seeks information.
SampleThe subset actually selected and observed when the entire population is not studied.
IndividualsThe entities represented by the rows or records: people, schools, products, days, animals, transactions, or other units.
VariablesThe characteristics recorded for each individual, such as commute time, grade level, battery life, or response category.
Data setThe organized collection of observed values, together with enough labels and documentation to interpret them.
StatisticA numerical summary calculated from the sample data.
ParameterThe corresponding numerical feature of the entire population, often unknown.
Conclusion in contextA statement that reconnects the analysis to the population, variable, and original question.

A miniature study

A public library wants to know the typical number of minutes adult visitors spend in the building on weekday afternoons. During randomly selected weekday afternoon periods, staff record the visit length of 120 adult visitors. The sample mean is 46.3 minutes.

ComponentIdentificationWhy it matters
Investigative questionWhat is the typical weekday-afternoon visit length for adult visitors to this library?It defines the purpose and context.
PopulationAll adult visitors to the library on weekday afternoons during the period of interest.This is the group the library wants to understand.
SampleThe 120 adult visitors whose visit lengths were recorded.These are the observed cases.
IndividualsIndividual adult library visits, or visitors if each visitor appears only once.Each row must represent one clearly defined unit.
VariableVisit length in minutes.This is the characteristic measured on each unit.
StatisticThe sample mean, 46.3 minutes.It summarizes the observed sample.
ParameterThe true mean visit length for the full population.It is the population quantity the library would like to know.

Notice that 46.3 is not simply “the average visit length.” A precise statement says it is the sample mean visit length for the 120 observed adult visits. Whether it estimates the population mean well depends on how the sample was selected and how visit length was recorded. Introducing Statistics trains this precision from the beginning.

Context is not decoration

In statistics, “in context” means matching every abstract term and calculation to the real-world component it represents. If a calculation produces 0.37, a contextual interpretation might be “37 percent of the sampled households reported using public transportation at least once during the previous week.” Merely writing “the proportion is 0.37” omits the group, behavior, and time frame.

Context also prevents unit errors. A mean of 12.4 could mean dollars, hours, points, kilograms, or customer complaints. A standard deviation of 3 has the same units as the variable it describes. A percentage must specify the denominator. A sample size must refer to the number of observational units, not automatically the number of measurements if repeated observations are present.

Introducing Statistics warningDo not identify components by copying the nearest noun. Read the study purpose. The population is the full group about which the question seeks information, not every group mentioned in the story. The sample is the subset actually observed for that purpose.
03
Whole and part

Population and sample

The population is the complete collection of items or individuals relevant to an investigative question. The symbol N commonly represents population size. The sample is the subset from which data are actually obtained, and n commonly represents sample size. A sample is not defined merely by being small; it is defined by being a part of a larger population of interest.

Population, size N

Every item or individual relevant to the question

Sample, size n

The subset actually observed

Defining the population precisely

A useful population definition usually includes more than a broad label. “Students” is rarely precise enough. Which students? At which schools? During what academic year? Enrolled at what point in time? Eligible under which conditions? The population might be “all students enrolled in grades 9–12 at North Valley High School on October 1, 2026.” That definition creates a boundary. A student who transferred after October 1 may not belong to the defined population even though the student later attends the school.

Populations can contain people, but they do not have to. A manufacturer may study all light bulbs produced by a machine during a week. An environmental scientist may study all daily air-quality readings in a city during a year. A streaming service may study all viewing sessions started by U.S. subscribers during a month. A botanist may study all trees of a species in a protected forest. Introducing Statistics uses “items or individuals” because the unit depends on the question.

Why use a sample?

A census attempts to collect data from every member of the population. A census can be appropriate when the population is small, every member can be reached, and the required measurements are practical. However, a census can be expensive, slow, or impossible. Some measurements destroy the item being tested. A company cannot measure the breaking strength of every helmet it intends to sell because the test would leave no helmets to sell. Some populations are continuously changing, such as all future customers or all future production units.

Sampling can provide timely information with fewer resources. Yet a sample introduces a central statistical issue: different possible samples can produce different results. That is sampling variability. A sample of 100 residents may report a different average commute time from another sample of 100 residents, even if both are selected carefully from the same city.

Target population and accessible population

The target population is the group the researcher ultimately wants to understand. The accessible population is the portion that can realistically be reached using the available list, place, period, or system. Suppose a university wants to learn about all graduates from the past ten years but only has current email addresses for 60 percent of them. The target population is all graduates in the defined period. The accessible population is the set with usable contact information. The difference can matter because reachable and unreachable graduates may differ.

Sample ⊂ Population    and usually    n < NThe sample is a subset of the population; sample size is usually smaller than population size.

Population and sample can change with the question

The same data source can support different populations if the investigative question changes. A data set of 500 restaurant orders collected on Fridays could be a sample from all Friday orders at one restaurant, or it might be the complete population of Friday orders during the specific data-collection month. Whether the records form a sample or a population depends on the scope of the question, not on the number of rows alone.

Introducing Statistics checkpointAsk two separate questions: “What full group does the study want to describe?” and “Which units actually supplied data?” The first answer is the population; the second is the sample.
Advertisement · Google AdSense in-content placement 3 reserved here
04
What each row represents

Individuals and observational units

An individual is an object, person, case, or other entity described by the data. The more general term observational unit is useful when the row is not literally a person. In a conventional rectangular data table, each row usually represents one observational unit and each column represents one variable. The word “usually” matters because repeated-measures and hierarchical data can use more complex structures.

The unit must match the question

Suppose a study records the number of emergency-department visits at 40 hospitals for each of 365 days. Is the observational unit a hospital, a day, or a hospital-day combination? The answer depends on the table. If each row is one hospital on one date, the observational unit is a hospital-day. There are up to 40 × 365 = 14,600 rows. If the data are first summarized so that each row contains a hospital’s annual average, the observational unit in the summarized data set is a hospital.

Misidentifying the unit can lead to a false sample size. A study with 50 students measured on 10 occasions has 500 measurements but only 50 student participants. Those measurements are not necessarily 500 independent individuals. Introducing Statistics does not yet develop dependence models, but it establishes the habit of asking what one row or record represents.

People

Student survey

Each individual may be one student. Variables might include grade level, travel time, and preferred learning format.

Objects

Battery test

Each observational unit may be one battery. Variables might include brand, production batch, and operating life.

Events

Customer transaction

Each observational unit may be one purchase. Variables might include price, time, channel, and product category.

Individual, subject, participant, case, and record

Different fields use different words. A subject or participant is usually a person involved in a study. A case may be a person, organization, event, or object. A record is the stored representation of an observational unit. The terms overlap, but the safest practice is to define the unit explicitly: “Each row represents one household interviewed once” or “Each row represents one shipment.”

Units can be nested

Students belong to classrooms, classrooms belong to schools, and schools belong to districts. Employees belong to teams and companies. Measurements may occur within patients over time. At Topic 1.1, the main goal is recognition: a study may contain multiple levels, and the population or sample statement should specify the level to which the question refers. A question about average school size has schools as the units, even though student counts appear in the data.

Introducing Statistics warningDo not use “sample size” to mean the number of cells in a spreadsheet. Sample size counts the observational units selected for the study at the relevant level.
05
Recorded evidence

Data and datasets

A datum is one recorded piece of information. Data is the plural form, although in everyday writing the word is often treated as a mass noun. A data set is an organized collection of data. Organization includes labels, units, definitions, and enough structure to connect values to individuals and variables.

Consider the row below from a study of school travel. The student identifier is ST104. The student’s travel mode is bus, travel time is 27 minutes, distance is 8.4 miles, and arrival status is on time. Each cell contains a datum. The complete table of many students and columns is the data set.

Student IDTravel modeTravel time (min)Distance (miles)Arrival status
ST104Bus278.4On time

Data values need variable definitions

The value 27 is not fully meaningful until the variable is defined. Does travel time begin when the student leaves home or when the bus arrives? Does it end at the school gate, classroom door, or official start time? Are waiting minutes included? Clear operational definitions make values comparable. Without them, two observers may record different numbers for the same journey.

A variable is a characteristic that can take different values across individuals or occasions. Topic 1.2 develops variable types in detail, but Topic 1.1 requires enough understanding to identify what was recorded. In the travel data, travel mode and arrival status are categorical variables; travel time and distance are quantitative variables. The important first step is to name the variable in context rather than simply pointing to a column.

A data set is not just numbers

Names, labels, categories, dates, yes/no responses, text codes, and missing-value indicators can all be part of a data set. Statistical analysis depends on the meanings attached to these entries. A code of 1 might mean “yes,” “male,” “treatment group,” or “first visit.” A data dictionary should explain the code. If a data set uses 999 to mean “not reported,” treating 999 as a genuine measurement could severely distort an average.

Data table

Rows

Rows generally identify observational units. A row might represent a student, a product, a household, a day, or a transaction.

Data table

Columns

Columns generally identify variables. Every column needs a name, definition, units when applicable, and valid coding rules.

Raw data and derived data

Raw data are values as initially recorded, although even “raw” data may already reflect measurement choices. Derived data are created from other variables. Age on an interview date can be derived from date of birth. Body mass index can be derived from height and weight. A pass/fail indicator can be derived from a score and a stated cutoff. Derived values can be useful, but the transformation should be documented.

Data quality begins before analysis

Introducing Statistics encourages basic questions about data quality: Are the intended individuals included? Are there duplicate rows? Are units consistent? Are impossible values present? Are missing values clearly marked? Were definitions applied in the same way? These questions are not advanced software tasks. They are part of understanding what the data set actually says.

Introducing Statistics checkpointA data set is an organized representation of observations, not a pile of detached values. To interpret a cell, you need its row meaning, column meaning, coding, units, and context.
Advertisement · Google AdSense in-content placement 4 reserved here
06
Population truth and sample summary

Parameters and statistics

A parameter is a numerical characteristic of a population. A statistic is a numerical characteristic calculated from a sample. Parameters and statistics often describe the same kind of feature—such as a mean or proportion—but they refer to different groups.

FeaturePopulation parameterSample statistic
MeanPopulation mean, often written μSample mean, often written x̄
ProportionPopulation proportion, often written pSample proportion, often written p̂
Standard deviationPopulation standard deviation, often written σSample standard deviation, often written s
SizePopulation size NSample size n

Example: screen time

A school district wants to know the mean weekday recreational screen time of all 10th-grade students in the district. The population mean, μ, is the parameter. It is fixed for the defined population and measurement rules, although it may be unknown. The district surveys a sample of 180 students and obtains a sample mean of x̄ = 3.6 hours. The value 3.6 hours is a statistic.

The sample statistic is used as an estimate of the population parameter. Another properly selected sample might yield 3.4 or 3.8 hours. This does not mean the population parameter changes every time a sample is drawn. It means the statistic varies from sample to sample.

Parameter does not mean “important number”

Students sometimes label a value a parameter because it appears in the question or because it seems official. The correct distinction is the group described. If 62 percent of the 250 surveyed customers were satisfied, 62 percent is a sample statistic. If company records include every customer in the defined population and show that 62 percent were satisfied, the value is a population parameter for that population.

A statistic can describe a complete observed population

In informal language, any numerical summary may be called a statistic. In introductory inference, however, the parameter-statistic distinction is tied to population and sample. When a school calculates the mean score of every student enrolled in a particular class, that mean is a parameter for the class population. The same class may be treated as a sample from a broader conceptual population, such as future students taught under the same method, but that broader claim requires careful definition.

Parameter: describes the population   |   Statistic: describes the sampleThe numerical operation may be identical; the group being summarized is different.

Estimand, estimate, and estimator

Three related terms deepen the distinction. An estimand is the precise population quantity the study aims to learn, such as the population mean difference in sleep duration between two groups. An estimator is the rule or formula used to estimate it, such as the difference between two sample means. An estimate is the numerical result obtained from the observed data. Topic 1.1 does not require advanced estimator theory, but the language shows why a question must specify the desired population feature.

Introducing Statistics warningThe sample size n is itself a fact about the sample, but in AP Statistics questions the word “statistic” usually refers to a calculated numerical summary such as a sample mean, median, proportion, standard deviation, or difference.
07
Why statistics is necessary

Statistical variability

Variability means that observations are not all identical and that repeated data-collection processes can produce different results. Variation is the reason statistical questions require data. If every individual had exactly the same value and every measurement were perfectly repeatable, many statistical problems would collapse into simple facts.

Variation among individuals

Students differ in commute time, sleep duration, course grades, and preferred learning methods. Manufactured parts differ slightly in length or strength. Daily temperatures change. Customer orders vary in amount. A statistical distribution records not only a typical value but also the pattern of differences across individuals.

Variation in measurement

The same characteristic may be recorded differently because instruments have limited precision, respondents interpret questions differently, observers use slightly different judgments, or conditions change. If a person’s pulse is measured twice, the results may differ because the pulse itself changes and because the measurement process is imperfect.

Variation from sample to sample

Suppose 40 percent of all students at a school use the bus. One random sample of 50 students may contain 18 bus users, giving a sample proportion of 36 percent. Another sample may contain 23 bus users, giving 46 percent. Both statistics can arise from the same population. Later units quantify this sampling variability through sampling distributions and standard errors.

Variation does not automatically mean error

Natural differences are not mistakes. Two batteries from the same production process can have different lifetimes. Two students can respond differently to the same lesson. Statistical reasoning distinguishes natural variation from measurement error, selection bias, recording mistakes, and systematic changes. Introducing Statistics begins with recognition: whenever a question concerns a variable characteristic, expect a distribution of possible values rather than one universal answer.

Question without variability

What is Maria’s student ID?

For a specified person and current record, the answer is a single fact. Data collection may be needed to look it up, but the question does not ask about a variable across a group.

Question with variability

How do student commute times vary at the school?

Different students are expected to have different times. Answering requires observations from a group and a description of the distribution.

Variability shapes the wording of conclusions

A sample average does not imply that every individual has that value. If the mean commute time is 24 minutes, some students may travel 5 minutes and others 60 minutes. If 72 percent of a sample supports a policy, the remaining 28 percent does not. Good statistical writing preserves this variation instead of turning a summary into a claim about every case.

Introducing Statistics checkpointA statistical question anticipates variation in the data. The question may ask about a typical value, a distribution, a proportion, a difference, or an association, but it expects that observations will not all be the same.
Advertisement · Google AdSense in-content placement 5 reserved here
08
The question controls the study

Investigative questions

An investigative question is a question that guides a statistical study. It identifies what the study seeks to learn and requires data from or about a group. A valid investigative question is answerable, connected to a population or process, and written so that the required variables can be identified.

The official Topic 1.1 skill is to determine a valid investigative question that requires a statistical investigation. That means recognizing more than punctuation. A sentence can end with a question mark and still fail as an investigative question. “Is exercise good?” is too vague. Good for whom, measured how, compared with what, and during what period? Introducing Statistics turns broad interests into questions that can guide evidence collection.

What a strong question reveals

  • Population: the group or process about which information is wanted.
  • Variable or variables: the characteristics that must be recorded.
  • Purpose: description, comparison, association, prediction, or another clearly stated aim.
  • Context: place, period, condition, or definition needed for a meaningful answer.
  • Variability: an expectation that observations may differ.
  • Feasibility: the needed data could reasonably be collected or obtained.

Five common families of investigative questions

Distribution

What values occur?

What is the distribution of weekday sleep duration among 11th-grade students at Pine High School?

Typical value

What is typical?

What is the median weekday wait time for Route 8 passengers during the morning rush?

Proportion

How common?

What proportion of households in Lake County used curbside recycling last month?

Comparison

How do groups differ?

How does average battery life compare between Brand A and Brand B earbuds tested under the same conditions?

Association

How are variables related?

Is weekly practice time associated with free-throw percentage among players in the regional league?

Change

How does a value change?

How did monthly library visits change from January through December 2026?

Descriptive and inferential versions

“What percentage of the 80 surveyed students bring lunch from home?” is descriptive because it asks about the observed sample. “What proportion of all students at the school bring lunch from home?” is inferential if only a sample is observed. The second question concerns a population parameter and requires a defensible sampling process if the result is to be generalized.

Questions create variable requirements

Consider the question, “Is weekly study time associated with mathematics exam score among first-year college students?” The study needs at least two variables for each student: weekly study time and exam score. It also needs a definition of first-year student, a time period for study hours, and a specified exam or score. If the question compares online and in-person sections, course format becomes another required variable.

Introducing Statistics principleBefore collecting data, underline the population, circle the variable or variables, and box the comparison or relationship. If those components cannot be identified, the question probably needs revision.
09
Does the question expect a distribution?

Statistical versus nonstatistical questions

A statistical question anticipates variability and is answered by collecting, analyzing, or summarizing data. A nonstatistical question seeks a single fixed fact, definition, or calculation that does not require understanding variation across observations. The distinction depends on the intended question, not simply on whether a number appears in the answer.

QuestionClassificationReason
How many seats are listed on Bus 24?NonstatisticalFor the specified bus and current configuration, the question seeks one fixed count.
How many passengers ride Bus 24 on weekday mornings?StatisticalPassenger counts vary by day, date, weather, and other conditions.
What is Jamal’s height today?Usually nonstatisticalIt asks for one measurement on one identified individual, although repeated measurement could introduce another question.
What is the distribution of heights among students in the class?StatisticalStudents have different heights, so data from a group must be summarized.
Does the school cafeteria open at 7:00 a.m.?NonstatisticalThe official schedule provides a fixed answer for the defined day.
At what time do students usually arrive at the cafeteria?StatisticalArrival time varies among students and days.

Examples and nonexamples

NonexampleWhat is the best phone?
Statistical versionAmong students at Central High, how do ratings of battery life, price, camera quality, and overall satisfaction vary across the three most commonly owned phone models?
NonexampleAre school lunches healthy?
Statistical versionWhat proportion of lunches served by the district during October meet the district’s stated limits for sodium and added sugar?
NonexampleDo students sleep enough?
Statistical versionWhat is the distribution of self-reported school-night sleep duration among 9th-grade students at West High during the spring semester?

A question can become statistical through repeated observations

“What is the temperature at noon today?” seeks one measurement. “How does the noon temperature vary across the 30 days of April?” is statistical. “What is the battery life of this device?” might seek one test result, while “What distribution of battery life is produced when this model is tested repeatedly under a standard protocol?” expects variability.

A statistical question is not automatically a good question

“What do people think?” anticipates variation, but the population and subject are undefined. “Is there a relationship between everything?” is impossible to operationalize. A valid statistical question must be sufficiently specific to identify data, units, and scope.

Introducing Statistics checkpointUse the variability test: Could reasonable observations differ? If yes, the question may be statistical. Then use the clarity test: Are the population, variables, and context defined well enough to collect relevant data?
Advertisement · Google AdSense in-content placement 6 reserved here
10
A practical writing method

How to write an investigative question

A strong investigative question can be built deliberately. Introducing Statistics uses the POP-V-C method: identify the POPulation, name the Variable or variables, and state the Comparison, connection, or contextual boundary.

Step 1: Start with a real purpose

Begin with the decision or understanding the study should support. A school may need to adjust bus schedules, a manufacturer may need to evaluate battery consistency, or a library may need to plan staffing. The purpose should not dictate the result, but it helps identify which information is relevant.

Step 2: Define the population

Replace broad words with an operational boundary. Instead of “teenagers,” write “students ages 14–18 enrolled at the three district high schools on September 15, 2026.” Instead of “customers,” write “customers who completed an online order from the company’s U.S. store during June 2027.”

Step 3: Name measurable variables

Replace abstract ideas with recordable characteristics. “Engagement” might become number of voluntary discussion posts, minutes active in the course platform, or a score from a defined survey. Different operational definitions answer different questions. A variable should include units or categories when necessary.

Step 4: State the intended statistical feature

Decide whether the question asks for a distribution, mean, median, proportion, difference, association, or trend. Topic 1.1 does not require choosing formulas, but a clear target helps the study collect the right data.

Step 5: Include context and time

Conditions can change. Commute time during the morning rush differs from late evening. Customer satisfaction may depend on the product version. A time period, location, or operating condition often belongs in the question.

Step 6: Check that variability is expected

If the answer is a fixed definition or a single known record, the question may not require statistical investigation. A valid statistical question should anticipate differences among individuals, occasions, or repeated samples.

Step 7: Remove loaded wording

“Why do irresponsible drivers speed?” assumes irresponsibility and that the behavior occurs. A more neutral question is, “What proportion of licensed drivers in the county report exceeding the posted limit by at least 10 miles per hour during the previous month?” Neutral wording does not guarantee unbiased measurement, but it avoids building a preferred conclusion into the question.

Among [population], what is the [distribution / mean / proportion / difference / association] of [variable or variables] during [context or time]?A flexible template, not a sentence that must be copied mechanically.

Question repair workshop

Weak questionProblemImproved investigative question
Do students like math?Undefined students, undefined “like,” no time or measurement.Among students enrolled in Algebra II at Hill High this semester, what is the distribution of ratings on a 1–5 mathematics-interest scale?
Is the bus late?One bus? Which route? What counts as late? Over what period?During October weekdays, what proportion of Route 12 arrivals at Central Station occur more than five minutes after the scheduled time?
Which class is better?“Better” is not operationally defined.How do final exam scores and course-satisfaction ratings compare between the online and in-person sections of Introductory Biology in spring 2027?
Does coffee help?Population, dose, outcome, and comparison are missing.Among adult volunteers who normally consume less than one cup of coffee per day, how does reaction time compare after a standardized caffeinated drink and a visually identical noncaffeinated drink?
Are phones distracting?“Distracting” needs measurement and context.Among 10th-grade students completing a 30-minute reading task, how does the number of comprehension questions answered correctly differ between a phone-visible condition and a phone-stored-away condition?

A self-check before data collection

  • Can I identify the population without guessing?
  • Can I identify each observational unit?
  • Can I name the variable or variables and how they are recorded?
  • Does the question anticipate variability?
  • Is the time, place, or condition clear enough?
  • Could data realistically answer the question?
  • Does the wording avoid assuming the conclusion?
  • Will the planned statistic or graph connect directly to the question?
Introducing Statistics writing ruleA strong question is specific enough to guide a study but not written so narrowly that it predetermines the answer. It names what will be learned, not what the researcher hopes will be true.
11
See every component working together

Complete worked cases

The following original cases show how Introducing Statistics concepts fit together. Each case begins with a purpose, states an investigative question, and maps the population, sample, individuals, variables, parameter, statistic, and expected sources of variability.

Worked case A

Morning commute to school

A district transportation office wants to review the morning schedule at East High. It selects 150 students from the current enrollment list and records the number of minutes from leaving home to entering the school building on one regular school day. The sample mean is 31.8 minutes, and the sample median is 27 minutes.

Investigative questionWhat is the distribution of morning commute time among students currently enrolled at East High on regular school days?
PopulationAll students currently enrolled at East High, under the stated regular-day context.
SampleThe 150 selected students whose commute times were recorded.
IndividualsIndividual East High students.
VariableMorning commute time in minutes, defined from leaving home to entering the school building.
StatisticsSample mean 31.8 minutes and sample median 27 minutes.
ParametersThe population mean and population median commute times for all enrolled students under the stated conditions.
VariabilityStudents live at different distances, use different travel modes, leave at different times, and encounter different traffic conditions.

The mean and median describe the sample, not every student. The difference between 31.8 and 27 suggests that longer commutes may pull the mean upward, but a graph would be needed to study the distribution properly. The sample selection method determines whether the statistics can credibly represent the population.

Worked case B

Wireless-earbud battery life

A quality-control team selects 60 pairs of earbuds from the 8,400 pairs produced during one week. Each pair is fully charged and played continuously at a fixed volume until shutdown. The sample mean operating time is 7.42 hours.

Investigative questionWhat is the mean continuous operating time of earbud pairs produced during the specified week under the standard test?
PopulationAll 8,400 earbud pairs produced during that week.
SampleThe 60 selected pairs tested.
Observational unitsIndividual earbud pairs.
VariableContinuous operating time in hours under the defined volume and shutdown rule.
Statisticx̄ = 7.42 hours.
Parameterμ, the mean operating time of all 8,400 pairs under the same test.
VariabilitySmall differences in cells, assembly, charging, electronics, and measurement conditions produce different lifetimes.

The test conditions are part of the variable definition. The result does not automatically describe battery life at every volume, with every device, or under intermittent use. Introducing Statistics emphasizes that a parameter belongs to a defined population under a defined measurement process.

Worked case C

Library program participation

A city library system wants to estimate the proportion of registered teen members who attended at least one library program during the summer. From the membership database, 400 teen members are selected; 148 attended at least one program.

Investigative questionWhat proportion of registered teen library members attended at least one program during the defined summer period?
PopulationAll registered teen members in the library system during that summer.
SampleThe 400 selected teen members.
IndividualsRegistered teen members.
VariableWhether the member attended at least one program during the summer: yes or no.
Statisticp̂ = 148/400 = 0.37, or 37 percent.
Parameterp, the true proportion of all registered teen members who attended.
VariabilityAttendance differs among members, and a different sample could contain a different proportion of attendees.

A contextual interpretation is: “In the sample, 37 percent of the 400 selected registered teen members attended at least one summer program.” Calling 37 percent the population proportion would be unjustified unless every member were included or a later inferential argument supported the estimate.

Worked case D

Comparing two garden fertilizers

A horticulture class grows 48 tomato plants of the same variety under similar greenhouse conditions. Twenty-four plants receive Fertilizer A and 24 receive Fertilizer B. After eight weeks, students record each plant’s increase in height.

Investigative questionHow does eight-week height increase compare between tomato plants receiving Fertilizer A and those receiving Fertilizer B under the greenhouse conditions?
Population or process of interestTomato plants of the specified variety grown under comparable greenhouse conditions.
Observed unitsThe 48 plants in the study.
VariablesFertilizer group and height increase in centimeters.
Possible statisticThe difference between the sample mean height increases for the two groups.
Possible parameterThe difference between mean height increases in the broader processes or populations represented by the two treatment conditions.
VariabilityPlants differ genetically and biologically, and greenhouse microconditions or measurement may vary.

This question is statistical because plant growth varies. It is comparative because it involves two treatment groups. Later lessons will determine what the design permits the class to conclude. Topic 1.1 focuses on identifying the components before analyzing the results.

Worked case E

Streaming-session duration

A service studies 10,000 viewing sessions selected from sessions started by U.S. subscribers during one month. Each record contains session duration, device type, content category, and whether the session ended voluntarily or because of an error.

Investigative questionHow does session duration vary by device type and ending status among U.S. subscriber sessions started during the month?
PopulationAll qualifying U.S. subscriber sessions started during the month.
SampleThe 10,000 selected sessions.
Observational unitsViewing sessions, not necessarily unique subscribers.
VariablesDuration, device type, content category, and ending status.
StatisticsSample medians, means, proportions, or group differences calculated from the 10,000 sessions.
ParametersCorresponding numerical features of all qualifying sessions in the month.
VariabilitySessions vary in viewer behavior, content length, device conditions, and technical performance.

The observational unit is a session because one subscriber can contribute multiple sessions. If the question concerned average monthly viewing per subscriber, the unit and data organization would need to change. Introducing Statistics always ties the unit to the question.

Advertisement · Google AdSense in-content placement 7 reserved here
12
Apply the language

Original Introducing Statistics exercises

These exercises are original and designed specifically for this lesson. They emphasize identification before calculation. For each scenario, write answers in complete contextual phrases. Do not answer “population = everyone” unless the study truly concerns everyone; name the exact group.

Introducing Statistics exercise methodRead the final purpose first. Then mark the full group of interest, the observed subset, the row unit, the recorded characteristic, the population quantity, the sample summary, and the question that connects them.

Introducing Statistics Exercise 1

A county health office wants the mean number of hours adults in the county sleep on a typical weeknight. It selects 320 adults and obtains a sample mean of 6.7 hours. Identify the population, sample, individual, variable, parameter, statistic, and investigative question.

Introducing Statistics Exercise 2

A bakery produced 2,400 loaves on Monday. Quality staff randomly inspect 80 loaves and find that 6 have an underweight package. Identify the population, sample, observational unit, variable, parameter, statistic, and investigative question.

Introducing Statistics Exercise 3

A principal records the final mathematics grade of every one of the 214 students in the graduating class. Is the group of 214 a sample or the population for the question, ‘What is the mean final mathematics grade of this graduating class?’ Explain.

Introducing Statistics Exercise 4

A researcher downloads 5,000 product reviews from all 68,000 verified reviews posted during 2026. Each row represents one review. Identify N, n, the population, the sample, and the observational unit.

Introducing Statistics Exercise 5

A study of 75 households records household size, monthly electricity use, and whether the home has solar panels. Identify the individuals and the variables.

Introducing Statistics Exercise 6

A survey of 180 students reports that 117 support extending library hours. Identify the sample statistic and state the corresponding population parameter in words.

Introducing Statistics Exercise 7

A delivery company states that the mean delivery time for all packages delivered last month was 2.8 days. The figure was calculated from the company’s complete database of every package in that month. Is 2.8 days a parameter or statistic for that defined population?

Introducing Statistics Exercise 8

A class asks, ‘What is the height of the school flagpole?’ Is this statistical or nonstatistical? Explain.

Introducing Statistics Exercise 9

A class asks, ‘How do the heights of flagpoles at public high schools in the state vary?’ Is this statistical or nonstatistical? Explain.

Introducing Statistics Exercise 10

Improve this question: ‘Do people use the internet too much?’ Write a measurable investigative question with a defined population, variable, and time context.

Introducing Statistics Exercise 11

A wildlife team tags 90 turtles from a lake and records shell length in centimeters. Its goal is to estimate the mean shell length of all adult turtles in the lake. Identify population, sample, individuals, variable, parameter, and statistic if the sample mean is 24.6 cm.

Introducing Statistics Exercise 12

A spreadsheet has 40 patients and five blood-pressure readings per patient. It contains 200 measurement rows. For a question about patients’ average blood pressure, what is the participant sample size? Why is 200 not automatically the number of independent individuals?

Introducing Statistics Exercise 13

A school data set contains one row per course enrollment. A student enrolled in six courses appears in six rows. What is the observational unit: student or course enrollment? Explain.

Introducing Statistics Exercise 14

A manufacturer tests 50 light bulbs until failure. The bulbs come from 10,000 bulbs produced in one shift. The sample median lifetime is 1,140 hours. Identify the parameter that the sample median might estimate.

Introducing Statistics Exercise 15

Classify the question: ‘How many pages are in the printed school handbook?’ Statistical or nonstatistical?

Introducing Statistics Exercise 16

Classify the question: ‘How many pages do students read from the handbook during the first week?’ Statistical or nonstatistical?

Introducing Statistics Exercise 17

A city wants to learn about commuting among all employed residents, but its contact list includes only residents who registered a vehicle. Distinguish the target population from the accessible population.

Introducing Statistics Exercise 18

A data table uses 999 to represent a missing response for weekly exercise minutes. Explain why treating 999 as an actual number could damage the analysis.

Introducing Statistics Exercise 19

A café records one row per transaction, including purchase amount, payment method, time, and whether a discount was used. Identify the individuals and variables.

Introducing Statistics Exercise 20

A teacher asks, ‘Which teaching method is best?’ Rewrite the question as a valid comparative investigative question.

Introducing Statistics Exercise 21

A sample proportion changes from 0.42 in one random sample to 0.47 in another. Does this alone prove that the population proportion changed? Explain using sampling variability.

Introducing Statistics Exercise 22

A study question is, ‘What percentage of the 250 surveyed voters favor Proposal A?’ Is the requested value a parameter or statistic? Explain.

Introducing Statistics Exercise 23

A different question asks, ‘What percentage of all registered voters in the city favor Proposal A?’ If only 250 voters are surveyed, what population quantity is being targeted?

Introducing Statistics Exercise 24

A student writes, ‘The population is the 100 people surveyed because those are all the people in the data.’ Correct the statement.

Introducing Statistics Exercise 25

Write an investigative question about school lunch wait time that clearly identifies the population, variable, and time period.

Introducing Statistics Exercise 26

Write an investigative question comparing battery life for two phone models under a common test condition.

Introducing Statistics Exercise 27

Write an investigative question about an association between two variables among high school students.

Introducing Statistics Exercise 28

For the question ‘What is the distribution of weekly paid-work hours among 12th-grade students at Lake High this semester?’ identify the population, individuals, and variable.

Introducing Statistics Exercise 29

A study collects ratings from 1 to 5 but does not explain what 1 and 5 mean. What data-documentation problem is present?

Introducing Statistics Exercise 30

Create a complete statistical study map for this situation: a community center selects 200 members to estimate the proportion who used the fitness room at least once in March; 86 selected members did so.

Advertisement · Google AdSense in-content placement 8 reserved here
13
Reasoning, not just labels

Complete solutions

Open each solution after attempting the exercise. The wording demonstrates how to keep the answer in context. Equivalent answers may be correct when they define the same population, sample, unit, and variable clearly.

Introducing Statistics Solution 1
  • Population: all adults in the county under the stated weeknight definition.
  • Sample: the 320 selected adults.
  • Individual: one selected adult.
  • Variable: typical weeknight sleep duration in hours.
  • Parameter: the population mean sleep duration for all county adults.
  • Statistic: x̄ = 6.7 hours.
  • Investigative question: What is the mean typical-weeknight sleep duration among adults in the county?
Introducing Statistics Solution 2
  • Population: all 2,400 loaves produced Monday.
  • Sample: the 80 inspected loaves.
  • Observational unit: one loaf.
  • Variable: whether the loaf has an underweight package, yes or no.
  • Parameter: the true proportion of all Monday loaves that are underweight.
  • Statistic: p̂ = 6/80 = 0.075, or 7.5 percent.
  • Investigative question: What proportion of Monday’s loaves are underweight?
Introducing Statistics Solution 3
  • For that exact question, the 214 students are the population because every member of the defined graduating class is included.
  • The calculated class mean is therefore a population parameter for that graduating class.
  • The same students could be treated as a sample only if the question were broadened to a larger conceptual population, which would require additional justification.
Introducing Statistics Solution 4
  • N = 68,000 reviews.
  • n = 5,000 reviews.
  • Population: all 68,000 verified reviews posted during 2026.
  • Sample: the 5,000 downloaded reviews.
  • Observational unit: one verified product review.
Introducing Statistics Solution 5
  • Individuals: the 75 households.
  • Variables: household size, monthly electricity use, and solar-panel status.
  • The data set should also document units for electricity use and the month or billing period represented.
Introducing Statistics Solution 6
  • Sample statistic: p̂ = 117/180 = 0.65, or 65 percent.
  • Population parameter: the true proportion of all students in the defined school population who support extending library hours.
Introducing Statistics Solution 7
  • It is a parameter for the defined population because every delivered package in that month was included.
  • If the company used the month as a sample of future operations, 2.8 could also function as a sample summary for a broader process, but that is a different question.
Introducing Statistics Solution 8
  • Nonstatistical in its ordinary form.
  • The question seeks one fixed measurement for one specified flagpole rather than a distribution across a group.
Introducing Statistics Solution 9
  • Statistical.
  • Different schools are expected to have flagpoles of different heights, so the answer requires data from many schools and a description of variability.
Introducing Statistics Solution 10
  • One valid answer: Among students enrolled at Central High during fall 2027, what is the distribution of self-reported non-school internet use in hours per weekday?
  • Other answers are acceptable if they define a population, measurable variable, and context without using the loaded phrase ‘too much.’
Introducing Statistics Solution 11
  • Population: all adult turtles in the lake.
  • Sample: the 90 tagged turtles.
  • Individuals: individual adult turtles.
  • Variable: shell length in centimeters.
  • Parameter: the population mean shell length of all adult turtles in the lake.
  • Statistic: x̄ = 24.6 cm.
Introducing Statistics Solution 12
  • The participant sample size is 40 patients.
  • There are 200 measurements, but the five readings from one patient are linked to the same person and may be more similar to one another than readings from different patients.
  • Counting 200 as independent individuals would confuse measurements with participants.
Introducing Statistics Solution 13
  • The observational unit in the stated table is a course enrollment because each row describes one student-course combination.
  • A separate student-level table would use one row per student.
Introducing Statistics Solution 14
  • The corresponding parameter is the population median lifetime of all 10,000 bulbs produced in that shift under the same test conditions.
Introducing Statistics Solution 15
  • Nonstatistical. The printed handbook has a fixed number of pages for that edition.
Introducing Statistics Solution 16
  • Statistical. Students are expected to read different numbers of pages, so the question concerns a distribution across students.
Introducing Statistics Solution 17
  • Target population: all employed residents of the city.
  • Accessible population: employed residents who are included on the vehicle-registration contact list.
  • Residents without registered vehicles may be systematically different in commuting behavior, so the accessible group may not represent the target population.
Introducing Statistics Solution 18
  • The value 999 is a missing-value code, not 999 minutes of exercise.
  • Including it as a genuine measurement would create extreme false values and could greatly inflate the mean and spread.
  • The code should be documented and treated as missing.
Introducing Statistics Solution 19
  • Individuals or observational units: transactions.
  • Variables: purchase amount, payment method, transaction time, and discount-use status.
  • Customers are not the row units unless the data are reorganized to one row per customer.
Introducing Statistics Solution 20
  • One valid answer: Among students enrolled in Algebra I at North High, how do end-of-unit test scores compare between classes using Method A and classes using Method B during the fall semester?
  • The response variable and population are defined, and the comparison is explicit.
Introducing Statistics Solution 21
  • No. Different random samples from the same population can produce different sample proportions.
  • The change from 0.42 to 0.47 may be ordinary sampling variability.
  • Evidence that the population changed would require a design and analysis that distinguish real change from sample-to-sample fluctuation.
Introducing Statistics Solution 22
  • Statistic, because the percentage is calculated for the 250 surveyed voters—the observed sample.
Introducing Statistics Solution 23
  • The targeted parameter is p, the true proportion of all registered voters in the city who favor Proposal A.
  • The sample proportion from 250 voters would be used as an estimate of that parameter.
Introducing Statistics Solution 24
  • The 100 surveyed people are the sample because they are the observed subset.
  • The population is the larger group the study seeks to describe, such as all eligible residents, customers, or students defined by the question.
Introducing Statistics Solution 25
  • One valid answer: What is the distribution of minutes from joining the lunch line to receiving food among students buying lunch at West High during regular school days in October?
Introducing Statistics Solution 26
  • One valid answer: Under continuous video playback at 50 percent screen brightness, how does battery life in hours compare between new Phone Model A and Phone Model B devices?
Introducing Statistics Solution 27
  • One valid answer: Among students in grades 9–12 at River High, is school-night sleep duration associated with first-period tardiness during the fall semester?
Introducing Statistics Solution 28
  • Population: all 12th-grade students enrolled at Lake High during the semester.
  • Individuals: individual 12th-grade students.
  • Variable: weekly paid-work hours.
Introducing Statistics Solution 29
  • The rating scale lacks a coding definition or data dictionary.
  • Readers cannot know whether 1 means very dissatisfied or very satisfied, and they cannot interpret the values reliably.
Introducing Statistics Solution 30
  • Investigative question: What proportion of all community-center members used the fitness room at least once in March?
  • Population: all community-center members during the defined period.
  • Sample: the 200 selected members.
  • Individuals: members.
  • Variable: whether each member used the fitness room at least once in March.
  • Statistic: p̂ = 86/200 = 0.43, or 43 percent.
  • Parameter: the true proportion of all center members who used the fitness room at least once in March.
  • Variability: members differ in use, and another sample could produce a different proportion.
14
Language that protects meaning

Common errors and precise corrections

Error 1: Calling the sample the population

A data set contains only the observed cases, so students often call those cases the population. Correct the error by returning to the question. If the study wants to describe all 8,000 district students but surveys 300, the 8,000 students form the population and the 300 surveyed students form the sample.

Error 2: Calling every reported number a parameter

A parameter describes the entire population. A sample mean, sample median, sample proportion, or sample standard deviation is a statistic. The size or authority of the organization reporting the number does not change this definition.

Error 3: Naming a topic instead of a variable

“Health” is a topic, not a sufficiently defined variable. Resting heart rate in beats per minute, number of sick days during the semester, or a score on a named health questionnaire are variables. Introducing Statistics expects a characteristic that can be recorded for each unit.

Error 4: Confusing individuals with variables

In a student survey, students are individuals; grade level and commute time are variables. In a table of schools, schools are individuals; enrollment and graduation rate are variables. A quick check is to ask: Does this label identify a row or a column?

Error 5: Treating repeated measurements as new individuals

Ten readings from one machine are ten observations, but they may belong to one machine unit. The sample size at the machine level is one. A study can analyze readings, machines, or machine-time combinations, but the unit must match the question and design.

Error 6: Writing a question that hides the denominator

“What percent preferred option A?” is incomplete unless the denominator is known. Was the percentage among all invited people, all respondents, all eligible participants, or all people who answered that item? A proportion always compares a count with a defined total.

Error 7: Using a summary as if it describes every individual

If the sample mean is 14.2, it does not follow that every person scored 14.2. If 60 percent answered yes, it does not follow that a typical person is “60 percent yes.” Summaries describe distributions or groups.

Error 8: Assuming more rows guarantee better evidence

A very large convenience sample can be systematically unrepresentative. Sample size influences random variability, but selection methods influence bias and generalizability. Topic 1.1 introduces the distinction; later topics develop sampling design.

Error 9: Forgetting the time frame

A population can change. “All employees” today may differ from all employees last year. A question about monthly spending needs a stated month or typical-month definition. Time is part of the context, not an optional detail.

Error 10: Writing a conclusion without the study noun

“The mean is 7.4” is incomplete. “The sample mean continuous operating time of the 60 tested earbud pairs was 7.42 hours under the standard test” is contextual. Introducing Statistics builds the habit of naming the variable, units, and sample.

Introducing Statistics precision testA reader who did not see the original problem should still be able to understand what your population, sample, variable, parameter, and statistic refer to.
Advertisement · Google AdSense in-content placement 9 reserved here
15
A complete identification challenge

Mini assessment: one story, many components

A regional transit authority operates 340 buses. It wants to estimate the proportion of buses that require an unscheduled repair during a typical 30-day operating period. Engineers select 70 buses at the beginning of April and follow each bus for 30 days. Fourteen selected buses require at least one unscheduled repair. The data set contains one row per selected bus and columns for route group, bus age, miles traveled, number of repairs, and whether at least one unscheduled repair occurred.

Questions

  1. State the investigative question.
  2. Identify the population and its size.
  3. Identify the sample and its size.
  4. Identify the observational units.
  5. List the variables named in the story.
  6. Identify the sample statistic most directly connected to the question.
  7. State the corresponding population parameter.
  8. Give two reasons the observed results may vary among buses.
  9. Explain why “14” is not the sample proportion.
  10. Write a one-sentence contextual description of the sample result.
Open the complete Introducing Statistics mini-assessment solution
  1. What proportion of the authority’s buses require at least one unscheduled repair during a 30-day operating period comparable to the study period?
  2. Population: all 340 buses operated by the authority; N = 340.
  3. Sample: the 70 selected buses; n = 70.
  4. Observational units: individual buses.
  5. Route group, bus age, miles traveled, number of repairs, and whether at least one unscheduled repair occurred.
  6. p̂ = 14/70 = 0.20, or 20 percent.
  7. p, the true proportion of all 340 buses that would require at least one unscheduled repair during the defined 30-day conditions.
  8. Possible sources include bus age, mileage, route conditions, maintenance history, driver use, component differences, and random failures.
  9. The number 14 is the count of selected buses with a repair. A proportion requires dividing that count by the sample total of 70.
  10. In the sample, 20 percent of the 70 selected buses required at least one unscheduled repair during the 30-day observation period.

Why this assessment matters

The story contains several numbers, but only some answer the investigative question. N = 340 and n = 70 describe group sizes. Fourteen is a count. The value 0.20 is the sample proportion. The unknown population proportion is the parameter. Introducing Statistics teaches how to assign each number a role rather than treating every number as interchangeable.

16
Keep the language connected

Introducing Statistics glossary

Accessible population
The part of the target population that can realistically be reached through the available list, location, system, or time period.
Census
An attempt to collect data from every member of a defined population.
Context
The real-world meaning that connects statistical terms and calculations to the people, objects, variables, units, place, and time in the study.
Data
Recorded pieces of information about observational units.
Data set
An organized collection of data with labels, definitions, units, and structure.
Datum
One recorded piece of information; the singular form of data.
Descriptive statistics
Methods used to organize, display, and summarize observed data.
Estimate
A numerical value calculated from sample data to approximate a population parameter.
Estimator
A rule or statistic used to estimate a parameter.
Estimand
The precise population quantity a study intends to estimate.
Individual
A person, object, case, or entity described by the data.
Investigative question
A question that guides a statistical study and requires data to answer.
N
A common symbol for population size.
n
A common symbol for sample size.
Nonstatistical question
A question seeking a fixed fact, definition, or single result without requiring analysis of variability across observations.
Observational unit
The entity represented by one record at the level relevant to the question.
Parameter
A numerical characteristic of a population.
Population
The complete group of items or individuals relevant to the investigative question.
Sample
The subset of the population from which data are obtained.
Sampling variability
The natural tendency for statistics calculated from different samples to differ.
Statistic
A numerical characteristic calculated from sample data.
Statistical inference
Using sample data to draw conclusions about a larger population while accounting for uncertainty.
Statistical question
A question that anticipates variability and requires data from a group or repeated process.
Statistical study
A study in which data are collected and analyzed to answer an investigative question about a population or process.
Target population
The full group the researcher ultimately wants to understand.
Variable
A characteristic recorded for each observational unit that can take different values.
Variability
Differences among observations, measurements, occasions, or possible samples.

The final concept map

Q

Question

Defines the population and needed variables.

S

Sample

Supplies the observed individuals and data.

T

Statistic

Summarizes the sample and may estimate a parameter.

C

Context

Returns the result to the original population and purpose.

Every later AP Statistics topic adds tools to this map. Graphs reveal distributions. Probability models random behavior. Confidence intervals estimate parameters. Hypothesis tests evaluate claims. Regression describes relationships. Yet those methods remain meaningful only when the population, sample, unit, variables, and investigative question are clear.

17
Clear answers to foundational questions

Frequently asked questions

What is the simplest meaning of statistics?

Statistics is the practice of learning from data while considering variation, context, uncertainty, and how the data were collected.

What is the difference between statistics and a statistic?

Statistics is the broader discipline. A statistic is a numerical summary calculated from sample data, such as a sample mean or sample proportion.

Can a population be infinite?

Yes. A population can be a conceptual process with indefinitely many possible outcomes, such as all future items produced under stable conditions. The population must still be defined clearly.

Is a large sample always representative?

No. A large sample can still be biased if the selection process systematically excludes or overrepresents parts of the population.

Can a sample contain the whole population?

If every population member is observed, the study is a census for that defined population rather than a sample survey. The same data can still be viewed as a sample from a broader conceptual population if the question changes.

Why is N uppercase and n lowercase?

Uppercase N is a common convention for population size, while lowercase n is commonly used for sample size. The symbols help distinguish the whole group from the observed subset.

Is every row always one person?

No. A row may represent a person, object, school, transaction, day, visit, or combined unit such as a patient-visit. The observational unit must be identified from the data structure and question.

What is the difference between an individual and a variable?

An individual is the entity described by a row. A variable is a characteristic recorded for individuals and usually represented by a column.

What is the difference between data and a data set?

Data are recorded pieces of information. A data set is the organized collection of those data with labels, definitions, and structure.

Is a percentage always a statistic?

A percentage calculated from a sample is a statistic. A percentage calculated from every member of the defined population is a parameter for that population.

Does a parameter change from sample to sample?

No. The parameter belongs to the defined population. Sample statistics change from sample to sample and are used to estimate the parameter.

Why do different samples give different answers?

Samples contain different individuals or outcomes. Natural population variation causes sample summaries to differ; this is sampling variability.

What makes a question statistical?

It anticipates variability and requires data from a group, repeated observations, or a random process to answer.

Can a yes-or-no question be statistical?

Yes. For example, ‘What proportion of households recycled last week?’ uses a yes-or-no variable across many households and expects variation.

Is ‘What is the average?’ a complete investigative question?

No. It should specify the average of which variable, for which population, and under what time or conditions.

What is the role of context?

Context tells the reader what each number, variable, group, and unit represents. It prevents technically correct but meaningless statements.

What is the quickest way to distinguish parameter and statistic?

Ask whether the number describes the full population or the observed sample. Population means parameter; sample means statistic.

What happens if the sample is the wrong group?

The statistic may accurately describe the observed sample but fail to answer the question about the intended population.

Can text responses be data?

Yes. Text can be recorded as data and later coded or analyzed, provided the collection and interpretation rules are clear.

Why introduce variables in Topic 1.1 if Topic 1.2 covers variables?

Topic 1.1 needs the basic idea that a variable is a recorded characteristic. Topic 1.2 develops variable types and classifications more fully.

What should a complete identification answer include?

Use contextual phrases: name the exact population, sample, observational unit, variable with units or categories, parameter, statistic, and investigative question.

What is the most common mistake in Introducing Statistics?

A common mistake is naming the observed sample as the population simply because only the sample appears in the data table.

Can the same number be a parameter in one question and a statistic in another?

Yes. A mean calculated from every student in one class is a parameter for that class. The class could be treated as a sample from a broader population in a different question.

Why must an investigative question avoid loaded language?

Loaded wording builds an assumption or preferred answer into the question and can influence data collection or interpretation.

What comes after Introducing Statistics?

The next AP Statistics topic studies variables in greater depth, including how characteristics are classified and represented.

18
One-sentence mastery review

Introducing Statistics concept recap

Use this final bank as a rapid conceptual check. Each statement condenses one habit developed in the chapter.

  • Introducing Statistics connects the population to the sample before interpreting a number.
  • Introducing Statistics asks what one row represents before counting observations.
  • Introducing Statistics distinguishes a population parameter from a sample statistic.
  • Introducing Statistics treats variability as expected information rather than an inconvenience.
  • Introducing Statistics requires every investigative question to identify measurable information.
  • Introducing Statistics keeps units, labels, and time frames attached to the data.
  • Introducing Statistics uses contextual language so a conclusion answers the original question.
  • Introducing Statistics recognizes that a large data set can still come from a weak selection process.
  • Introducing Statistics separates a fixed factual question from a question that expects a distribution.
  • Introducing Statistics turns broad interests into populations, variables, and answerable comparisons.
  • Introducing Statistics connects the population to the sample before interpreting a number.
  • Introducing Statistics asks what one row represents before counting observations.
  • Introducing Statistics distinguishes a population parameter from a sample statistic.
  • Introducing Statistics treats variability as expected information rather than an inconvenience.
  • Introducing Statistics requires every investigative question to identify measurable information.
  • Introducing Statistics keeps units, labels, and time frames attached to the data.
  • Introducing Statistics uses contextual language so a conclusion answers the original question.
  • Introducing Statistics recognizes that a large data set can still come from a weak selection process.
  • Introducing Statistics separates a fixed factual question from a question that expects a distribution.
Advertisement · Google AdSense bottom placement reserved here

Introducing Statistics is the language of every later method

You now have a complete map for reading a statistical study: begin with the investigative question, identify the population and sample, define the observational units and variables, organize the data set, separate the sample statistic from the population parameter, and expect variability. These habits protect meaning before any graph or formula is used.

Previous course chapter: What Is the AP Statistics Exam?

Next course topic: Topic 1.2, Variables. The next visual chapter will develop categorical and quantitative variables, roles of variables, units, coding, and classification.

Course alignment note: This independent educational chapter follows the required ideas of AP Statistics Topic 1.1 in the Course and Exam Description effective Fall 2026. All scenarios, exercises, explanations, diagrams, and solutions in this post are original. AP® and Advanced Placement® are registered trademarks of the College Board. College Board does not sponsor or endorse this independent resource.

Back to the top ↑

Need help applying this to your own data?

Salar Cafe can help interpret output, clean datasets, review assumptions, build dashboards and explain statistical results ethically.

Need help interpreting your data analysis results?

Contact Salar Cafe
Engr. Muhammad Yar Saqib author profile photo

Engr. Muhammad Yar Saqib

WhatsApp Get Data Analysis Help