Introducing Statistics
What can we learn from data? Begin with the language of a statistical study: the population we care about, the sample we observe, the individuals that supply information, the variables recorded, and the question that gives every number a purpose.
Introducing Statistics begins with a simple idea that is easy to say but powerful to apply: data do not explain themselves. A spreadsheet can contain thousands of values, yet those values become statistical evidence only when we know who or what was observed, which larger group matters, what was measured, how the observations vary, and which question the study is intended to answer.
Imagine that a school principal is shown the number 18.7. Is it encouraging, worrying, or meaningless? It might be the average number of minutes students wait for a bus, the percentage of seniors absent on a particular day, the mean score on a twenty-point quiz, or the number of books borrowed per student during a semester. The same numerical symbol can support completely different interpretations. Statistical language supplies the missing context.
This chapter treats a statistical study as a connected story. The story starts with an investigative question. The question points to a population and to one or more characteristics that must be recorded. Because observing the entire population may be impossible, expensive, slow, or unnecessary, researchers often examine a sample. Each item or person observed is an individual or observational unit. The recorded pieces of information become data, and the organized collection becomes a data set. A numerical description of the population is a parameter; a numerical description calculated from the sample is a statistic.
The official AP Statistics framework places Topic 1.1 at the beginning because every later graph, probability, confidence interval, test, and regression model depends on this vocabulary. Introducing Statistics is therefore not a list of definitions to memorize. It is training in how to read, design, and explain a study without losing the real-world meaning of the evidence.
What statistics means
Statistics is the science and practice of learning from data while accounting for variation, uncertainty, context, and the way the data were produced. The word also has a second, narrower meaning: a statistic is a numerical value calculated from sample data. These two meanings are related but not identical. A course in statistics teaches a way of reasoning; a statistic is one output within that reasoning.
Introducing Statistics should change how a reader reacts to claims. Instead of asking only, “What number was reported?” a statistically minded reader asks, “What was the question? Who or what did the data describe? How were observations selected? What exactly was measured? How much did the observations differ? Does the reported number describe the sample or claim something about a larger population?” These questions prevent a polished calculation from being mistaken for strong evidence.
Statistics is often divided into two broad activities. Descriptive statistics organize and summarize the data that were actually observed. Tables, graphs, percentages, averages, and measures of spread can all describe a sample or a complete population. Statistical inference uses sample data to learn about a larger population, while acknowledging that another sample could have produced a different result. Topic 1.1 introduces the language required for both activities, even though formal inference is developed later in the course.
Statistics is more than arithmetic
Arithmetic can tell us that the mean of 4, 6, and 8 is 6. Statistics asks whether those three values came from the people or objects relevant to the question, whether the measurement was meaningful, whether 6 is an appropriate summary, how the values varied around 6, and whether the result can be generalized. A correct calculation can coexist with a poor study. Conversely, a thoughtfully designed study may produce a result with unavoidable uncertainty. Statistics makes that uncertainty visible rather than pretending it does not exist.
For example, suppose a student surveys five close friends and finds that all five prefer online homework. The sample proportion is 1.00, or 100 percent. The arithmetic is correct. However, if the intended population is all 1,200 students in the school, the selection process is unlikely to represent the population. The result is a statistic describing those five friends. It is not automatically a trustworthy estimate of the population parameter.
Statistics is a complete investigation cycle
Question
State what you want to learn and about whom or what.
Data
Decide what information is needed and obtain it responsibly.
Analysis
Represent, summarize, compare, or model the observations.
Conclusion
Answer the question in context and state the limits.
In practice, these stages interact. A preliminary look at data may reveal that a variable was recorded ambiguously. A planned analysis may show that the question needs a clearer population. A conclusion may lead to a better follow-up question. Introducing Statistics teaches the vocabulary needed to move around this cycle without confusing one component with another.
The anatomy of a statistical study
A statistical study can be read as a set of connected components. If one component is missing, the meaning of the study becomes uncertain. Introducing Statistics uses the following map throughout the chapter.
From a real-world question to statistical evidence
A miniature study
A public library wants to know the typical number of minutes adult visitors spend in the building on weekday afternoons. During randomly selected weekday afternoon periods, staff record the visit length of 120 adult visitors. The sample mean is 46.3 minutes.
| Component | Identification | Why it matters |
|---|---|---|
| Investigative question | What is the typical weekday-afternoon visit length for adult visitors to this library? | It defines the purpose and context. |
| Population | All adult visitors to the library on weekday afternoons during the period of interest. | This is the group the library wants to understand. |
| Sample | The 120 adult visitors whose visit lengths were recorded. | These are the observed cases. |
| Individuals | Individual adult library visits, or visitors if each visitor appears only once. | Each row must represent one clearly defined unit. |
| Variable | Visit length in minutes. | This is the characteristic measured on each unit. |
| Statistic | The sample mean, 46.3 minutes. | It summarizes the observed sample. |
| Parameter | The true mean visit length for the full population. | It is the population quantity the library would like to know. |
Notice that 46.3 is not simply “the average visit length.” A precise statement says it is the sample mean visit length for the 120 observed adult visits. Whether it estimates the population mean well depends on how the sample was selected and how visit length was recorded. Introducing Statistics trains this precision from the beginning.
Context is not decoration
In statistics, “in context” means matching every abstract term and calculation to the real-world component it represents. If a calculation produces 0.37, a contextual interpretation might be “37 percent of the sampled households reported using public transportation at least once during the previous week.” Merely writing “the proportion is 0.37” omits the group, behavior, and time frame.
Context also prevents unit errors. A mean of 12.4 could mean dollars, hours, points, kilograms, or customer complaints. A standard deviation of 3 has the same units as the variable it describes. A percentage must specify the denominator. A sample size must refer to the number of observational units, not automatically the number of measurements if repeated observations are present.
Population and sample
The population is the complete collection of items or individuals relevant to an investigative question. The symbol N commonly represents population size. The sample is the subset from which data are actually obtained, and n commonly represents sample size. A sample is not defined merely by being small; it is defined by being a part of a larger population of interest.
Population, size N
Every item or individual relevant to the question
Sample, size n
The subset actually observed
Defining the population precisely
A useful population definition usually includes more than a broad label. “Students” is rarely precise enough. Which students? At which schools? During what academic year? Enrolled at what point in time? Eligible under which conditions? The population might be “all students enrolled in grades 9–12 at North Valley High School on October 1, 2026.” That definition creates a boundary. A student who transferred after October 1 may not belong to the defined population even though the student later attends the school.
Populations can contain people, but they do not have to. A manufacturer may study all light bulbs produced by a machine during a week. An environmental scientist may study all daily air-quality readings in a city during a year. A streaming service may study all viewing sessions started by U.S. subscribers during a month. A botanist may study all trees of a species in a protected forest. Introducing Statistics uses “items or individuals” because the unit depends on the question.
Why use a sample?
A census attempts to collect data from every member of the population. A census can be appropriate when the population is small, every member can be reached, and the required measurements are practical. However, a census can be expensive, slow, or impossible. Some measurements destroy the item being tested. A company cannot measure the breaking strength of every helmet it intends to sell because the test would leave no helmets to sell. Some populations are continuously changing, such as all future customers or all future production units.
Sampling can provide timely information with fewer resources. Yet a sample introduces a central statistical issue: different possible samples can produce different results. That is sampling variability. A sample of 100 residents may report a different average commute time from another sample of 100 residents, even if both are selected carefully from the same city.
Target population and accessible population
The target population is the group the researcher ultimately wants to understand. The accessible population is the portion that can realistically be reached using the available list, place, period, or system. Suppose a university wants to learn about all graduates from the past ten years but only has current email addresses for 60 percent of them. The target population is all graduates in the defined period. The accessible population is the set with usable contact information. The difference can matter because reachable and unreachable graduates may differ.
Population and sample can change with the question
The same data source can support different populations if the investigative question changes. A data set of 500 restaurant orders collected on Fridays could be a sample from all Friday orders at one restaurant, or it might be the complete population of Friday orders during the specific data-collection month. Whether the records form a sample or a population depends on the scope of the question, not on the number of rows alone.
Individuals and observational units
An individual is an object, person, case, or other entity described by the data. The more general term observational unit is useful when the row is not literally a person. In a conventional rectangular data table, each row usually represents one observational unit and each column represents one variable. The word “usually” matters because repeated-measures and hierarchical data can use more complex structures.
The unit must match the question
Suppose a study records the number of emergency-department visits at 40 hospitals for each of 365 days. Is the observational unit a hospital, a day, or a hospital-day combination? The answer depends on the table. If each row is one hospital on one date, the observational unit is a hospital-day. There are up to 40 × 365 = 14,600 rows. If the data are first summarized so that each row contains a hospital’s annual average, the observational unit in the summarized data set is a hospital.
Misidentifying the unit can lead to a false sample size. A study with 50 students measured on 10 occasions has 500 measurements but only 50 student participants. Those measurements are not necessarily 500 independent individuals. Introducing Statistics does not yet develop dependence models, but it establishes the habit of asking what one row or record represents.
Student survey
Each individual may be one student. Variables might include grade level, travel time, and preferred learning format.
Battery test
Each observational unit may be one battery. Variables might include brand, production batch, and operating life.
Customer transaction
Each observational unit may be one purchase. Variables might include price, time, channel, and product category.
Individual, subject, participant, case, and record
Different fields use different words. A subject or participant is usually a person involved in a study. A case may be a person, organization, event, or object. A record is the stored representation of an observational unit. The terms overlap, but the safest practice is to define the unit explicitly: “Each row represents one household interviewed once” or “Each row represents one shipment.”
Units can be nested
Students belong to classrooms, classrooms belong to schools, and schools belong to districts. Employees belong to teams and companies. Measurements may occur within patients over time. At Topic 1.1, the main goal is recognition: a study may contain multiple levels, and the population or sample statement should specify the level to which the question refers. A question about average school size has schools as the units, even though student counts appear in the data.
Data and datasets
A datum is one recorded piece of information. Data is the plural form, although in everyday writing the word is often treated as a mass noun. A data set is an organized collection of data. Organization includes labels, units, definitions, and enough structure to connect values to individuals and variables.
Consider the row below from a study of school travel. The student identifier is ST104. The student’s travel mode is bus, travel time is 27 minutes, distance is 8.4 miles, and arrival status is on time. Each cell contains a datum. The complete table of many students and columns is the data set.
| Student ID | Travel mode | Travel time (min) | Distance (miles) | Arrival status |
|---|---|---|---|---|
| ST104 | Bus | 27 | 8.4 | On time |
Data values need variable definitions
The value 27 is not fully meaningful until the variable is defined. Does travel time begin when the student leaves home or when the bus arrives? Does it end at the school gate, classroom door, or official start time? Are waiting minutes included? Clear operational definitions make values comparable. Without them, two observers may record different numbers for the same journey.
A variable is a characteristic that can take different values across individuals or occasions. Topic 1.2 develops variable types in detail, but Topic 1.1 requires enough understanding to identify what was recorded. In the travel data, travel mode and arrival status are categorical variables; travel time and distance are quantitative variables. The important first step is to name the variable in context rather than simply pointing to a column.
A data set is not just numbers
Names, labels, categories, dates, yes/no responses, text codes, and missing-value indicators can all be part of a data set. Statistical analysis depends on the meanings attached to these entries. A code of 1 might mean “yes,” “male,” “treatment group,” or “first visit.” A data dictionary should explain the code. If a data set uses 999 to mean “not reported,” treating 999 as a genuine measurement could severely distort an average.
Rows
Rows generally identify observational units. A row might represent a student, a product, a household, a day, or a transaction.
Columns
Columns generally identify variables. Every column needs a name, definition, units when applicable, and valid coding rules.
Raw data and derived data
Raw data are values as initially recorded, although even “raw” data may already reflect measurement choices. Derived data are created from other variables. Age on an interview date can be derived from date of birth. Body mass index can be derived from height and weight. A pass/fail indicator can be derived from a score and a stated cutoff. Derived values can be useful, but the transformation should be documented.
Data quality begins before analysis
Introducing Statistics encourages basic questions about data quality: Are the intended individuals included? Are there duplicate rows? Are units consistent? Are impossible values present? Are missing values clearly marked? Were definitions applied in the same way? These questions are not advanced software tasks. They are part of understanding what the data set actually says.
Parameters and statistics
A parameter is a numerical characteristic of a population. A statistic is a numerical characteristic calculated from a sample. Parameters and statistics often describe the same kind of feature—such as a mean or proportion—but they refer to different groups.
| Feature | Population parameter | Sample statistic |
|---|---|---|
| Mean | Population mean, often written μ | Sample mean, often written x̄ |
| Proportion | Population proportion, often written p | Sample proportion, often written p̂ |
| Standard deviation | Population standard deviation, often written σ | Sample standard deviation, often written s |
| Size | Population size N | Sample size n |
Example: screen time
A school district wants to know the mean weekday recreational screen time of all 10th-grade students in the district. The population mean, μ, is the parameter. It is fixed for the defined population and measurement rules, although it may be unknown. The district surveys a sample of 180 students and obtains a sample mean of x̄ = 3.6 hours. The value 3.6 hours is a statistic.
The sample statistic is used as an estimate of the population parameter. Another properly selected sample might yield 3.4 or 3.8 hours. This does not mean the population parameter changes every time a sample is drawn. It means the statistic varies from sample to sample.
Parameter does not mean “important number”
Students sometimes label a value a parameter because it appears in the question or because it seems official. The correct distinction is the group described. If 62 percent of the 250 surveyed customers were satisfied, 62 percent is a sample statistic. If company records include every customer in the defined population and show that 62 percent were satisfied, the value is a population parameter for that population.
A statistic can describe a complete observed population
In informal language, any numerical summary may be called a statistic. In introductory inference, however, the parameter-statistic distinction is tied to population and sample. When a school calculates the mean score of every student enrolled in a particular class, that mean is a parameter for the class population. The same class may be treated as a sample from a broader conceptual population, such as future students taught under the same method, but that broader claim requires careful definition.
Estimand, estimate, and estimator
Three related terms deepen the distinction. An estimand is the precise population quantity the study aims to learn, such as the population mean difference in sleep duration between two groups. An estimator is the rule or formula used to estimate it, such as the difference between two sample means. An estimate is the numerical result obtained from the observed data. Topic 1.1 does not require advanced estimator theory, but the language shows why a question must specify the desired population feature.
Statistical variability
Variability means that observations are not all identical and that repeated data-collection processes can produce different results. Variation is the reason statistical questions require data. If every individual had exactly the same value and every measurement were perfectly repeatable, many statistical problems would collapse into simple facts.
Variation among individuals
Students differ in commute time, sleep duration, course grades, and preferred learning methods. Manufactured parts differ slightly in length or strength. Daily temperatures change. Customer orders vary in amount. A statistical distribution records not only a typical value but also the pattern of differences across individuals.
Variation in measurement
The same characteristic may be recorded differently because instruments have limited precision, respondents interpret questions differently, observers use slightly different judgments, or conditions change. If a person’s pulse is measured twice, the results may differ because the pulse itself changes and because the measurement process is imperfect.
Variation from sample to sample
Suppose 40 percent of all students at a school use the bus. One random sample of 50 students may contain 18 bus users, giving a sample proportion of 36 percent. Another sample may contain 23 bus users, giving 46 percent. Both statistics can arise from the same population. Later units quantify this sampling variability through sampling distributions and standard errors.
Variation does not automatically mean error
Natural differences are not mistakes. Two batteries from the same production process can have different lifetimes. Two students can respond differently to the same lesson. Statistical reasoning distinguishes natural variation from measurement error, selection bias, recording mistakes, and systematic changes. Introducing Statistics begins with recognition: whenever a question concerns a variable characteristic, expect a distribution of possible values rather than one universal answer.
What is Maria’s student ID?
For a specified person and current record, the answer is a single fact. Data collection may be needed to look it up, but the question does not ask about a variable across a group.
How do student commute times vary at the school?
Different students are expected to have different times. Answering requires observations from a group and a description of the distribution.
Variability shapes the wording of conclusions
A sample average does not imply that every individual has that value. If the mean commute time is 24 minutes, some students may travel 5 minutes and others 60 minutes. If 72 percent of a sample supports a policy, the remaining 28 percent does not. Good statistical writing preserves this variation instead of turning a summary into a claim about every case.
Investigative questions
An investigative question is a question that guides a statistical study. It identifies what the study seeks to learn and requires data from or about a group. A valid investigative question is answerable, connected to a population or process, and written so that the required variables can be identified.
The official Topic 1.1 skill is to determine a valid investigative question that requires a statistical investigation. That means recognizing more than punctuation. A sentence can end with a question mark and still fail as an investigative question. “Is exercise good?” is too vague. Good for whom, measured how, compared with what, and during what period? Introducing Statistics turns broad interests into questions that can guide evidence collection.
What a strong question reveals
- Population: the group or process about which information is wanted.
- Variable or variables: the characteristics that must be recorded.
- Purpose: description, comparison, association, prediction, or another clearly stated aim.
- Context: place, period, condition, or definition needed for a meaningful answer.
- Variability: an expectation that observations may differ.
- Feasibility: the needed data could reasonably be collected or obtained.
Five common families of investigative questions
What values occur?
What is the distribution of weekday sleep duration among 11th-grade students at Pine High School?
What is typical?
What is the median weekday wait time for Route 8 passengers during the morning rush?
How common?
What proportion of households in Lake County used curbside recycling last month?
How do groups differ?
How does average battery life compare between Brand A and Brand B earbuds tested under the same conditions?
How are variables related?
Is weekly practice time associated with free-throw percentage among players in the regional league?
How does a value change?
How did monthly library visits change from January through December 2026?
Descriptive and inferential versions
“What percentage of the 80 surveyed students bring lunch from home?” is descriptive because it asks about the observed sample. “What proportion of all students at the school bring lunch from home?” is inferential if only a sample is observed. The second question concerns a population parameter and requires a defensible sampling process if the result is to be generalized.
Questions create variable requirements
Consider the question, “Is weekly study time associated with mathematics exam score among first-year college students?” The study needs at least two variables for each student: weekly study time and exam score. It also needs a definition of first-year student, a time period for study hours, and a specified exam or score. If the question compares online and in-person sections, course format becomes another required variable.
Statistical versus nonstatistical questions
A statistical question anticipates variability and is answered by collecting, analyzing, or summarizing data. A nonstatistical question seeks a single fixed fact, definition, or calculation that does not require understanding variation across observations. The distinction depends on the intended question, not simply on whether a number appears in the answer.
| Question | Classification | Reason |
|---|---|---|
| How many seats are listed on Bus 24? | Nonstatistical | For the specified bus and current configuration, the question seeks one fixed count. |
| How many passengers ride Bus 24 on weekday mornings? | Statistical | Passenger counts vary by day, date, weather, and other conditions. |
| What is Jamal’s height today? | Usually nonstatistical | It asks for one measurement on one identified individual, although repeated measurement could introduce another question. |
| What is the distribution of heights among students in the class? | Statistical | Students have different heights, so data from a group must be summarized. |
| Does the school cafeteria open at 7:00 a.m.? | Nonstatistical | The official schedule provides a fixed answer for the defined day. |
| At what time do students usually arrive at the cafeteria? | Statistical | Arrival time varies among students and days. |
Examples and nonexamples
A question can become statistical through repeated observations
“What is the temperature at noon today?” seeks one measurement. “How does the noon temperature vary across the 30 days of April?” is statistical. “What is the battery life of this device?” might seek one test result, while “What distribution of battery life is produced when this model is tested repeatedly under a standard protocol?” expects variability.
A statistical question is not automatically a good question
“What do people think?” anticipates variation, but the population and subject are undefined. “Is there a relationship between everything?” is impossible to operationalize. A valid statistical question must be sufficiently specific to identify data, units, and scope.
How to write an investigative question
A strong investigative question can be built deliberately. Introducing Statistics uses the POP-V-C method: identify the POPulation, name the Variable or variables, and state the Comparison, connection, or contextual boundary.
Step 1: Start with a real purpose
Begin with the decision or understanding the study should support. A school may need to adjust bus schedules, a manufacturer may need to evaluate battery consistency, or a library may need to plan staffing. The purpose should not dictate the result, but it helps identify which information is relevant.
Step 2: Define the population
Replace broad words with an operational boundary. Instead of “teenagers,” write “students ages 14–18 enrolled at the three district high schools on September 15, 2026.” Instead of “customers,” write “customers who completed an online order from the company’s U.S. store during June 2027.”
Step 3: Name measurable variables
Replace abstract ideas with recordable characteristics. “Engagement” might become number of voluntary discussion posts, minutes active in the course platform, or a score from a defined survey. Different operational definitions answer different questions. A variable should include units or categories when necessary.
Step 4: State the intended statistical feature
Decide whether the question asks for a distribution, mean, median, proportion, difference, association, or trend. Topic 1.1 does not require choosing formulas, but a clear target helps the study collect the right data.
Step 5: Include context and time
Conditions can change. Commute time during the morning rush differs from late evening. Customer satisfaction may depend on the product version. A time period, location, or operating condition often belongs in the question.
Step 6: Check that variability is expected
If the answer is a fixed definition or a single known record, the question may not require statistical investigation. A valid statistical question should anticipate differences among individuals, occasions, or repeated samples.
Step 7: Remove loaded wording
“Why do irresponsible drivers speed?” assumes irresponsibility and that the behavior occurs. A more neutral question is, “What proportion of licensed drivers in the county report exceeding the posted limit by at least 10 miles per hour during the previous month?” Neutral wording does not guarantee unbiased measurement, but it avoids building a preferred conclusion into the question.
Question repair workshop
| Weak question | Problem | Improved investigative question |
|---|---|---|
| Do students like math? | Undefined students, undefined “like,” no time or measurement. | Among students enrolled in Algebra II at Hill High this semester, what is the distribution of ratings on a 1–5 mathematics-interest scale? |
| Is the bus late? | One bus? Which route? What counts as late? Over what period? | During October weekdays, what proportion of Route 12 arrivals at Central Station occur more than five minutes after the scheduled time? |
| Which class is better? | “Better” is not operationally defined. | How do final exam scores and course-satisfaction ratings compare between the online and in-person sections of Introductory Biology in spring 2027? |
| Does coffee help? | Population, dose, outcome, and comparison are missing. | Among adult volunteers who normally consume less than one cup of coffee per day, how does reaction time compare after a standardized caffeinated drink and a visually identical noncaffeinated drink? |
| Are phones distracting? | “Distracting” needs measurement and context. | Among 10th-grade students completing a 30-minute reading task, how does the number of comprehension questions answered correctly differ between a phone-visible condition and a phone-stored-away condition? |
A self-check before data collection
- Can I identify the population without guessing?
- Can I identify each observational unit?
- Can I name the variable or variables and how they are recorded?
- Does the question anticipate variability?
- Is the time, place, or condition clear enough?
- Could data realistically answer the question?
- Does the wording avoid assuming the conclusion?
- Will the planned statistic or graph connect directly to the question?
Complete worked cases
The following original cases show how Introducing Statistics concepts fit together. Each case begins with a purpose, states an investigative question, and maps the population, sample, individuals, variables, parameter, statistic, and expected sources of variability.
Morning commute to school
A district transportation office wants to review the morning schedule at East High. It selects 150 students from the current enrollment list and records the number of minutes from leaving home to entering the school building on one regular school day. The sample mean is 31.8 minutes, and the sample median is 27 minutes.
The mean and median describe the sample, not every student. The difference between 31.8 and 27 suggests that longer commutes may pull the mean upward, but a graph would be needed to study the distribution properly. The sample selection method determines whether the statistics can credibly represent the population.
Wireless-earbud battery life
A quality-control team selects 60 pairs of earbuds from the 8,400 pairs produced during one week. Each pair is fully charged and played continuously at a fixed volume until shutdown. The sample mean operating time is 7.42 hours.
The test conditions are part of the variable definition. The result does not automatically describe battery life at every volume, with every device, or under intermittent use. Introducing Statistics emphasizes that a parameter belongs to a defined population under a defined measurement process.
Library program participation
A city library system wants to estimate the proportion of registered teen members who attended at least one library program during the summer. From the membership database, 400 teen members are selected; 148 attended at least one program.
A contextual interpretation is: “In the sample, 37 percent of the 400 selected registered teen members attended at least one summer program.” Calling 37 percent the population proportion would be unjustified unless every member were included or a later inferential argument supported the estimate.
Comparing two garden fertilizers
A horticulture class grows 48 tomato plants of the same variety under similar greenhouse conditions. Twenty-four plants receive Fertilizer A and 24 receive Fertilizer B. After eight weeks, students record each plant’s increase in height.
This question is statistical because plant growth varies. It is comparative because it involves two treatment groups. Later lessons will determine what the design permits the class to conclude. Topic 1.1 focuses on identifying the components before analyzing the results.
Streaming-session duration
A service studies 10,000 viewing sessions selected from sessions started by U.S. subscribers during one month. Each record contains session duration, device type, content category, and whether the session ended voluntarily or because of an error.
The observational unit is a session because one subscriber can contribute multiple sessions. If the question concerned average monthly viewing per subscriber, the unit and data organization would need to change. Introducing Statistics always ties the unit to the question.
Original Introducing Statistics exercises
These exercises are original and designed specifically for this lesson. They emphasize identification before calculation. For each scenario, write answers in complete contextual phrases. Do not answer “population = everyone” unless the study truly concerns everyone; name the exact group.
Introducing Statistics Exercise 1
A county health office wants the mean number of hours adults in the county sleep on a typical weeknight. It selects 320 adults and obtains a sample mean of 6.7 hours. Identify the population, sample, individual, variable, parameter, statistic, and investigative question.
Introducing Statistics Exercise 2
A bakery produced 2,400 loaves on Monday. Quality staff randomly inspect 80 loaves and find that 6 have an underweight package. Identify the population, sample, observational unit, variable, parameter, statistic, and investigative question.
Introducing Statistics Exercise 3
A principal records the final mathematics grade of every one of the 214 students in the graduating class. Is the group of 214 a sample or the population for the question, ‘What is the mean final mathematics grade of this graduating class?’ Explain.
Introducing Statistics Exercise 4
A researcher downloads 5,000 product reviews from all 68,000 verified reviews posted during 2026. Each row represents one review. Identify N, n, the population, the sample, and the observational unit.
Introducing Statistics Exercise 5
A study of 75 households records household size, monthly electricity use, and whether the home has solar panels. Identify the individuals and the variables.
Introducing Statistics Exercise 6
A survey of 180 students reports that 117 support extending library hours. Identify the sample statistic and state the corresponding population parameter in words.
Introducing Statistics Exercise 7
A delivery company states that the mean delivery time for all packages delivered last month was 2.8 days. The figure was calculated from the company’s complete database of every package in that month. Is 2.8 days a parameter or statistic for that defined population?
Introducing Statistics Exercise 8
A class asks, ‘What is the height of the school flagpole?’ Is this statistical or nonstatistical? Explain.
Introducing Statistics Exercise 9
A class asks, ‘How do the heights of flagpoles at public high schools in the state vary?’ Is this statistical or nonstatistical? Explain.
Introducing Statistics Exercise 10
Improve this question: ‘Do people use the internet too much?’ Write a measurable investigative question with a defined population, variable, and time context.
Introducing Statistics Exercise 11
A wildlife team tags 90 turtles from a lake and records shell length in centimeters. Its goal is to estimate the mean shell length of all adult turtles in the lake. Identify population, sample, individuals, variable, parameter, and statistic if the sample mean is 24.6 cm.
Introducing Statistics Exercise 12
A spreadsheet has 40 patients and five blood-pressure readings per patient. It contains 200 measurement rows. For a question about patients’ average blood pressure, what is the participant sample size? Why is 200 not automatically the number of independent individuals?
Introducing Statistics Exercise 13
A school data set contains one row per course enrollment. A student enrolled in six courses appears in six rows. What is the observational unit: student or course enrollment? Explain.
Introducing Statistics Exercise 14
A manufacturer tests 50 light bulbs until failure. The bulbs come from 10,000 bulbs produced in one shift. The sample median lifetime is 1,140 hours. Identify the parameter that the sample median might estimate.
Introducing Statistics Exercise 15
Classify the question: ‘How many pages are in the printed school handbook?’ Statistical or nonstatistical?
Introducing Statistics Exercise 16
Classify the question: ‘How many pages do students read from the handbook during the first week?’ Statistical or nonstatistical?
Introducing Statistics Exercise 17
A city wants to learn about commuting among all employed residents, but its contact list includes only residents who registered a vehicle. Distinguish the target population from the accessible population.
Introducing Statistics Exercise 18
A data table uses 999 to represent a missing response for weekly exercise minutes. Explain why treating 999 as an actual number could damage the analysis.
Introducing Statistics Exercise 19
A café records one row per transaction, including purchase amount, payment method, time, and whether a discount was used. Identify the individuals and variables.
Introducing Statistics Exercise 20
A teacher asks, ‘Which teaching method is best?’ Rewrite the question as a valid comparative investigative question.
Introducing Statistics Exercise 21
A sample proportion changes from 0.42 in one random sample to 0.47 in another. Does this alone prove that the population proportion changed? Explain using sampling variability.
Introducing Statistics Exercise 22
A study question is, ‘What percentage of the 250 surveyed voters favor Proposal A?’ Is the requested value a parameter or statistic? Explain.
Introducing Statistics Exercise 23
A different question asks, ‘What percentage of all registered voters in the city favor Proposal A?’ If only 250 voters are surveyed, what population quantity is being targeted?
Introducing Statistics Exercise 24
A student writes, ‘The population is the 100 people surveyed because those are all the people in the data.’ Correct the statement.
Introducing Statistics Exercise 25
Write an investigative question about school lunch wait time that clearly identifies the population, variable, and time period.
Introducing Statistics Exercise 26
Write an investigative question comparing battery life for two phone models under a common test condition.
Introducing Statistics Exercise 27
Write an investigative question about an association between two variables among high school students.
Introducing Statistics Exercise 28
For the question ‘What is the distribution of weekly paid-work hours among 12th-grade students at Lake High this semester?’ identify the population, individuals, and variable.
Introducing Statistics Exercise 29
A study collects ratings from 1 to 5 but does not explain what 1 and 5 mean. What data-documentation problem is present?
Introducing Statistics Exercise 30
Create a complete statistical study map for this situation: a community center selects 200 members to estimate the proportion who used the fitness room at least once in March; 86 selected members did so.
Complete solutions
Open each solution after attempting the exercise. The wording demonstrates how to keep the answer in context. Equivalent answers may be correct when they define the same population, sample, unit, and variable clearly.
Introducing Statistics Solution 1
- Population: all adults in the county under the stated weeknight definition.
- Sample: the 320 selected adults.
- Individual: one selected adult.
- Variable: typical weeknight sleep duration in hours.
- Parameter: the population mean sleep duration for all county adults.
- Statistic: x̄ = 6.7 hours.
- Investigative question: What is the mean typical-weeknight sleep duration among adults in the county?
Introducing Statistics Solution 2
- Population: all 2,400 loaves produced Monday.
- Sample: the 80 inspected loaves.
- Observational unit: one loaf.
- Variable: whether the loaf has an underweight package, yes or no.
- Parameter: the true proportion of all Monday loaves that are underweight.
- Statistic: p̂ = 6/80 = 0.075, or 7.5 percent.
- Investigative question: What proportion of Monday’s loaves are underweight?
Introducing Statistics Solution 3
- For that exact question, the 214 students are the population because every member of the defined graduating class is included.
- The calculated class mean is therefore a population parameter for that graduating class.
- The same students could be treated as a sample only if the question were broadened to a larger conceptual population, which would require additional justification.
Introducing Statistics Solution 4
- N = 68,000 reviews.
- n = 5,000 reviews.
- Population: all 68,000 verified reviews posted during 2026.
- Sample: the 5,000 downloaded reviews.
- Observational unit: one verified product review.
Introducing Statistics Solution 5
- Individuals: the 75 households.
- Variables: household size, monthly electricity use, and solar-panel status.
- The data set should also document units for electricity use and the month or billing period represented.
Introducing Statistics Solution 6
- Sample statistic: p̂ = 117/180 = 0.65, or 65 percent.
- Population parameter: the true proportion of all students in the defined school population who support extending library hours.
Introducing Statistics Solution 7
- It is a parameter for the defined population because every delivered package in that month was included.
- If the company used the month as a sample of future operations, 2.8 could also function as a sample summary for a broader process, but that is a different question.
Introducing Statistics Solution 8
- Nonstatistical in its ordinary form.
- The question seeks one fixed measurement for one specified flagpole rather than a distribution across a group.
Introducing Statistics Solution 9
- Statistical.
- Different schools are expected to have flagpoles of different heights, so the answer requires data from many schools and a description of variability.
Introducing Statistics Solution 10
- One valid answer: Among students enrolled at Central High during fall 2027, what is the distribution of self-reported non-school internet use in hours per weekday?
- Other answers are acceptable if they define a population, measurable variable, and context without using the loaded phrase ‘too much.’
Introducing Statistics Solution 11
- Population: all adult turtles in the lake.
- Sample: the 90 tagged turtles.
- Individuals: individual adult turtles.
- Variable: shell length in centimeters.
- Parameter: the population mean shell length of all adult turtles in the lake.
- Statistic: x̄ = 24.6 cm.
Introducing Statistics Solution 12
- The participant sample size is 40 patients.
- There are 200 measurements, but the five readings from one patient are linked to the same person and may be more similar to one another than readings from different patients.
- Counting 200 as independent individuals would confuse measurements with participants.
Introducing Statistics Solution 13
- The observational unit in the stated table is a course enrollment because each row describes one student-course combination.
- A separate student-level table would use one row per student.
Introducing Statistics Solution 14
- The corresponding parameter is the population median lifetime of all 10,000 bulbs produced in that shift under the same test conditions.
Introducing Statistics Solution 15
- Nonstatistical. The printed handbook has a fixed number of pages for that edition.
Introducing Statistics Solution 16
- Statistical. Students are expected to read different numbers of pages, so the question concerns a distribution across students.
Introducing Statistics Solution 17
- Target population: all employed residents of the city.
- Accessible population: employed residents who are included on the vehicle-registration contact list.
- Residents without registered vehicles may be systematically different in commuting behavior, so the accessible group may not represent the target population.
Introducing Statistics Solution 18
- The value 999 is a missing-value code, not 999 minutes of exercise.
- Including it as a genuine measurement would create extreme false values and could greatly inflate the mean and spread.
- The code should be documented and treated as missing.
Introducing Statistics Solution 19
- Individuals or observational units: transactions.
- Variables: purchase amount, payment method, transaction time, and discount-use status.
- Customers are not the row units unless the data are reorganized to one row per customer.
Introducing Statistics Solution 20
- One valid answer: Among students enrolled in Algebra I at North High, how do end-of-unit test scores compare between classes using Method A and classes using Method B during the fall semester?
- The response variable and population are defined, and the comparison is explicit.
Introducing Statistics Solution 21
- No. Different random samples from the same population can produce different sample proportions.
- The change from 0.42 to 0.47 may be ordinary sampling variability.
- Evidence that the population changed would require a design and analysis that distinguish real change from sample-to-sample fluctuation.
Introducing Statistics Solution 22
- Statistic, because the percentage is calculated for the 250 surveyed voters—the observed sample.
Introducing Statistics Solution 23
- The targeted parameter is p, the true proportion of all registered voters in the city who favor Proposal A.
- The sample proportion from 250 voters would be used as an estimate of that parameter.
Introducing Statistics Solution 24
- The 100 surveyed people are the sample because they are the observed subset.
- The population is the larger group the study seeks to describe, such as all eligible residents, customers, or students defined by the question.
Introducing Statistics Solution 25
- One valid answer: What is the distribution of minutes from joining the lunch line to receiving food among students buying lunch at West High during regular school days in October?
Introducing Statistics Solution 26
- One valid answer: Under continuous video playback at 50 percent screen brightness, how does battery life in hours compare between new Phone Model A and Phone Model B devices?
Introducing Statistics Solution 27
- One valid answer: Among students in grades 9–12 at River High, is school-night sleep duration associated with first-period tardiness during the fall semester?
Introducing Statistics Solution 28
- Population: all 12th-grade students enrolled at Lake High during the semester.
- Individuals: individual 12th-grade students.
- Variable: weekly paid-work hours.
Introducing Statistics Solution 29
- The rating scale lacks a coding definition or data dictionary.
- Readers cannot know whether 1 means very dissatisfied or very satisfied, and they cannot interpret the values reliably.
Introducing Statistics Solution 30
- Investigative question: What proportion of all community-center members used the fitness room at least once in March?
- Population: all community-center members during the defined period.
- Sample: the 200 selected members.
- Individuals: members.
- Variable: whether each member used the fitness room at least once in March.
- Statistic: p̂ = 86/200 = 0.43, or 43 percent.
- Parameter: the true proportion of all center members who used the fitness room at least once in March.
- Variability: members differ in use, and another sample could produce a different proportion.
Common errors and precise corrections
Error 1: Calling the sample the population
A data set contains only the observed cases, so students often call those cases the population. Correct the error by returning to the question. If the study wants to describe all 8,000 district students but surveys 300, the 8,000 students form the population and the 300 surveyed students form the sample.
Error 2: Calling every reported number a parameter
A parameter describes the entire population. A sample mean, sample median, sample proportion, or sample standard deviation is a statistic. The size or authority of the organization reporting the number does not change this definition.
Error 3: Naming a topic instead of a variable
“Health” is a topic, not a sufficiently defined variable. Resting heart rate in beats per minute, number of sick days during the semester, or a score on a named health questionnaire are variables. Introducing Statistics expects a characteristic that can be recorded for each unit.
Error 4: Confusing individuals with variables
In a student survey, students are individuals; grade level and commute time are variables. In a table of schools, schools are individuals; enrollment and graduation rate are variables. A quick check is to ask: Does this label identify a row or a column?
Error 5: Treating repeated measurements as new individuals
Ten readings from one machine are ten observations, but they may belong to one machine unit. The sample size at the machine level is one. A study can analyze readings, machines, or machine-time combinations, but the unit must match the question and design.
Error 6: Writing a question that hides the denominator
“What percent preferred option A?” is incomplete unless the denominator is known. Was the percentage among all invited people, all respondents, all eligible participants, or all people who answered that item? A proportion always compares a count with a defined total.
Error 7: Using a summary as if it describes every individual
If the sample mean is 14.2, it does not follow that every person scored 14.2. If 60 percent answered yes, it does not follow that a typical person is “60 percent yes.” Summaries describe distributions or groups.
Error 8: Assuming more rows guarantee better evidence
A very large convenience sample can be systematically unrepresentative. Sample size influences random variability, but selection methods influence bias and generalizability. Topic 1.1 introduces the distinction; later topics develop sampling design.
Error 9: Forgetting the time frame
A population can change. “All employees” today may differ from all employees last year. A question about monthly spending needs a stated month or typical-month definition. Time is part of the context, not an optional detail.
Error 10: Writing a conclusion without the study noun
“The mean is 7.4” is incomplete. “The sample mean continuous operating time of the 60 tested earbud pairs was 7.42 hours under the standard test” is contextual. Introducing Statistics builds the habit of naming the variable, units, and sample.
Mini assessment: one story, many components
A regional transit authority operates 340 buses. It wants to estimate the proportion of buses that require an unscheduled repair during a typical 30-day operating period. Engineers select 70 buses at the beginning of April and follow each bus for 30 days. Fourteen selected buses require at least one unscheduled repair. The data set contains one row per selected bus and columns for route group, bus age, miles traveled, number of repairs, and whether at least one unscheduled repair occurred.
Questions
- State the investigative question.
- Identify the population and its size.
- Identify the sample and its size.
- Identify the observational units.
- List the variables named in the story.
- Identify the sample statistic most directly connected to the question.
- State the corresponding population parameter.
- Give two reasons the observed results may vary among buses.
- Explain why “14” is not the sample proportion.
- Write a one-sentence contextual description of the sample result.
Open the complete Introducing Statistics mini-assessment solution
- What proportion of the authority’s buses require at least one unscheduled repair during a 30-day operating period comparable to the study period?
- Population: all 340 buses operated by the authority; N = 340.
- Sample: the 70 selected buses; n = 70.
- Observational units: individual buses.
- Route group, bus age, miles traveled, number of repairs, and whether at least one unscheduled repair occurred.
- p̂ = 14/70 = 0.20, or 20 percent.
- p, the true proportion of all 340 buses that would require at least one unscheduled repair during the defined 30-day conditions.
- Possible sources include bus age, mileage, route conditions, maintenance history, driver use, component differences, and random failures.
- The number 14 is the count of selected buses with a repair. A proportion requires dividing that count by the sample total of 70.
- In the sample, 20 percent of the 70 selected buses required at least one unscheduled repair during the 30-day observation period.
Why this assessment matters
The story contains several numbers, but only some answer the investigative question. N = 340 and n = 70 describe group sizes. Fourteen is a count. The value 0.20 is the sample proportion. The unknown population proportion is the parameter. Introducing Statistics teaches how to assign each number a role rather than treating every number as interchangeable.
Introducing Statistics glossary
- Accessible population
- The part of the target population that can realistically be reached through the available list, location, system, or time period.
- Census
- An attempt to collect data from every member of a defined population.
- Context
- The real-world meaning that connects statistical terms and calculations to the people, objects, variables, units, place, and time in the study.
- Data
- Recorded pieces of information about observational units.
- Data set
- An organized collection of data with labels, definitions, units, and structure.
- Datum
- One recorded piece of information; the singular form of data.
- Descriptive statistics
- Methods used to organize, display, and summarize observed data.
- Estimate
- A numerical value calculated from sample data to approximate a population parameter.
- Estimator
- A rule or statistic used to estimate a parameter.
- Estimand
- The precise population quantity a study intends to estimate.
- Individual
- A person, object, case, or entity described by the data.
- Investigative question
- A question that guides a statistical study and requires data to answer.
- N
- A common symbol for population size.
- n
- A common symbol for sample size.
- Nonstatistical question
- A question seeking a fixed fact, definition, or single result without requiring analysis of variability across observations.
- Observational unit
- The entity represented by one record at the level relevant to the question.
- Parameter
- A numerical characteristic of a population.
- Population
- The complete group of items or individuals relevant to the investigative question.
- Sample
- The subset of the population from which data are obtained.
- Sampling variability
- The natural tendency for statistics calculated from different samples to differ.
- Statistic
- A numerical characteristic calculated from sample data.
- Statistical inference
- Using sample data to draw conclusions about a larger population while accounting for uncertainty.
- Statistical question
- A question that anticipates variability and requires data from a group or repeated process.
- Statistical study
- A study in which data are collected and analyzed to answer an investigative question about a population or process.
- Target population
- The full group the researcher ultimately wants to understand.
- Variable
- A characteristic recorded for each observational unit that can take different values.
- Variability
- Differences among observations, measurements, occasions, or possible samples.
The final concept map
Question
Defines the population and needed variables.
Sample
Supplies the observed individuals and data.
Statistic
Summarizes the sample and may estimate a parameter.
Context
Returns the result to the original population and purpose.
Every later AP Statistics topic adds tools to this map. Graphs reveal distributions. Probability models random behavior. Confidence intervals estimate parameters. Hypothesis tests evaluate claims. Regression describes relationships. Yet those methods remain meaningful only when the population, sample, unit, variables, and investigative question are clear.
Frequently asked questions
What is the simplest meaning of statistics?
Statistics is the practice of learning from data while considering variation, context, uncertainty, and how the data were collected.
What is the difference between statistics and a statistic?
Statistics is the broader discipline. A statistic is a numerical summary calculated from sample data, such as a sample mean or sample proportion.
Can a population be infinite?
Yes. A population can be a conceptual process with indefinitely many possible outcomes, such as all future items produced under stable conditions. The population must still be defined clearly.
Is a large sample always representative?
No. A large sample can still be biased if the selection process systematically excludes or overrepresents parts of the population.
Can a sample contain the whole population?
If every population member is observed, the study is a census for that defined population rather than a sample survey. The same data can still be viewed as a sample from a broader conceptual population if the question changes.
Why is N uppercase and n lowercase?
Uppercase N is a common convention for population size, while lowercase n is commonly used for sample size. The symbols help distinguish the whole group from the observed subset.
Is every row always one person?
No. A row may represent a person, object, school, transaction, day, visit, or combined unit such as a patient-visit. The observational unit must be identified from the data structure and question.
What is the difference between an individual and a variable?
An individual is the entity described by a row. A variable is a characteristic recorded for individuals and usually represented by a column.
What is the difference between data and a data set?
Data are recorded pieces of information. A data set is the organized collection of those data with labels, definitions, and structure.
Is a percentage always a statistic?
A percentage calculated from a sample is a statistic. A percentage calculated from every member of the defined population is a parameter for that population.
Does a parameter change from sample to sample?
No. The parameter belongs to the defined population. Sample statistics change from sample to sample and are used to estimate the parameter.
Why do different samples give different answers?
Samples contain different individuals or outcomes. Natural population variation causes sample summaries to differ; this is sampling variability.
What makes a question statistical?
It anticipates variability and requires data from a group, repeated observations, or a random process to answer.
Can a yes-or-no question be statistical?
Yes. For example, ‘What proportion of households recycled last week?’ uses a yes-or-no variable across many households and expects variation.
Is ‘What is the average?’ a complete investigative question?
No. It should specify the average of which variable, for which population, and under what time or conditions.
What is the role of context?
Context tells the reader what each number, variable, group, and unit represents. It prevents technically correct but meaningless statements.
What is the quickest way to distinguish parameter and statistic?
Ask whether the number describes the full population or the observed sample. Population means parameter; sample means statistic.
What happens if the sample is the wrong group?
The statistic may accurately describe the observed sample but fail to answer the question about the intended population.
Can text responses be data?
Yes. Text can be recorded as data and later coded or analyzed, provided the collection and interpretation rules are clear.
Why introduce variables in Topic 1.1 if Topic 1.2 covers variables?
Topic 1.1 needs the basic idea that a variable is a recorded characteristic. Topic 1.2 develops variable types and classifications more fully.
What should a complete identification answer include?
Use contextual phrases: name the exact population, sample, observational unit, variable with units or categories, parameter, statistic, and investigative question.
What is the most common mistake in Introducing Statistics?
A common mistake is naming the observed sample as the population simply because only the sample appears in the data table.
Can the same number be a parameter in one question and a statistic in another?
Yes. A mean calculated from every student in one class is a parameter for that class. The class could be treated as a sample from a broader population in a different question.
Why must an investigative question avoid loaded language?
Loaded wording builds an assumption or preferred answer into the question and can influence data collection or interpretation.
What comes after Introducing Statistics?
The next AP Statistics topic studies variables in greater depth, including how characteristics are classified and represented.
Introducing Statistics concept recap
Use this final bank as a rapid conceptual check. Each statement condenses one habit developed in the chapter.
- Introducing Statistics connects the population to the sample before interpreting a number.
- Introducing Statistics asks what one row represents before counting observations.
- Introducing Statistics distinguishes a population parameter from a sample statistic.
- Introducing Statistics treats variability as expected information rather than an inconvenience.
- Introducing Statistics requires every investigative question to identify measurable information.
- Introducing Statistics keeps units, labels, and time frames attached to the data.
- Introducing Statistics uses contextual language so a conclusion answers the original question.
- Introducing Statistics recognizes that a large data set can still come from a weak selection process.
- Introducing Statistics separates a fixed factual question from a question that expects a distribution.
- Introducing Statistics turns broad interests into populations, variables, and answerable comparisons.
- Introducing Statistics connects the population to the sample before interpreting a number.
- Introducing Statistics asks what one row represents before counting observations.
- Introducing Statistics distinguishes a population parameter from a sample statistic.
- Introducing Statistics treats variability as expected information rather than an inconvenience.
- Introducing Statistics requires every investigative question to identify measurable information.
- Introducing Statistics keeps units, labels, and time frames attached to the data.
- Introducing Statistics uses contextual language so a conclusion answers the original question.
- Introducing Statistics recognizes that a large data set can still come from a weak selection process.
- Introducing Statistics separates a fixed factual question from a question that expects a distribution.