This is a printable AP Statistics practice worksheet with 100 original questions: 70 multiple-choice questions and 30 multi-part long-response questions. Every question is grouped by topic and every answer includes the calculation, reasoning, conditions, and contextual conclusion needed to study the method rather than memorize a letter.
What this worksheet is
The historical 2019 search phrase identifies reader intent, not the origin of these questions. The questions below were written specifically for this worksheet. They use hypothetical school, manufacturing, health, polling, and service contexts; none reports real research findings or pretends that invented values are public data.
The worksheet emphasizes durable introductory-statistics reasoning: define the parameter, identify the design, check conditions, select a procedure, compute carefully, and interpret in context. Where the revised 2026-27 AP Statistics framework differs from the older course, the current framework controls the main practice set. A short historical comparison near the end explains removed topics so students do not confuse an old secure practice exam with the May 2027 design.
For a timed session, complete Questions 1-42 in 90 minutes to rehearse the revised multiple-choice pace. The remaining MCQs provide extra topic coverage. For free response, select four long-response questions that match the revised roles: study design, data analysis, inference, and a mixed investigation. Then use the other solutions as targeted review.
100-question blueprint
| Topic | MCQs | Long response | Total |
|---|---|---|---|
| Data, descriptive statistics, study design | 10 | 4 | 14 |
| Distributions and normal models | 10 | 4 | 14 |
| Probability and random variables | 10 | 4 | 14 |
| Sampling distributions | 10 | 4 | 14 |
| Confidence intervals | 10 | 4 | 14 |
| Hypothesis tests | 10 | 4 | 14 |
| Chi-square procedures | 5 | 2 | 7 |
| Regression and correlation | 5 | 2 | 7 |
| Mixed investigations | 0 | 2 | 2 |
| Total | 70 | 30 | 100 |
Formula coverage
The formula index shows where each major calculation is used. The goal is not to substitute formulas blindly: identify the parameter and sampling design before computing.
| Formula or principle | Mathematical form | Where it is practiced |
|---|---|---|
| Mean | MCQ 1; FRQ 71 | |
| Standardized score | MCQ 11, 17-19; FRQ 75-76 | |
| Addition rule | MCQ 22 | |
| Conditional probability | MCQ 23; FRQ 79 | |
| Binomial probability | MCQ 25; FRQ 81 | |
| Expected value | MCQ 27; FRQ 80 and 82 | |
| Sample-mean standard error | MCQ 32, 35, 38; FRQ 78, 84, 86 | |
| Sample-proportion standard error | MCQ 31, 33; FRQ 83 | |
| One-proportion interval | FRQ 87 | |
| One-mean t interval | MCQ 46; FRQ 88 | |
| One-proportion z test | MCQ 55; FRQ 91 | |
| One-mean t test | MCQ 56; FRQ 92 | |
| Two-proportion z test | MCQ 58; FRQ 93 | |
| Chi-square expected count | MCQ 62; FRQ 95 | |
| Chi-square contribution | MCQ 63; FRQ 96 | |
| Regression | MCQ 66-67; FRQ 97-98 |
Questions 1-70: multiple choice
Directions: Select the best of four answers. Calculators may be used. Do not open the solutions until the set is complete.
Data, Descriptive Statistics, and Study Design
Question 1 Easy
A sample of five commute times is 12, 15, 18, 20, and 25 minutes. What is the sample mean?
- A. 17
- B. 18
- C. 19
- D. 20
Question 2 Easy
The ordered data are 3, 4, 7, 8, 9, 12, 16. Which value is the median?
- A. 7
- B. 8
- C. 9
- D. 12
Question 3 Easy
For the ordered data 2, 5, 7, 9, 12, 14, 18, 21, what is the interquartile range using the median-of-halves convention?
- A. 8
- B. 9
- C. 10
- D. 11
Question 4 Easy
Every temperature in a data set measured in degrees Celsius is converted to Fahrenheit by F=1.8C+32. If the Celsius standard deviation is 4, what is the Fahrenheit standard deviation?
- A. 4
- B. 7.2
- C. 36
- D. 39.2
Question 5 Easy
A school numbers all 1,240 students and uses a random-number generator to select 80 different numbers. Which sampling method is used?
- A. Convenience sample
- B. Simple random sample
- C. Stratified random sample
- D. Cluster sample
Question 6 Easy
In an experiment comparing two review apps, students are randomly assigned to an app. What is the main purpose of random assignment?
- A. To make the sample represent the population
- B. To eliminate all measurement error
- C. To balance lurking variables across treatment groups in the long run
- D. To guarantee equal sample means
Question 7 Easy
A researcher separates voters into four age groups and takes a simple random sample from each group. Which design is this?
- A. Stratified random sample
- B. Cluster sample
- C. Systematic sample
- D. Matched-pairs experiment
Question 8 Tough
A study finds that students who carry water bottles earn higher test scores. Which conclusion is justified from this observational association?
- A. Water bottles cause higher scores
- B. Removing water bottles will lower scores
- C. There is an association, but confounding prevents a causal conclusion
- D. The association proves the groups were randomly assigned
Question 9 Tough
For a strongly right-skewed income distribution with several very high values, which pair best describes center and spread?
- A. Mean and standard deviation
- B. Median and IQR
- C. Mean and range
- D. Mode and variance
Question 10 Toughest
A matched-pairs experiment compares two keyboard layouts. Which pairing is most defensible?
- A. Pair the fastest typist with the slowest typist
- B. Let participants choose a layout
- C. Have each participant use both layouts in randomized order
- D. Assign the first half to layout A and the second half to layout B
Distributions, Normal Models, and Standardized Scores
Question 11 Easy
A test score is 78 in a distribution with mean 70 and standard deviation 4. What is the z-score?
- A. 1
- B. 2
- C. 4
- D. 8
Question 12 Easy
A normal distribution has mean 50 and standard deviation 6. Approximately what proportion lies between 44 and 56?
- A. 0.50
- B. 0.68
- C. 0.95
- D. 0.997
Question 13 Easy
The 90th percentile of a normal distribution with mean 100 and standard deviation 15 is closest to which value?
- A. 109.6
- B. 115.0
- C. 119.2
- D. 124.7
Question 14 Easy
If X is normal with mean 40 and standard deviation 5, what is P(X<45), to the nearest hundredth?
- A. 0.16
- B. 0.50
- C. 0.84
- D. 0.98
Question 15 Easy
A density curve is entirely above the horizontal axis and has total area 1. What does an area of 0.27 over an interval represent?
- A. The interval width
- B. The median
- C. The proportion of observations in the interval
- D. The standard deviation
Question 16 Easy
A distribution has mean 30 and standard deviation 5. Each value is transformed by Y=10-2X. What are the mean and standard deviation of Y?
- A. Mean -50, SD 10
- B. Mean -50, SD -10
- C. Mean 70, SD 10
- D. Mean -40, SD 5
Question 17 Tough
A normal probability plot is close to a straight line except for two points far above the line at the upper end. What is the best interpretation?
- A. The data are perfectly normal
- B. The data may contain unusually large observations
- C. The median must equal zero
- D. The sample size is necessarily too small
Question 18 Tough
Two students have z-scores 1.4 and -0.6 on different exams. Which statement is correct?
- A. The first student scored 2 raw points higher
- B. The first student is 2 standard deviations higher relative to the respective distributions
- C. The second student has the larger percentile
- D. Raw scores can be compared directly from z-scores
Question 19 Tough
For a standard normal variable Z, P(Z>2.10) is closest to which value?
- A. 0.0179
- B. 0.0358
- C. 0.4821
- D. 0.9821
Question 20 Toughest
A variable is strongly left-skewed. Samples of size 4 are repeatedly drawn and their means recorded. Which claim is safest?
- A. The sampling distribution is exactly normal
- B. The sampling distribution has the same left skew as individual observations
- C. A normal model is not guaranteed because n=4 is small and the population is strongly skewed
- D. The sampling distribution has standard deviation four times the population SD
Probability and Random Variables
Question 21 Easy
If P(A)=0.42, what is P(not A)?
- A. 0.42
- B. 0.48
- C. 0.58
- D. 1.42
Question 22 Easy
Suppose P(A)=0.55, P(B)=0.40, and P(A and B)=0.20. What is P(A or B)?
- A. 0.35
- B. 0.60
- C. 0.75
- D. 0.95
Question 23 Easy
Among 80 students, 32 play an instrument and 12 of those 32 also sing. What is P(sings | plays an instrument)?
- A. 0.15
- B. 0.375
- C. 0.40
- D. 0.55
Question 24 Easy
Events A and B have probabilities 0.30 and 0.50. If they are independent, what is P(A and B)?
- A. 0.15
- B. 0.20
- C. 0.50
- D. 0.80
Question 25 Easy
For X~Binomial(8,0.25), what is P(X=3), to the nearest thousandth?
- A. 0.086
- B. 0.208
- C. 0.311
- D. 0.422
Question 26 Easy
For X~Binomial(40,0.30), what are the mean and standard deviation?
- A. 12 and 2.90
- B. 12 and 8.40
- C. 28 and 2.90
- D. 28 and 8.40
Question 27 Tough
A game pays $8 with probability 0.20 and loses $3 otherwise. What is the expected payoff per play?
- A. -$0.80
- B. $0.80
- C. $1.60
- D. $5.00
Question 28 Tough
A random variable X has standard deviation 3. For Y=5X-7, what is SD(Y)?
- A. 3
- B. 8
- C. 15
- D. 22
Question 29 Tough
A device fails independently with probability 0.10 on each trial. What is the probability of at least two failures in six trials?
- A. 0.114
- B. 0.354
- C. 0.469
- D. 0.886
Question 30 Toughest
A simulation uses two random digits per trial. The event has probability 0.23. Which assignment is valid?
- A. 00-22 represent success
- B. 00-23 represent success
- C. 01-23 represent success and 00 is ignored
- D. 00-22 represent failure
Sampling Distributions
Question 31 Easy
For random samples of size 100 from a population with proportion p=0.36, what is the mean of p-hat?
- A. 0.0036
- B. 0.036
- C. 0.36
- D. 36
Question 32 Easy
A population has standard deviation 18. What is the standard deviation of x-bar for random samples of size 36?
- A. 0.5
- B. 3
- C. 6
- D. 18
Question 33 Easy
For p=0.40 and n=150, what is SD(p-hat), to three decimals?
- A. 0.016
- B. 0.040
- C. 0.063
- D. 0.240
Question 34 Tough
Which condition supports an approximately normal sampling distribution for p-hat?
- A. np>=10 and n(1-p)>=10
- B. n>=30 only
- C. p=0.5 exactly
- D. The population must be normal
Question 35 Tough
If sample size increases from 64 to 256 with the population unchanged, how does SD(x-bar) change?
- A. It doubles
- B. It is halved
- C. It is quartered
- D. It is unchanged
Question 36 Tough
Which statistic is unbiased for the population mean mu?
- A. The sample range
- B. The sample maximum
- C. The sample mean
- D. The square of the sample mean
Question 37 Tough
Independent samples estimate p1 and p2. Which expression is the standard deviation of p-hat1-p-hat2?
- A. sqrt[p1(1-p1)/n1 + p2(1-p2)/n2]
- B. sqrt[(p1-p2)(1-p1+p2)/(n1+n2)]
- C. p1/n1-p2/n2
- D. sqrt[p1p2/(n1n2)]
Question 38 Tough
A population is normal with mean 72 and SD 12. For n=16, which describes x-bar?
- A. Exactly normal with mean 72 and SD 3
- B. Approximately normal with mean 72 and SD 12
- C. Exactly normal with mean 4.5 and SD 3
- D. Unknown shape because n<30
Question 39 Toughest
A sample of 80 is drawn without replacement from a population of 500. Which concern is most relevant for treating observations as approximately independent?
- A. The sample exceeds 10 percent of the population
- B. The population is too large
- C. The sample mean is biased
- D. The central limit theorem requires n=100
Question 40 Toughest
A highly skewed population has finite mean and SD. Which change most directly improves the normal approximation for x-bar?
- A. Increase sample size
- B. Decrease sample size
- C. Replace x-bar with the sample maximum
- D. Use convenience sampling
Confidence Intervals
Question 41 Easy
A 95% confidence interval for a population proportion is (0.42,0.50). What is its margin of error?
- A. 0.02
- B. 0.04
- C. 0.08
- D. 0.46
Question 42 Easy
Which is a correct interpretation of a 95% confidence interval for a population mean?
- A. 95% of observations lie in the interval
- B. There is a 95% chance the fixed mean changes into the interval
- C. The method captures the true mean in about 95% of repeated samples
- D. The sample mean has a 95% chance of equaling the population mean
Question 43 Tough
Holding confidence level and estimated proportion fixed, what happens to margin of error when sample size is multiplied by 4?
- A. It doubles
- B. It is halved
- C. It is quartered
- D. It stays the same
Question 44 Tough
With no prior estimate of p, what minimum sample size is needed for a 95% margin of error at most 0.03?
- A. 534
- B. 752
- C. 1,068
- D. 2,135
Question 45 Tough
Which interval is appropriate for estimating the mean change in blood pressure measured on the same patients before and after treatment?
- A. One-sample t interval on paired differences
- B. Two-sample z interval for proportions
- C. Chi-square interval
- D. One-proportion z interval
Question 46 Tough
A sample mean is based on n=25 with unknown population SD. Which critical distribution is normally used for a mean interval?
- A. Standard normal z
- B. t with 24 degrees of freedom
- C. Chi-square with 25 degrees of freedom
- D. Binomial with n=25
Question 47 Toughest
A 90% interval and a 99% interval are computed from the same sample and method. Which is wider?
- A. The 90% interval
- B. The 99% interval
- C. They have equal width
- D. Width cannot be compared
Question 48 Toughest
For a one-proportion z interval, which counts are checked using sample data?
- A. n p0 and n(1-p0)
- B. n p-hat and n(1-p-hat)
- C. Population size only
- D. Degrees of freedom and expected counts
Question 49 Toughest
Two independent groups produce a confidence interval for p1-p2 of (-0.03,0.11). Which conclusion is supported?
- A. p1 is definitely larger
- B. p2 is definitely larger
- C. A zero difference remains plausible at this confidence level
- D. Both proportions equal zero
Question 50 Toughest
A poll's interval accounts only for random sampling error. Which issue is not repaired by increasing sample size?
- A. Large standard error
- B. Nonresponse bias
- C. A wide interval from small n
- D. Sampling variability
Hypothesis Tests
Question 51 Easy
A claim says a population proportion exceeds 0.60. Which alternative hypothesis matches the claim?
- A. p=0.60
- B. p<0.60
- C. p>0.60
- D. p not equal to 0
Question 52 Easy
A test gives p-value 0.018 at alpha=0.05. What is the decision?
- A. Fail to reject H0
- B. Reject H0
- C. Accept H0 as true
- D. Increase alpha after seeing the data
Question 53 Tough
Which describes a Type I error when testing H0:p=0.40 against Ha:p>0.40?
- A. Concluding p>0.40 when p=0.40
- B. Failing to conclude p>0.40 when p=0.55
- C. Concluding p=0.40 when p=0.40
- D. Using a sample that is too large
Question 54 Tough
What does a p-value measure?
- A. The probability H0 is true
- B. The probability of results at least as extreme as observed, assuming H0 is true
- C. The probability the study is unbiased
- D. The long-run confidence level
Question 55 Tough
In a sample of 200, p-hat=0.56. Testing H0:p=0.50, what is the one-proportion z statistic?
- A. 0.85
- B. 1.20
- C. 1.70
- D. 2.40
Question 56 Tough
A one-sample t test has t=2.31 and df=19. Which statement is correct?
- A. The population SD must be known
- B. The statistic measures the sample mean's standardized distance from the null mean using s
- C. The p-value must equal 2.31
- D. The sample size is 19
Question 57 Toughest
Which change generally increases power for detecting a fixed effect?
- A. Decrease sample size
- B. Use a smaller alpha while holding all else fixed
- C. Increase sample size
- D. Add measurement noise
Question 58 Toughest
For a two-proportion z test of H0:p1=p2, which estimate is used in the null standard error?
- A. Separate sample proportions only
- B. The pooled proportion
- C. The midpoint of the confidence interval
- D. The population mean
Question 59 Toughest
A study measures reaction time before and after caffeine for each participant. Which test matches the design?
- A. Paired t test on differences
- B. Two independent-sample t test
- C. One-proportion z test
- D. Chi-square test of independence
Question 60 Toughest
A statistically significant result has a very small estimated effect in a sample of 50,000. What is the best interpretation?
- A. The effect must be practically important
- B. Statistical significance and practical importance must be evaluated separately
- C. The null has probability zero
- D. The data prove causation
Chi-Square Procedures
Question 61 Easy
A 3-by-4 two-way table is analyzed with a chi-square test of independence. What are the degrees of freedom?
- A. 5
- B. 6
- C. 7
- D. 12
Question 62 Tough
A table has row total 80, column total 45, and grand total 200. What is the expected count for their cell under independence?
- A. 9
- B. 18
- C. 22.5
- D. 36
Question 63 Tough
A cell has observed count 26 and expected count 20. What is its chi-square contribution?
- A. 0.30
- B. 1.20
- C. 1.80
- D. 6.00
Question 64 Toughest
A chi-square test of independence yields p=0.004. Which conclusion is appropriate?
- A. The categorical variables are associated in the population
- B. The variables have a strong linear correlation
- C. Every cell differs from expectation
- D. One variable causes the other
Question 65 Toughest
Which condition problem most directly undermines a chi-square approximation?
- A. Expected counts are mostly below 5
- B. The sample size is even
- C. Rows and columns have labels
- D. Observed counts are integers
Regression and Correlation
Question 66 Easy
A least-squares line is y-hat=12+3.5x. What is the predicted change in y for a one-unit increase in x?
- A. 3.5 decrease
- B. 3.5 increase
- C. 12 increase
- D. 15.5 increase
Question 67 Tough
For y-hat=5+2x, an observation has x=4 and y=15. What is its residual?
- A. -2
- B. 2
- C. 8
- D. 13
Question 68 Tough
A regression has r^2=0.64. Which interpretation is correct?
- A. 64% of observations lie on the line
- B. 64% of variability in the response is explained by the linear model with the predictor
- C. The slope is 0.64
- D. The correlation is necessarily -0.64
Question 69 Toughest
A point has an extreme x-value and lies close to the existing regression trend. What effect can it have?
- A. It can be influential even with a small residual
- B. It can never affect slope
- C. It must make correlation zero
- D. It proves the relationship is causal
Question 70 Toughest
A residual plot shows a clear U-shaped pattern. What is the best conclusion?
- A. A linear model is appropriate
- B. The response has no variability
- C. A nonlinear relationship remains in the residuals
- D. The correlation must be exactly 1
Questions 71-100: long response
Directions: Show method, conditions, calculations, and a conclusion in context. Blank space is included for printing; add paper when a response needs more room.
Data, Descriptive Statistics, and Study Design
Question 71 Easy
The daily numbers of customers served at a help desk over eight days were 42, 48, 51, 51, 54, 56, 60, and 78. (a) Compute the mean and median. (b) Compute the IQR using the median-of-halves convention. (c) Identify any 1.5-IQR outliers and recommend summaries of center and spread.
Question 72 Tough
A district wants to estimate the proportion of its 9,600 students who have reliable home internet. The roster records grade level. Design a stratified random sample of 480 students, explain why grade is a useful stratum, and state the population and parameter.
Question 73 Toughest
A teacher compares a retrieval-practice app with ordinary rereading. Sixty volunteers are available, and prior achievement strongly predicts the outcome. Design a randomized block experiment and explain what causal conclusion would be justified.
Question 74 Toughest
A news article reports that neighborhoods with more public libraries have higher literacy rates and concludes that building a library will raise every resident's literacy. Evaluate the conclusion and propose a stronger study.
Distributions, Normal Models, and Standardized Scores
Question 75 Easy
Package weights are approximately normal with mean 500 g and standard deviation 8 g. Find (a) P(X<492), (b) P(492<X<508), and (c) the 95th percentile.
Question 76 Tough
Two placement exams use different scales. Lina earns 82 on Exam A (mean 70, SD 8) and 610 on Exam B (mean 550, SD 50). Compare her relative standing and explain a limitation.
Question 77 Toughest
A variable X has mean 14 and SD 3. Define Y=40-2.5X. Find the mean and SD of Y, and explain why the sign of the multiplier does not make the SD negative.
Question 78 Toughest
A right-skewed population has mean 120 and SD 30. Describe the sampling distribution of x-bar for n=100 and estimate P(x-bar>126). State the assumptions.
Probability and Random Variables
Question 79 Easy
A screening rule flags 8% of records. Among flagged records, 30% contain an actual error. Among unflagged records, 2% contain an error. Draw a probability tree and find the overall probability that a record contains an error.
Question 80 Tough
A fair spinner has outcomes 1, 2, 3, 4 with probabilities 0.10, 0.20, 0.30, 0.40. Find the mean and standard deviation of X.
Question 81 Toughest
A component passes inspection independently with probability 0.92. In a batch of 20, find the probability exactly 18 pass, the expected number that pass, and the standard deviation.
Question 82 Toughest
A game costs $4. It pays $20 with probability 0.10, $6 with probability 0.25, and $0 otherwise. Define net gain, compute its expected value, and decide whether the game is favorable to the player in the long run.
Sampling Distributions
Question 83 Easy
In a population, 35% support a proposal. For an SRS of 200, describe the sampling distribution of p-hat and find P(p-hat>0.40).
Question 84 Tough
A population has mean 64 and SD 20. Compare the standard errors of x-bar for n=25 and n=100, and explain the precision change.
Question 85 Toughest
Independent samples have p1=0.60, n1=150 and p2=0.45, n2=200. Describe the sampling distribution of p-hat1-p-hat2 and approximate P(p-hat1-p-hat2<0.08).
Question 86 Toughest
Explain why a sample mean can be unbiased but still have large variability. Use a numerical example with population SD 50 and compare n=4 with n=100.
Confidence Intervals
Question 87 Easy
In an SRS, 138 of 240 students prefer a later start. Construct a 95% confidence interval for the population proportion and interpret it.
Question 88 Tough
A random sample of 36 battery lives has mean 41.2 hours and SD 4.8 hours. Construct a 95% t interval using t*=2.030 and interpret it.
Question 89 Toughest
Two independent random samples give 84 of 140 users succeeding with interface A and 99 of 180 with interface B. Construct a 95% interval for pA-pB.
Question 90 Toughest
Determine the minimum sample size for estimating a population proportion with 99% confidence and margin of error at most 0.025 when a pilot estimate is 0.38. Use z*=2.576.
Hypothesis Tests
Question 91 Easy
A company claims more than 70% of orders arrive within two days. In a random sample of 180 orders, 137 do. Test at alpha=0.05.
Question 92 Tough
A machine should fill bags to mean 500 g. A random sample of 25 bags has mean 496.8 g and SD 7.5 g. Test H0:mu=500 against Ha:mu<500 using t=-2.133 and report the conclusion if p=0.021.
Question 93 Toughest
In independent samples, treatment succeeds for 72 of 120 patients and control succeeds for 54 of 120. Test whether the treatment success proportion is higher at alpha=0.05.
Question 94 Toughest
Explain Type I error, Type II error, and power for a medical alarm test of H0: the mean contaminant level is safe versus Ha: it exceeds the safe threshold. Discuss the effect of lowering alpha.
Chi-Square Procedures
Question 95 Tough
A survey cross-tabulates preferred study method (video, textbook, group) by grade level (9, 10, 11, 12). State hypotheses, degrees of freedom, expected-count calculation, and conclusion language for p=0.032.
Question 96 Toughest
For a 2-by-3 table, one cell has O=42 and E=30, while another has O=18 and E=24. Compute both chi-square contributions and explain how residual direction and contribution size differ.
Regression and Correlation
Question 97 Tough
A regression predicting quiz score from study hours is y-hat=58+4.2x with r^2=0.49. Interpret the slope, intercept, and r-squared, and predict for x=6.
Question 98 Toughest
A fitted line changes from y-hat=20+1.8x to y-hat=23+1.1x when one high-x observation is removed. Its residual under the original fit is small. Explain leverage, residual, and influence.
Mixed Investigation
Question 99 Toughest
Design an investigation of whether a new review schedule improves AP Statistics performance. Include the question, design, response, analysis, and scope of inference.
Question 100 Toughest
A school reports a 95% interval of (2.1,6.7) points for the mean paired score improvement after a workshop and a two-sided p-value of 0.001. Reconcile the interval and test, discuss practical importance, and identify one design threat.
Worked solutions for Questions 1-70
Each solution displays the controlling formula or principle and explains the answer. A correct numerical value without the appropriate model or condition is not a complete statistical argument.
Data, Descriptive Statistics, and Study Design
Question 1: B (Easy)
Add the five values to obtain 90 and divide by 5. The sample mean is 18 minutes.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 2: B (Easy)
There are seven observations, so the median is the fourth ordered value. That value is 8.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 3: C (Easy)
The lower-half median is (5+7)/2=6 and the upper-half median is (14+18)/2=16. Therefore IQR=16-6=10. The correct choice is 10.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 4: B (Easy)
Adding 32 changes location but not spread. Multiplying every observation by 1.8 multiplies the standard deviation by 1.8, giving 7.2.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 5: B (Easy)
Every set of 80 students has the same chance to be selected when 80 distinct labels are chosen uniformly from the complete roster. This is a simple random sample.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 6: C (Easy)
Random assignment creates comparable treatment groups by distributing known and unknown confounders across groups in the long run. It does not make a convenience sample representative.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 7: A (Easy)
Age groups are strata. Sampling independently inside every stratum is stratified random sampling.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 8: C (Tough)
An observational study can establish association, not causation. Study habits, course load, or other variables could affect both bottle use and scores.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 9: B (Tough)
The median and IQR resist extreme high incomes, while the mean and standard deviation are pulled upward. Resistant summaries are preferred for a skewed distribution with outliers.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 10: C (Toughest)
Using each person for both layouts controls person-to-person typing ability. Randomizing order limits practice and fatigue effects.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Distributions, Normal Models, and Standardized Scores
Question 11: B (Easy)
Standardize by subtracting the mean and dividing by the standard deviation: (78-70)/4=2.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 12: B (Easy)
The interval is one standard deviation below to one above the mean. The empirical rule gives approximately 68 percent.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 13: C (Easy)
The standard-normal 90th percentile is z=1.282. Transform back: 100+(1.282)(15)=119.2.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 14: C (Easy)
The standardized value is z=(45-40)/5=1. The normal cumulative probability is 0.8413, which rounds to 0.84.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 15: C (Easy)
For a continuous distribution, probability is represented by area under the density curve. An area of 0.27 corresponds to a proportion of 0.27.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 16: A (Easy)
The transformed mean is 10-2(30)=-50. Standard deviation is multiplied by the absolute slope, so it is 2(5)=10.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 17: B (Tough)
Departures at the upper end indicate observations larger than a normal model would predict. The plot suggests high-end outliers or a heavy right tail.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 18: B (Tough)
Their relative positions differ by 1.4-(-0.6)=2 standard deviations. The raw-score difference cannot be recovered without means and standard deviations.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 19: A (Tough)
The left-tail probability at 2.10 is 0.9821; subtract from 1 to obtain the right tail 0.0179.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 20: C (Toughest)
The central limit theorem does not guarantee approximate normality for a very small sample from a strongly skewed population. The standard deviation of the sample mean is sigma/sqrt(n), not n times sigma.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Probability and Random Variables
Question 21: C (Easy)
An event and its complement have probabilities that sum to 1, so 1-0.42=0.58.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 22: C (Easy)
Use the general addition rule: 0.55+0.40-0.20=0.75.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 23: B (Easy)
Restrict the denominator to the 32 instrument players. The conditional proportion is 12/32=0.375.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 24: A (Easy)
For independent events, multiply their probabilities: 0.30(0.50)=0.15.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 25: B (Easy)
Use the binomial formula: C(8,3)(0.25)^3(0.75)^5=0.2076, which rounds to 0.208.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 26: A (Easy)
The mean is np=12. The standard deviation is sqrt(np(1-p))=sqrt(8.4)=2.90.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 27: A (Tough)
Compute the probability-weighted average: 8(0.20)+(-3)(0.80)=1.60-2.40=-0.80 dollars.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 28: C (Tough)
Subtracting 7 does not affect spread. Multiplying by 5 multiplies standard deviation by |5|, giving 15.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 29: A (Tough)
Use the complement: 1-P(0)-P(1)=1-(0.9)^6-6(0.1)(0.9)^5=0.1143.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 30: A (Toughest)
There are 100 equally likely pairs from 00 through 99. Assigning 23 pairs, 00 through 22, to success represents probability 23/100.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Sampling Distributions
Question 31: C (Easy)
The sample proportion is an unbiased estimator of p, so the sampling-distribution mean equals 0.36.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 32: B (Easy)
The standard error of the sample mean is 18/sqrt(36)=3.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 33: B (Easy)
Compute sqrt[p(1-p)/n]=sqrt[(0.4)(0.6)/150]=0.0400, which rounds to 0.040.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 34: A (Tough)
At least 10 expected successes and 10 expected failures is the standard large-counts check for a sample proportion.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 35: B (Tough)
Standard error is proportional to 1/sqrt(n). Multiplying n by 4 divides the standard error by 2.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 36: C (Tough)
Across repeated random samples, the mean of x-bar equals the population mean. Therefore x-bar is unbiased for mu.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 37: A (Tough)
Variances add for independent sample proportions, so add the two variance terms and take the square root.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 38: A (Tough)
A mean from a normal population is exactly normal for any sample size. Its mean is 72 and its SD is 12/sqrt(16)=3.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 39: A (Toughest)
Because 80/500=0.16, the sample is more than 10 percent of the population. The usual 10 percent condition is not satisfied.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 40: A (Toughest)
The central limit theorem gives a better normal approximation for the sample mean as n increases, assuming independent observations and no pathological infinite-variance behavior.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Confidence Intervals
Question 41: B (Easy)
The margin of error is half the interval width: (0.50-0.42)/2=0.04.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 42: C (Easy)
Confidence describes the long-run success rate of the interval-producing method. The population mean is fixed; intervals vary across samples.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 43: B (Tough)
Standard error contains 1/sqrt(n). Multiplying n by 4 divides standard error and margin of error by 2.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 44: C (Tough)
Use p*=0.5 for the conservative maximum variance. n=0.25(1.96/0.03)^2=1067.11; round up to 1068.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 45: A (Tough)
Measurements are paired within patients. Compute each patient's difference and use a one-sample t interval for the mean difference.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 46: B (Tough)
When sigma is unknown, standardize with the sample SD and use a t distribution with n-1=24 degrees of freedom.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 47: B (Toughest)
Higher confidence requires a larger critical value, so the 99% interval is wider when the estimate and standard error are unchanged.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 48: B (Toughest)
An interval has no null value p0, so use the observed successes n p-hat and failures n(1-p-hat) for the large-counts check.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 49: C (Toughest)
Because the interval includes 0, no difference is a plausible population value. The interval does not prove equality, but it does not establish a directional difference.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 50: B (Toughest)
A larger sample reduces sampling variability but does not automatically remove systematic nonresponse bias. Design and follow-up procedures address nonresponse.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Hypothesis Tests
Question 51: C (Easy)
The word exceeds specifies a right-tailed alternative: p>0.60. Equality belongs in the null hypothesis.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 52: B (Easy)
Because 0.018<0.05, reject the null hypothesis. The conclusion should then be stated in context as evidence for the alternative.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 53: A (Tough)
A Type I error rejects a true null. Here it means claiming the proportion exceeds 0.40 when the true proportion is 0.40.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 54: B (Tough)
A p-value is computed under the null model. It is the probability of the observed statistic or one more extreme in the alternative's direction, assuming H0.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 55: C (Tough)
Under H0, SE=sqrt[(0.5)(0.5)/200]=0.0354. Thus z=(0.56-0.50)/SE=1.70.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 56: B (Tough)
The t statistic uses the sample SD: (x-bar-mu0)/(s/sqrt(n)). With df=n-1=19, the sample size is 20.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 57: C (Toughest)
A larger sample reduces standard error, making a fixed effect easier to distinguish from the null. Smaller alpha and more noise generally reduce power.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 58: B (Toughest)
Under the null, both groups share a common proportion. Pool successes and sample sizes to estimate that common value.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 59: A (Toughest)
The same participant provides both measurements, so observations are paired. Test the mean of participant-level differences.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 60: B (Toughest)
Very large samples can detect effects too small to matter in practice. Report the estimate and uncertainty, then judge practical importance in context.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Chi-Square Procedures
Question 61: B (Easy)
Degrees of freedom are (rows-1)(columns-1)=(3-1)(4-1)=6.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 62: B (Tough)
Multiply the row and column totals and divide by the grand total: 80(45)/200=18.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 63: C (Tough)
The contribution is (26-20)^2/20=36/20=1.8.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 64: A (Toughest)
The small p-value is evidence against independence, so the categorical variables are associated. The test alone does not establish causation or linear correlation.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 65: A (Toughest)
The chi-square reference distribution requires sufficiently large expected counts. Many expected counts below 5 can make the approximation unreliable.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Regression and Correlation
Question 66: B (Easy)
The slope is 3.5, so predicted y increases by 3.5 units for each one-unit increase in x.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 67: B (Tough)
The predicted value is 5+2(4)=13. Residual=observed-predicted=15-13=2.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 68: B (Tough)
The coefficient of determination is the proportion of response variation explained by the least-squares linear relationship with x.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 69: A (Toughest)
A high-leverage point can substantially affect the fitted slope even when its residual relative to the fitted line is small. Influence is assessed by refitting without the point.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Question 70: C (Toughest)
A systematic curve in residuals shows that the linear model misses structure. A nonlinear model or transformed relationship may be more appropriate, although departures-from-linearity analysis is not a named revised-course topic.
Why the alternatives fail: The other choices represent a common arithmetic, condition, interpretation, or method-selection error; compare each one with the displayed formula and the defined parameter.
Worked solutions for Questions 71-100
Each long-response solution includes a compact four-point guide. These are instructional rubrics for this original worksheet, not official College Board scoring guidelines.
Data, Descriptive Statistics, and Study Design
Question 71: worked solution (Easy)
The sum is 440, so the mean is 440/8=55. The median is the average of the fourth and fifth observations: (51+54)/2=52.5. The lower-half median is (48+51)/2=49.5 and the upper-half median is (56+60)/2=58, so IQR=8.5. The upper fence is 58+1.5(8.5)=70.75; 78 is an outlier. The lower fence is 49.5-12.75=36.75, so there is no low outlier. Because the distribution has a high outlier and right skew, report the median 52.5 and IQR 8.5 as the primary resistant summaries.
Four-point scoring guide
- 1 point: correct mean and median
- 1 point: correct quartiles and IQR
- 1 point: correct fences and outlier
- 1 point: justified median/IQR recommendation
Question 72: worked solution (Tough)
The population is all 9,600 district students. The parameter is p, the true proportion of those students with reliable home internet. Divide the roster into grade-level strata. Allocate the 480 selections proportionally: for each grade, multiply 480 by that grade's enrollment divided by 9,600, then randomly select the resulting number from that grade roster without replacement. Grade is useful if internet access or response patterns differ by age while students within a grade are comparatively similar. Sampling every grade can reduce variability and ensures representation. Follow up consistently across strata and record nonresponse, because random selection alone does not eliminate nonresponse bias.
Four-point scoring guide
- 1 point: population and parameter
- 1 point: correct stratification
- 1 point: random sampling within every stratum
- 1 point: valid precision/representation explanation
Question 73: worked solution (Toughest)
Block the 60 volunteers by a defensible prior-achievement measure, such as forming pairs of students with similar pretest scores or forming narrow achievement bands. Within each pair or block, randomly assign equal numbers to the app and rereading. Keep study time, content, testing conditions, and instructor contact constant. The response variable should be defined before the experiment, such as score on a common delayed assessment. Compare treatment means overall while respecting the block structure. Random assignment permits a cause-and-effect conclusion for participants under the implemented conditions. Because the participants volunteered rather than being randomly sampled, generalization to all students requires caution.
Four-point scoring guide
- 1 point: valid blocks based on prior achievement
- 1 point: random assignment within blocks
- 1 point: controlled conditions and defined response
- 1 point: correct causation/generalization distinction
Question 74: worked solution (Toughest)
The reported data establish a neighborhood-level association, not an individual causal effect. Income, school funding, urban density, educational attainment, and civic investment could confound the relationship. The ecological design also risks attributing a group pattern to individuals. A stronger observational study would prespecify the response, sample comparable neighborhoods, measure major confounders, and use a matched or adjusted analysis while still avoiding a causal claim. If feasible and ethical, a randomized intervention could assign comparable communities to an enhanced library-access program or usual services, then compare changes in a validated literacy outcome. Random assignment supports causation; random sampling or broad site selection supports generalization.
Four-point scoring guide
- 1 point: association is not causation
- 1 point: plausible confounders/ecological issue
- 1 point: stronger observational design
- 1 point: randomized intervention and scope of inference
Distributions, Normal Models, and Standardized Scores
Question 75: worked solution (Easy)
For 492 g, z=(492-500)/8=-1, so P(X<492)=Phi(-1)=0.1587. The interval 492 to 508 corresponds to -1<Z<1, so the probability is 0.6827. The 95th-percentile standard-normal value is about 1.645; transform back to obtain 500+1.645(8)=513.16 g, or about 513.2 g. The numerical answers rely on the stated normal model.
Four-point scoring guide
- 1 point: first z-score and probability
- 1 point: middle probability
- 1 point: correct percentile z
- 1 point: back-transformed weight with units
Question 76: worked solution (Tough)
Exam A gives z_A=(82-70)/8=1.50. Exam B gives z_B=(610-550)/50=1.20. Relative to the respective distributions, Lina's Exam A result is 0.30 standard deviations farther above its mean. If both distributions are reasonably modeled and percentile comparisons are appropriate, Exam A is the stronger relative performance. A z-score comparison does not establish that the exams measure identical skills, have equal reliability, or have normal distributions. Raw-point differences across scales are meaningless.
Four-point scoring guide
- 1 point: correct z_A
- 1 point: correct z_B
- 1 point: correct relative comparison
- 1 point: substantive limitation
Question 77: worked solution (Toughest)
Linearity of expectation gives E(Y)=40-2.5E(X)=40-35=5. For spread, SD(Y)=|-2.5|SD(X)=2.5(3)=7.5. Standard deviation measures a nonnegative magnitude. The negative multiplier reverses the ordering of observations around the mean but scales every distance from the mean by 2.5; distances do not retain a negative sign. Variance would be (-2.5)^2(3^2)=56.25, and its square root is 7.5.
Four-point scoring guide
- 1 point: transformed mean
- 1 point: transformed SD
- 1 point: correct absolute-value reasoning
- 1 point: variance cross-check
Question 78: worked solution (Toughest)
For independent random observations, E(x-bar)=120 and SD(x-bar)=30/sqrt(100)=3. Although the population is right-skewed, n=100 is large enough for a central-limit approximation unless the distribution is extraordinarily heavy-tailed. Standardize 126: z=(126-120)/3=2. Therefore P(x-bar>126) is approximately P(Z>2)=0.0228. The calculation assumes random sampling or a design that supports independence, the 10 percent condition when sampling without replacement from a finite population, a finite population variance, and a sufficiently accurate normal approximation.
Four-point scoring guide
- 1 point: center and standard error
- 1 point: CLT justification
- 1 point: z-score and probability
- 1 point: assumptions/conditions
Probability and Random Variables
Question 79: worked solution (Easy)
The first branches are flagged with probability 0.08 and unflagged with probability 0.92. Conditional error probabilities are 0.30 and 0.02. The joint probability of flagged and error is 0.08(0.30)=0.024. The joint probability of unflagged and error is 0.92(0.02)=0.0184. Add the disjoint paths: P(error)=0.024+0.0184=0.0424. Thus about 4.24 percent of records contain an error under the model.
Four-point scoring guide
- 1 point: correct branches
- 1 point: flagged-error joint probability
- 1 point: unflagged-error joint probability
- 1 point: correct total with interpretation
Question 80: worked solution (Tough)
The mean is E(X)=1(0.10)+2(0.20)+3(0.30)+4(0.40)=3.00. Next E(X^2)=1(0.10)+4(0.20)+9(0.30)+16(0.40)=10.00. Therefore Var(X)=E(X^2)-[E(X)]^2=10-9=1, and SD(X)=1. The probabilities sum to 1, so the distribution is valid.
Four-point scoring guide
- 1 point: valid probability check
- 1 point: correct expected value
- 1 point: correct second moment/variance
- 1 point: correct standard deviation
Question 81: worked solution (Toughest)
Let X be the number that pass. The fixed number of trials, two outcomes, independence, and constant success probability justify X~Binomial(20,0.92). P(X=18)=C(20,18)(0.92)^18(0.08)^2, approximately 0.270. The mean is np=18.4. The standard deviation is sqrt[np(1-p)]=sqrt[20(0.92)(0.08)]=sqrt(1.472)=1.213. Independence is an explicit modeling assumption; if components share production conditions, it should be checked.
Four-point scoring guide
- 1 point: binomial model justification
- 1 point: correct exact probability setup/value
- 1 point: correct mean
- 1 point: correct SD and contextual caveat
Question 82: worked solution (Toughest)
Net gains after the $4 cost are $16, $2, and -$4 with probabilities 0.10, 0.25, and 0.65. The expected net gain is 16(0.10)+2(0.25)-4(0.65)=1.60+0.50-2.60=-$0.50 per play. In a long sequence of plays under a stable independent model, the player's average net gain should approach a loss of about 50 cents per play. A negative expected value does not mean every play loses; it describes the long-run average.
Four-point scoring guide
- 1 point: correct net-gain distribution
- 1 point: correct expected-value expression
- 1 point: correct -$0.50 result
- 1 point: long-run interpretation
Sampling Distributions
Question 83: worked solution (Easy)
The center is p=0.35. The standard deviation is sqrt[(0.35)(0.65)/200]=0.03373. The large-counts values are np=70 and n(1-p)=130, both at least 10, so a normal approximation is appropriate. Standardize 0.40: z=(0.40-0.35)/0.03373=1.482. Thus P(p-hat>0.40) is approximately 0.069. Random sampling and the 10 percent condition support independence.
Four-point scoring guide
- 1 point: center and standard error
- 1 point: normal-condition check
- 1 point: z calculation
- 1 point: tail probability and independence condition
Question 84: worked solution (Tough)
For n=25, SE=20/sqrt(25)=4. For n=100, SE=20/sqrt(100)=2. Quadrupling sample size halves the standard error. The sampling distribution for n=100 is therefore more concentrated around the same population mean 64. This is a precision statement, not a claim that every larger sample estimate will be closer to 64.
Four-point scoring guide
- 1 point: SE for n=25
- 1 point: SE for n=100
- 1 point: correct ratio
- 1 point: careful precision interpretation
Question 85: worked solution (Toughest)
The mean difference is 0.60-0.45=0.15. The standard deviation is sqrt[(0.60)(0.40)/150+(0.45)(0.55)/200]=sqrt(0.0028375)=0.05327. All four expected success/failure counts exceed 10, so a normal approximation is appropriate. For 0.08, z=(0.08-0.15)/0.05327=-1.314. The left-tail probability is approximately 0.094. The two samples and observations within each sample must be independent.
Four-point scoring guide
- 1 point: mean difference
- 1 point: standard error
- 1 point: normal checks
- 1 point: z and probability
Question 86: worked solution (Toughest)
Unbiasedness concerns the center across repeated samples: E(x-bar)=mu. Variability concerns how far individual sample means tend to fall from that center. With population SD 50, n=4 gives SE=50/sqrt(4)=25, while n=100 gives SE=50/sqrt(100)=5. Both sampling distributions are centered at mu, so both estimators are unbiased, but the n=4 estimator is much less precise. Bias and variability are separate dimensions of estimator quality.
Four-point scoring guide
- 1 point: defines unbiasedness
- 1 point: SE for n=4
- 1 point: SE for n=100
- 1 point: explains bias-versus-variability distinction
Confidence Intervals
Question 87: worked solution (Easy)
The estimate is p-hat=138/240=0.575. Successes 138 and failures 102 exceed 10; random sampling and the 10 percent condition should also hold. SE=sqrt[(0.575)(0.425)/240]=0.03191. With z*=1.96, ME=0.06254. The interval is (0.512,0.638), rounded to three decimals. We are 95% confident that the true proportion of students in the sampled population who prefer a later start lies between about 51.2% and 63.8%.
Four-point scoring guide
- 1 point: estimate and conditions
- 1 point: standard error/margin
- 1 point: correct interval
- 1 point: contextual interpretation
Question 88: worked solution (Tough)
Use a one-sample t interval because the population SD is unknown. SE=4.8/sqrt(36)=0.8. The margin of error is 2.030(0.8)=1.624. The interval is 41.2+/-1.624=(39.576,42.824), or about (39.6,42.8) hours. This method requires random/independent observations and a population distribution without severe skew or outliers; n=36 also gives some robustness. We are 95% confident the population mean battery life lies in this interval.
Four-point scoring guide
- 1 point: method/conditions
- 1 point: SE and margin
- 1 point: interval
- 1 point: contextual interpretation
Question 89: worked solution (Toughest)
The estimates are 0.600 and 0.550, so the observed difference is 0.050. Each sample has at least 10 successes and failures. The unpooled interval SE is sqrt[(0.6)(0.4)/140+(0.55)(0.45)/180]=0.05558. The 95% margin is 1.96(0.05558)=0.10894. The interval is (-0.059,0.159). Because 0 is included, the data do not establish a population difference at the corresponding two-sided 5% level. This is not proof that the interfaces are identical.
Four-point scoring guide
- 1 point: estimates/conditions
- 1 point: unpooled SE
- 1 point: interval
- 1 point: conclusion about zero
Question 90: worked solution (Toughest)
Use n=p*(1-p*)(z*/ME)^2 with p*=0.38. This gives n=(0.38)(0.62)(2.576/0.025)^2=2502.0 approximately; using unrounded arithmetic yields about 2502.05, so round up to 2,503. Rounding up is required because rounding down can produce a margin of error larger than requested. If the pilot estimate were not credible, using p*=0.50 would be the conservative choice and would require a larger sample.
Four-point scoring guide
- 1 point: correct formula
- 1 point: substitution
- 1 point: correct upward rounding
- 1 point: conservative-estimate explanation
Hypothesis Tests
Question 91: worked solution (Easy)
Set H0:p=0.70 and Ha:p>0.70, where p is the true on-time proportion. The sample estimate is 137/180=0.7611. Under H0, np0=126 and n(1-p0)=54, so the normal condition holds. The standard error is sqrt[(0.70)(0.30)/180]=0.03416, giving z=(0.7611-0.70)/0.03416=1.789. The right-tail p-value is about 0.0368. Because 0.0368<0.05, reject H0; the sample provides evidence that the true proportion exceeds 0.70, assuming the sampling conditions hold.
Four-point scoring guide
- 1 point: hypotheses/parameter
- 1 point: conditions
- 1 point: statistic and p-value
- 1 point: contextual decision
Question 92: worked solution (Tough)
The parameter mu is the true mean fill weight. A one-sample t test is appropriate if the sample is random/independent and the fill-weight distribution has no severe skew or outliers. The statistic can be verified: (496.8-500)/(7.5/sqrt(25))=-3.2/1.5=-2.133 with df=24. The supplied left-tail p-value is 0.021. At alpha=0.05, reject H0 and conclude there is statistically significant evidence that the machine's mean fill is below 500 g. The result does not state the probability that H0 is true.
Four-point scoring guide
- 1 point: parameter/hypotheses
- 1 point: method/conditions
- 1 point: t calculation/p-value
- 1 point: contextual conclusion
Question 93: worked solution (Toughest)
Let pT and pC be the population success proportions. Test H0:pT=pC against Ha:pT>pC. The estimates are 0.60 and 0.45, difference 0.15. The pooled estimate is (72+54)/(240)=0.525. The null SE is sqrt[(0.525)(0.475)(1/120+1/120)]=0.06447. Thus z=0.15/0.06447=2.327. The right-tail p-value is about 0.010. Large counts and independent randomization/sampling must hold. Reject H0; there is evidence that treatment has a higher success proportion. Causation requires random treatment assignment.
Four-point scoring guide
- 1 point: hypotheses
- 1 point: pooled SE/conditions
- 1 point: z and p-value
- 1 point: conclusion and scope
Question 94: worked solution (Toughest)
A Type I error occurs when the system triggers an alarm even though the true mean is at the safe threshold; this may cause unnecessary shutdown or investigation. A Type II error occurs when the system fails to trigger even though the true mean exceeds the safe threshold; this may expose people to unsafe conditions. Power is the probability that the test triggers for a specified truly unsafe mean. Lowering alpha reduces the probability of a Type I error, but with sample size and effect fixed it generally increases Type II error and reduces power. The choice of alpha should reflect the relative consequences, not be changed after seeing data.
Four-point scoring guide
- 1 point: contextual Type I
- 1 point: contextual Type II
- 1 point: power definition
- 1 point: correct alpha tradeoff
Chi-Square Procedures
Question 95: worked solution (Tough)
H0 states that preferred study method and grade level are independent in the population; Ha states they are associated. The table has 3 rows and 4 columns, so df=(3-1)(4-1)=6. For any cell, expected count=(its row total)(its column total)/(grand total). Check that expected counts are sufficiently large and that observations come from a random/independent design. At alpha=0.05, p=0.032 leads to rejection of H0. Conclude that the data provide evidence of an association between grade and preferred study method. Do not claim that grade causes preference.
Four-point scoring guide
- 1 point: hypotheses
- 1 point: df
- 1 point: expected-count rule/conditions
- 1 point: contextual conclusion
Question 96: worked solution (Toughest)
The first contribution is (42-30)^2/30=144/30=4.8. Its residual O-E is +12, so the category is overrepresented relative to independence. The second contribution is (18-24)^2/24=36/24=1.5. Its residual is -6, so it is underrepresented. Chi-square contributions are nonnegative because residuals are squared; direction must be recovered from O-E or a standardized residual. The first cell contributes more to the overall statistic. A full conclusion still depends on every cell, the degrees of freedom, and the p-value.
Four-point scoring guide
- 1 point: first contribution
- 1 point: second contribution
- 1 point: correct residual directions
- 1 point: contribution-versus-direction explanation
Regression and Correlation
Question 97: worked solution (Tough)
The slope says each additional study hour is associated with a predicted increase of 4.2 quiz points, on average, within the range of observed hours. The intercept predicts 58 points at zero study hours; it is meaningful only if zero lies in the data range and the linear model remains sensible there. The r-squared value means 49% of the sample variability in quiz scores is explained by the linear regression on study hours. At x=6, y-hat=58+4.2(6)=83.2 points. None of these statements alone establishes causation.
Four-point scoring guide
- 1 point: slope
- 1 point: intercept caveat
- 1 point: r-squared
- 1 point: prediction and causation caveat
Question 98: worked solution (Toughest)
The observation has high leverage because its x-value is far from the center of the predictor values. Its small residual means it lies close vertically to the original fitted line; small residual does not imply small influence. Removing it changes the slope from 1.8 to 1.1 and the intercept from 20 to 23, a substantial change, so the point is influential. The correct investigation is to verify the observation, report analyses with and without it, and explain why conclusions differ. A valid observation should not be deleted merely because it affects the fit.
Four-point scoring guide
- 1 point: leverage
- 1 point: residual
- 1 point: influence from refitting
- 1 point: responsible analysis
Mixed Investigation
Question 99: worked solution (Toughest)
Ask whether the review schedule causes a higher mean score on a prespecified assessment than the existing schedule for eligible students. Recruit students under a documented rule, block by prior performance or teacher, and randomly assign within blocks to new or existing review. Keep content exposure and total study time comparable. Use a common assessment scored without knowledge of assignment. Analyze the treatment difference with a blocked or two-sample method appropriate to the design, report an interval and effect size, and inspect assumptions. Random assignment supports causation for participants; representative sampling or diverse sites are needed for broad generalization. Record attrition and analyze whether it differs by treatment.
Four-point scoring guide
- 1 point: clear question/response
- 1 point: randomized blocked design
- 1 point: appropriate analysis
- 1 point: scope and attrition
Question 100: worked solution (Toughest)
The interval excludes zero and contains only positive mean improvements, so it agrees with rejecting H0:mu_d=0 in a two-sided test; p=0.001 indicates that a mean difference at least this incompatible with zero would be rare under the null model. The estimated mean improvement is the interval midpoint, 4.4 points, with margin of error 2.3 points. Statistical significance does not decide whether a 2.1-to-6.7 point improvement is educationally important; that requires a subject-matter benchmark and consideration of costs. Without a randomized comparison group, maturation, concurrent instruction, or regression to the mean could explain some improvement.
Four-point scoring guide
- 1 point: interval-test agreement
- 1 point: estimate and margin
- 1 point: practical-importance distinction
- 1 point: credible design threat
How to score the worksheet
Score each MCQ as one point and each long-response problem as four points, for a maximum of 190 practice points. This total is a diagnostic score only; it is not an AP score conversion and should not be presented as one. College Board sets AP score standards through a formal process, and a classroom worksheet cannot predict a future cut score.
| Diagnostic area | Evidence of strength | Recommended next action |
|---|---|---|
| Method selection | The named procedure matches the parameter and design | Redo missed questions by writing the parameter before the formula |
| Conditions | Randomness, independence, distribution, and count checks are stated when relevant | Build a condition checklist for intervals and tests |
| Computation | Substitution, standard error, statistic, and probability agree | Recalculate without looking at the answer; then compare line by line |
| Interpretation | The conclusion names the population, parameter, direction, and uncertainty | Rewrite bare numerical answers as complete contextual sentences |
| Design | Sampling and assignment are distinguished | Mark whether each conclusion supports generalization, causation, both, or neither |
For a balanced second attempt, choose one easy, one moderate, one tough, and one toughest problem from every major topic. Delay the second attempt by at least a day so the score measures retrieval and reasoning rather than immediate recognition.
2019 versus the revised 2027 scope
The 2019 framework organized content differently and the older exam used 40 multiple-choice questions with five answer choices and six free-response questions. The revised May 2027 exam uses 42 multiple-choice questions with four choices and four 10-point free-response questions, all completed in Bluebook.
| Historical topic or feature | Status for the revised 2026-27 course | How this worksheet handles it |
|---|---|---|
| Geometric distribution | Removed from the revised required framework | Not included in the 100-question core |
| Combining random variables as a named topic | Removed | Only basic linear transformations and expected value are used |
| Chi-square goodness-of-fit | Removed | The set uses independence/association ideas instead |
| Inference for a regression slope | Removed | Regression questions address interpretation, residuals, r-squared, and influence |
| Paper free-response booklet | Replaced by fully digital responses in May 2027 | Prompts emphasize concise typed notation and contextual sentences |
Students who need authentic released AP questions should use College Board’s public past-question page. Authorized teachers can use AP Classroom for secure older materials. This worksheet intentionally protects that boundary while still providing substantial practice.
Official sources
| Official source | What it verifies |
|---|---|
| AP Statistics revisions | Verifies the five-unit revision, removed topics, 42 MCQs, four FRQs, and fully digital May 2027 exam |
| AP Statistics assessment | Verifies the 90-minute sections, 50/50 weighting, and four FRQ roles |
| Released AP Statistics questions | Explains public access to recent released questions and secure access through AP Classroom |
| Course and Exam Description effective Fall 2026 | Defines the revised content and statistical practices |