12 complete FRQ sets + 55 topic-by-topic MCQ/FRQ practice banks
Keep the existing 12-set, 48-FRQ library and five-unit AP Statistics navigation while adding 55 searchable practice microtopics, MCQ routes, full tests, formula review, and progress tools.
55 AP Statistics practice microtopics with MCQ + FRQ links
Search by concept, filter by practice track, mark topics practiced on this device, or jump to a random visible topic.
Categorical Versus Quantitative Variable Identification
Distinguish categorical labels from numerical measurements, including numeric-looking category codes.
Category Labels and Qualitative Data Values
Recognize category names, qualitative values, and when arithmetic on labels is meaningless.
Frequency Counts for Categories
Turn raw categorical observations into accurate counts and check the total frequency.
Relative Frequency as a Category Proportion
Convert category counts to proportions and interpret a relative frequency in context.
Percentage Frequency Conversion
Move correctly between proportions, decimals, and percentage frequencies.
One-Way Frequency Tables
Read and construct one-way frequency tables without omitting categories or totals.
One-Way Relative-Frequency Tables
Build relative-frequency tables and verify that proportions sum to approximately 1.
Distribution of a Categorical Variable
Describe the distribution of one categorical variable using dominant, rare, and comparative categories.
Frequency Bar Charts
Read and create frequency bar charts with correct categories, heights, labels, and scale.
Relative-Frequency Bar Charts
Use relative-frequency bar charts to compare category proportions rather than raw sample sizes.
Categorical Dotplot Displays
Interpret categorical dot-style displays and distinguish their purpose from numerical dotplots.
Pie-Chart Part-to-Whole Displays
Read part-to-whole relationships from pie charts and connect sector size to percentage.
Appropriate Use of Pie Charts
Decide when a pie chart is justified because categories partition one meaningful whole.
Inappropriate Pie-Chart Situations When Categories Are Not Parts of One Whole
Reject misleading pie-chart use when categories overlap or do not form parts of one whole.
Subset Proportions After Excluding Categories
Recalculate proportions after removing categories and clearly identify the new denominator.
Side-by-Side Bar Charts for Categorical Comparisons
Compare categorical distributions with side-by-side bars while separating count from proportion.
Graph Labeling and Scaling for Categorical Displays
Detect labeling, scaling, baseline, and ordering problems that can distort categorical graphs.
Definition of a Quantitative Variable
Define a quantitative variable and explain why numerical operations have real meaning.
Discrete Quantitative Variables
Identify discrete quantitative variables that arise from counts or separated possible values.
Continuous Quantitative Variables
Identify continuous variables measured on a continuum and discuss practical measurement precision.
Counted Quantities Versus Measured Quantities
Separate counted quantities from measured quantities and justify the classification in context.
Quantitative Frequency Tables
Organize numerical values into a quantitative frequency table and interpret repeated values or classes.
Quantitative Relative-Frequency Tables
Translate quantitative frequencies into relative frequencies and compare distributions fairly.
Dotplots for Numerical Data
Construct and interpret numerical dotplots while retaining each individual observation.
Histogram Construction Concepts
Choose bins, boundaries, and widths for a histogram without creating artificial visual patterns.
Histogram Interval or Bin Interpretation
Determine which observations belong in histogram intervals and handle boundary conventions correctly.
Frequency Histograms
Read frequency histograms and connect bar height or area to observation counts.
Relative-Frequency Histograms
Interpret relative-frequency histograms and compare groups with different sample sizes.
Relative Area as Relative Frequency in a Histogram
Use histogram area as relative frequency when bin widths and vertical scaling make that interpretation valid.
Using Relative Frequencies to Compare Different-Sized Groups
Choose relative frequencies for fair comparison when groups contain different numbers of observations.
Limits of Exact-Value Recovery from Histograms
Explain why a histogram usually cannot recover exact individual data values from binned bars.
Stemplots and Stem-and-Leaf Structure
Build and read stem-and-leaf plots while preserving the original numerical observations.
Stemplot Keys
Interpret a stemplot key correctly, including decimal place and measurement units.
Required Empty Stems to Reveal Gaps
Include empty stems when needed so gaps and distribution shape are not visually hidden.
Cumulative Relative Frequency
Compute and interpret cumulative relative frequency as the running proportion at or below a value.
Ogive or Cumulative Relative-Frequency Plots
Construct and read an ogive using cumulative relative frequencies and ordered class boundaries.
Reading proportions from an ogive
Read a proportion at or below a chosen value from an ogive with appropriate interpolation caution.
Reading the median from a cumulative plot
Locate the median on a cumulative plot near the 50th percentile and translate it back to the data scale.
Reading quartile locations from a cumulative plot
Locate first and third quartiles near the 25th and 75th percentiles on a cumulative graph.
Dotplot advantage of retaining individual observations
Explain why dotplots are useful when individual observations and small-sample detail matter.
Dotplot limitation for very large data sets
Recognize when a dotplot becomes cluttered or inefficient for a very large data set.
Stemplot advantage of retaining individual observations
Explain how stemplots retain exact values while also revealing shape, center, and spread.
Stemplot limitation for very large data sets
Recognize when too many stems or observations make a stemplot impractical.
Histogram advantage for large data sets
Explain why histograms summarize large numerical data sets efficiently and reveal overall shape.
Histogram disadvantage of hiding individual data values
Explain the trade-off of histogram binning: clear shape but loss of exact individual values.
SOCS framework for distribution descriptions
Use SOCS—shape, outliers, center, spread—to produce a complete contextual distribution description.
CUSS framework for distribution descriptions
Use the CUSS framework consistently when the course or teacher expects that description structure.
Context in a distribution description
Write distribution descriptions in the variable's real context rather than as disconnected graph vocabulary.
Symmetric distribution shape
Identify symmetry and explain what balanced tails imply for center comparisons.
Right-skewed distribution shape
Identify right skew from the longer right tail and reason about its effect on mean versus median.
Left-skewed distribution shape
Identify left skew from the longer left tail and reason about resistant versus nonresistant summaries.
Bell-shaped distribution form
Recognize bell-shaped form without assuming normality unless additional evidence supports that model.
Uniform distribution form
Recognize a roughly uniform distribution and explain what similar frequencies across intervals mean.
Unimodal distributions
Identify a unimodal distribution by one clear dominant peak while still describing skew and spread separately.
Bimodal distributions
Identify bimodality and discuss how two peaks may suggest distinct subgroups or generating processes.
Move from topic mastery to full AP Statistics practice
Once a topic feels secure, use mixed practice so the method is not given away by the page title.
A four-pass workflow for every topic
Identify
Name the variable type, display, probability structure, or statistical task before touching a formula or calculator.
Execute
Compute, construct, or compare carefully. Track denominators, units, graph scales, and assumptions rather than relying on pattern matching.
Interpret
Translate the result back into the real context. For graphs, describe what the display supports and what it cannot reveal.
Transfer
Repeat the idea inside mixed MCQs and FRQs where the correct method is no longer announced by the topic heading.
What to practice inside each topic bank
Questions about using this MCQ + FRQ hub
Is this page only for AP Statistics FRQ practice?
No. The page keeps the existing free-response practice bank and adds direct access to MCQ practice, full tests, and 55 dedicated practice-microtopic pages.
How should I use the 55 practice microtopic pages?
Start with a topic you cannot explain confidently, complete its MCQs before checking answers, then write the FRQ response in full sentences with conditions, calculations, and a contextual conclusion.
What is the AP Statistics exam format for 2027?
The current exam structure is fully digital with 42 multiple-choice questions in 90 minutes and 4 free-response questions in 90 minutes, with each section worth 50 percent of the exam score.
Should I practice MCQs or FRQs first?
Use both. MCQs are efficient for concept selection and error detection, while FRQs expose weaknesses in statistical reasoning, notation, condition checks, and contextual interpretation.
Does the topic progress tracker save my personal data?
The tracker uses browser local storage only. It does not require an account and does not send practice progress to the server.
Will activating this plugin change the page URL or delete the old post?
No. The plugin leaves the existing URL, post, publication data, title, and canonical handling unchanged. Deactivating the plugin returns the original frontend content.
Practice the actual 2027 architecture: 4 free-response questions in 90 minutes, 10 raw points each, fully typed in Bluebook. This bank contains 12 original four-question sets = 48 FRQs, with current-course content only, worked answers, diagnostic rubrics, response boxes, self-scoring, and a timed set mode.
The revised 2027 FRQ section is a different assessment design
The change is not simply “two fewer questions.” Each 2027 FRQ collects more evidence and is worth 10 raw points. The old six-question pattern and 0–4 holistic scoring used in historical resources should not be copied as the current exam blueprint.
Formulate Questions + Collect Data
Investigative questions, populations and variables, sampling, experiments, randomization, bias, and scope of inference.
Analyze Data + Interpret Results
Graphical/numerical evidence, probability or model output, comparisons, regression interpretation, and conclusions grounded in context.
Inference
Classify the setting, choose the correct current procedure, check conditions, calculate, and conclude about the population parameter.
Multiple Content Areas
Connect design, probability, data analysis, sampling distributions, regression, or categorical reasoning without relying on one memorized routine.
Choose a complete 2027 four-question set
Work one set without opening the answers. Your typed drafts and self-scores are stored only in your browser. Use the 90-minute timer for a full section rehearsal, then reveal answers and score each question from 0–10.
Your active-set breakdown
The four roles diagnose different weaknesses. A low total is less useful than knowing where points were lost.
Score the four questions after reviewing their worked answers to generate a targeted next step.
Practice Set 1
Sampling, descriptive statistics, one-proportion inference, and regression
School breakfast participation study
A school district wants to know whether offering a grab-and-go breakfast before first period changes the percentage of students who eat breakfast on school days. The district has 1,800 students across three high schools. Administrators first want a representative estimate of current breakfast participation and then want evidence about whether the new program causes a change.
- Write an investigative question that identifies the population and response variable.
- Describe a stratified random sample that would estimate current participation while representing all three schools.
- Describe a randomized experiment for evaluating the grab-and-go program. State the experimental units, treatments, response, and random-assignment mechanism.
- Explain what random sampling and random assignment each permit the district to conclude.
Reveal worked answer + practice scoring guide
Worked answer
- A suitable question is: Among students enrolled in the district’s three high schools, what proportion eat breakfast on a typical school day, and does access to grab-and-go breakfast increase that proportion? The population is all 1,800 district high-school students; the response is whether a student eats breakfast on the measured school day.
- Separate the roster by high school, randomly sample students within each school, and choose sample sizes proportional to each school’s enrollment. That preserves representation while keeping selection random within strata.
- Use students (or naturally intact classrooms if the treatment must be delivered by classroom) as experimental units. Randomly assign comparable units to grab-and-go access or usual breakfast access, keep measurement rules the same, and compare breakfast participation.
- Random sampling supports generalization to the population from which the sample was randomly selected. Random assignment supports a cause-and-effect conclusion for the experimental units, assuming implementation is otherwise comparable. Having both gives the strongest combination of generalizability and causal inference.
10-point practice rubric
2 pts investigative question/population/response · 2 pts valid stratified sample · 3 pts randomized experiment details · 3 pts scope-of-inference explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Commute-time distribution with a possible outlier
Nine students report one-way commute times, in minutes:
12, 14, 14, 15, 16, 17, 18, 22, 30
- Calculate the mean, median, sample standard deviation, Q1, Q3, and IQR.
- Use the 1.5×IQR rule to identify any outlier.
- Which pair of summary statistics—mean and standard deviation or median and IQR—better represents a typical commute here? Justify.
- If the 30-minute value were replaced by 42 minutes, describe how the mean, median, standard deviation, and IQR would change.
Reveal worked answer + practice scoring guide
Worked answer
The mean is 17.56 minutes, the median is 16, and the sample standard deviation is about 5.48. With the median excluded when splitting the ordered list, Q1 = 14, Q3 = 20, so IQR = 6.
The fences are 14 − 1.5(6) = 5 and 20 + 1.5(6) = 29; therefore 30 is a high outlier by this rule. Median and IQR are the more resistant summaries for a distribution with this unusually large value. Replacing 30 with 42 increases the mean and standard deviation substantially, while the median remains 16 and the quartile positions—and therefore IQR—remain unchanged for this ordered sample.
10-point practice rubric
4 pts correct numerical summaries · 2 pts outlier rule/fences · 2 pts resistant-summary justification · 2 pts correct effect of replacing the maximum
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.One-proportion confidence interval for library use
A simple random sample of 140 students at a large university finds that 84 used the library during the previous week.
- Identify the parameter and calculate the sample proportion.
- Check the conditions for a one-proportion z interval, including the large-count condition and the 10% condition.
- Construct a 95% confidence interval for the population proportion.
- Interpret the interval in context without saying that there is a 95% probability the fixed parameter lies inside this particular interval.
Reveal worked answer + practice scoring guide
Worked answer
Let p be the true proportion of all students at the university who used the library during the previous week. The estimate is p̂ = 84/140 = 0.600.
The sample was random; if the university has at least 1,400 students, the 10% condition is satisfied. The sample contains 84 successes and 56 failures, both at least 10. The standard error is about 0.0414, so the 95% interval is 0.600 ± 1.96(0.0414) = (0.519, 0.681).
Interpretation: We are 95% confident that the true proportion of all students at this university who used the library during the previous week is between about 51.9% and 68.1%.
10-point practice rubric
2 pts parameter/estimate · 2 pts conditions with evidence · 3 pts correct interval calculation · 3 pts contextual confidence interpretation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Study time and quiz performance
For six students, weekly study hours x and quiz score y are:
x: 1, 2, 3, 4, 5, 6 | y: 62, 65, 69, 74, 78, 80
- Find the least-squares regression line and correlation.
- Interpret the slope in context and report r².
- Predict the score for 4.5 study hours. For the student who studied 4 hours and scored 74, calculate the residual.
- Explain why this strong linear association, by itself, does not prove that increasing study time causes a higher quiz score.
Reveal worked answer + practice scoring guide
Worked answer
The least-squares line is ŷ = 57.933 + 3.829x, with r ≈ 0.995 and r² ≈ 0.989. Within this observed range, each additional study hour is associated with an increase of about 3.83 predicted quiz points. About 98.9% of the observed variation in quiz scores is explained by the linear relationship with study hours.
At x = 4.5, the predicted score is about 75.16. At x = 4, ŷ ≈ 73.25, so the residual is 74 − 73.25 ≈ +0.75. Because this is observational, students who study more may differ in prior preparation, motivation, course attendance, or other variables; association does not establish causation.
10-point practice rubric
3 pts line/correlation · 2 pts slope and r² interpretations · 3 pts prediction/residual · 2 pts causal limitation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 2
Randomized blocks, two-way tables, t intervals, and binomial/sampling distributions
Study-app reminder experiment
A high school wants to test whether nightly study-app reminders improve completion of assigned review problems. Because ninth- and twelfth-grade students have very different schedules, the researchers want grade level accounted for in the design.
- State a statistical question the experiment could answer.
- Describe a randomized block design using grade level as the blocking variable.
- Name one variable that should be measured consistently across treatment groups and explain why.
- State the strongest causal and generalization conclusions the school could make if volunteers, rather than a random sample, participate.
Reveal worked answer + practice scoring guide
Worked answer
A suitable question asks whether receiving nightly reminders changes the mean proportion of assigned review problems completed. Create separate ninth-grade and twelfth-grade blocks, then randomly assign students within each block to reminders or no reminders. Use the same assignment length, scoring rule, and observation period for both treatment groups so treatment is not confounded with measurement.
Random assignment can support a causal comparison for the participating students. Because volunteers were not randomly sampled from the whole school, broad generalization to every student is not automatically justified.
10-point practice rubric
2 pts clear statistical question · 3 pts correct block/randomization design · 2 pts controlled measurement explanation · 3 pts causal/generalization scope
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Program participation and course completion
Among 72 students who joined a voluntary after-school support program, 54 completed the course successfully. Among 78 students who did not join, 45 completed successfully.
- Compute the conditional completion proportion for each group.
- Compute the overall completion proportion.
- Describe the association using an absolute percentage-point difference.
- Explain why these data alone do not establish that the program caused the higher completion rate.
Reveal worked answer + practice scoring guide
Worked answer
Participants: 54/72 = 0.750. Nonparticipants: 45/78 ≈ 0.577. Overall: 99/150 = 0.660. The participant completion rate is about 17.3 percentage points higher.
This is an association, but participation was voluntary. Motivation, prior achievement, available time, and other variables can differ between those who chose the program and those who did not. Without random assignment, the observed difference does not by itself establish causation.
10-point practice rubric
3 pts two conditional proportions · 2 pts overall proportion · 2 pts correct association size · 3 pts confounding/causation explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Mean response-time confidence interval
A random sample of 36 service requests has mean response time 72.4 minutes and sample standard deviation 10.8 minutes. Assume the distribution has no extreme outliers.
- Identify the parameter and appropriate inference procedure.
- Check conditions for a one-sample t interval.
- Construct a 95% confidence interval for the population mean response time.
- Interpret the interval in context.
Reveal worked answer + practice scoring guide
Worked answer
Let μ be the population mean response time. A one-sample t interval is appropriate because σ is unknown and is estimated by s. The random sample supports independence; if sampling without replacement, the population should be at least 360 requests. With n = 36 and no extreme outliers, the t procedure is reasonable.
SE = 10.8/√36 = 1.8. With df = 35, t* ≈ 2.030. The interval is 72.4 ± 2.030(1.8) = (68.75, 76.05) minutes. We are 95% confident that the population mean response time lies between about 68.8 and 76.1 minutes.
10-point practice rubric
2 pts parameter/procedure · 2 pts conditions · 3 pts interval mechanics · 3 pts interpretation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Rare defects: binomial model and sample proportion
A production process historically has defect probability p = 0.08 for independently produced items.
- For 20 items, find P(X ≤ 1) if X is the number of defects. State the binomial conditions.
- Find the mean and standard deviation of X.
- For a separate random sample of 100 items, give the approximate sampling distribution of p̂ and estimate P(p̂ > 0.12).
- Explain why the normal approximation for p̂ needs a large-count check even though the exact binomial calculation in part (a) does not.
Reveal worked answer + practice scoring guide
Worked answer
For X ~ Binomial(20, 0.08), P(X ≤ 1) ≈ 0.5169. The binomial model requires binary outcomes, independent trials, fixed n, and constant p. E(X) = np = 1.6 and SD(X) = √np(1−p) ≈ 1.213.
For n = 100, μp̂ = 0.08 and SDp̂ = √[0.08(0.92)/100] ≈ 0.02713. Using a normal approximation gives z ≈ 1.474 for p̂ = 0.12, so P(p̂ > 0.12) ≈ 0.0702. The exact binomial model does not require a normal shape; the p̂ normal approximation does, so np and n(1−p) must be sufficiently large.
10-point practice rubric
3 pts exact binomial probability/conditions · 2 pts mean/SD · 3 pts p̂ distribution/probability · 2 pts approximation explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 3
Sampling bias, Normal models, two-proportion testing, and regression interpretation
City-park satisfaction survey
A city wants to estimate satisfaction with public parks among adult residents. An initial plan is to post a QR code only at the city’s largest downtown park.
- Identify two sources of bias in the QR-code plan.
- Describe a stratified random sample using geographic districts.
- Write an investigative question that avoids vague wording such as “Do you like our excellent parks?”
- Explain how nonresponse could still bias the final estimate even after a good random sample is selected.
Reveal worked answer + practice scoring guide
Worked answer
The QR-code plan undercovers residents who do not visit that park and creates voluntary-response bias because people who choose to scan may have unusually strong opinions. A stronger plan uses a resident sampling frame, divides residents into geographic districts, randomly samples within every district, and weights only if the sampling plan requires it.
A neutral question could ask, “On a scale from 1 to 5, how satisfied are you with the condition and accessibility of city public parks during the past six months?” Nonresponse can still bias results if responders and nonresponders systematically differ in satisfaction.
10-point practice rubric
3 pts two biases correctly explained · 3 pts valid stratified sample · 2 pts neutral measurable question · 2 pts nonresponse mechanism
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Normal model for placement scores
Placement scores are modeled as Normal with mean 68 and standard deviation 8.
- Find the z-score for a score of 80 and its percentile.
- Find P(60 < X < 76).
- A student says that 80 is “93.3% above the mean.” Correct the interpretation.
- State one reason a Normal-model probability could be inappropriate even if a z-score can still be calculated.
Reveal worked answer + practice scoring guide
Worked answer
For x = 80, z = (80−68)/8 = 1.50, and Φ(1.50) ≈ 0.9332, so the score is around the 93.3rd percentile. For 60 to 76, the z-scores are −1 and +1, giving probability ≈ 0.6827.
The correct statement is that about 93.3% of values in the Normal model are at or below 80—not that 80 is “93.3% above the mean.” Standardization is always arithmetic, but a Normal probability statement depends on a defensible Normal model for the distribution.
10-point practice rubric
3 pts z/percentile · 2 pts interval probability · 3 pts interpretation correction · 2 pts model-validity explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Two-proportion test for message response
In independent random samples, 72 of 120 people shown message A respond positively, compared with 54 of 120 shown message B. Test H0: pA = pB against Ha: pA ≠ pB at α = 0.05.
- Define the two population parameters and check conditions.
- Compute the pooled proportion and z statistic.
- Find the two-sided p-value and make a decision.
- State the conclusion in context without saying that the null hypothesis has a probability of being true.
Reveal worked answer + practice scoring guide
Worked answer
p̂A = 0.600, p̂B = 0.450, and the pooled estimate under H0 is 126/240 = 0.525. The random/independent samples and large expected success/failure counts support the two-proportion z test.
SE under H0 ≈ 0.06447, z ≈ 2.327, and the two-sided p-value ≈ 0.0200. Reject H0 at α = 0.05. The data provide convincing evidence that the population positive-response proportions differ between message A and message B.
10-point practice rubric
2 pts parameters/conditions · 3 pts pooled SE/z · 2 pts p-value/decision · 3 pts contextual conclusion
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Training time and task completion
A small observational study records training hours x and tasks completed y:
x: 2, 4, 6, 8, 10 | y: 18, 25, 31, 37, 44
- Find the least-squares line, r, and r².
- Predict y when x = 7. For the observation (8, 37), calculate the residual.
- Interpret the slope and r² in context.
- Explain why using the line to predict at x = 25 hours would be much less defensible than predicting at x = 7.
Reveal worked answer + practice scoring guide
Worked answer
The line is ŷ = 11.8 + 3.2x, r ≈ 0.9995, and r² ≈ 0.9990. At x = 7, ŷ = 34.2. At x = 8, ŷ = 37.4, so the residual is −0.4.
Within the observed range, each additional training hour is associated with about 3.2 more predicted tasks completed. Roughly 99.9% of observed variation in y is explained by the linear relationship with x. Predicting at 25 hours is extreme extrapolation far beyond the observed 2–10 hour range; the relationship need not continue there.
10-point practice rubric
3 pts regression/r/r² · 2 pts prediction/residual · 3 pts interpretations · 2 pts extrapolation warning
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 4
Matched pairs, sampling distributions, paired inference, and chi-square reasoning
Standing-desk matched-pairs experiment
A researcher wants to compare typing accuracy while seated versus while using a standing desk. Each of 30 volunteers can complete two comparable typing tasks.
- Explain why a matched-pairs design is natural here.
- Describe how to randomize the order of the two conditions for each participant.
- Define the paired difference so a positive value has a clear interpretation.
- State one potential carryover effect and how the design could reduce it.
Reveal worked answer + practice scoring guide
Worked answer
Each participant serves as their own control, which removes much person-to-person variation in baseline typing ability. Randomly assign each participant to seated-first or standing-first, using a fair random mechanism. Define d = accuracy while standing − accuracy while seated, so d > 0 means better accuracy while standing.
A carryover effect could occur if fatigue or practice from the first task influences the second. Use comparable but different tasks, allow a washout/rest period, and randomize order to balance order effects.
10-point practice rubric
2 pts why pairing fits · 3 pts randomized order · 2 pts clear difference definition · 3 pts carryover/control explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Sampling distribution of a mean
A population has mean 120 and standard deviation 24. Independent random samples of n = 64 are taken.
- State the mean and standard deviation of the sampling distribution of x̄.
- Find P(x̄ > 126) if the sampling distribution is approximately Normal.
- Explain why the standard deviation of x̄ is smaller than the population standard deviation.
- If n increases to 144, what happens to the standard error?
Reveal worked answer + practice scoring guide
Worked answer
μx̄ = 120 and σx̄ = 24/√64 = 3. For x̄ = 126, z = 2, so P(x̄ > 126) ≈ 0.0228.
Averages vary less from sample to sample than individual observations vary in the population; averaging n independent observations reduces standard deviation by √n. At n = 144, the standard error becomes 24/12 = 2.
10-point practice rubric
2 pts center/spread · 3 pts probability · 3 pts conceptual explanation · 2 pts new standard error
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Paired t test for time saved
For 12 matched participants, define d = old-system completion time − new-system completion time, in minutes. The sample has d̄ = 3.4 and sd = 2.1. Test H0: μd = 0 versus Ha: μd > 0.
- Explain what μd represents and check the paired-t conditions.
- Calculate the test statistic and degrees of freedom.
- Find the one-sided p-value and make a decision at α = 0.05.
- Give a 95% confidence interval for μd and connect it to the test conclusion.
Reveal worked answer + practice scoring guide
Worked answer
μd is the population mean time saved by the new system under the chosen difference definition. The data must be matched, the pairs should be independent, and the distribution of differences should have no severe skew/outliers for n = 12.
t = 3.4/(2.1/√12) ≈ 5.609, df = 11, and the one-sided p-value is about 0.00008. Reject H0. There is very strong evidence that the new system reduces mean completion time. A 95% CI is approximately (2.07, 4.73) minutes saved, which lies entirely above 0 and agrees with the test.
10-point practice rubric
2 pts parameter/conditions · 2 pts t/df · 3 pts p-value/decision · 3 pts CI and contextual link
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Device type and preferred study format
A random sample produces the following table of device type (tablet/other) by preferred study format (video/text/mixed):
Tablet: 42, 18, 10 | Other: 28, 25, 17
- Find the expected counts under independence.
- Compute χ² and degrees of freedom.
- Use the p-value to interpret evidence of association at α = 0.05.
- Explain why expected counts—not observed counts—are checked for the chi-square condition, and why a random sample supports generalization but not causation.
Reveal worked answer + practice scoring guide
Worked answer
Row totals are both 70; column totals are 70, 43, and 27, with N = 140. Expected counts for each row are 35, 21.5, 13.5. The chi-square statistic is χ² ≈ 5.754 with df = 2, giving p ≈ 0.0563.
At α = 0.05, fail to reject independence; the sample does not provide quite enough evidence of an association between device type and preferred study format. The approximation is based on the null-model expected cell counts, so those—not the observed counts—control the large-count check. Random sampling can support generalization to the sampled population, but the variables were not randomly assigned, so causation is not justified.
10-point practice rubric
3 pts expected counts · 2 pts χ²/df · 3 pts p-value/conclusion · 2 pts condition/scope explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 5
Survey quality, comparative distributions, two-sample t intervals, and discrete random variables
Public-transit satisfaction survey
A transit agency wants to estimate satisfaction among all weekday riders. A manager proposes emailing only riders who bought monthly passes online.
- Explain the undercoverage problem in the proposed frame.
- Describe a better sampling plan that can include cash, mobile-ticket, and pass users.
- Write one neutral satisfaction question with a clearly defined time window.
- Describe how you would investigate whether nonresponse is threatening the estimate.
Reveal worked answer + practice scoring guide
Worked answer
The email list excludes riders who pay by cash, use anonymous cards, or buy tickets through channels not linked to email, so it does not represent all weekday riders. A stronger plan samples trips, stops, or rider records across routes and times, then randomly selects riders within the chosen frame so each major rider type has a chance to be selected.
A neutral question might ask, “During your weekday transit trips in the past 30 days, how satisfied were you with on-time performance on a 1–5 scale?” Compare response rates across sampled route/time/rider strata and follow up with a random subset of nonresponders if feasible; systematic differences in response can bias the result.
10-point practice rubric
2 pts undercoverage · 3 pts broader random-sampling plan · 2 pts neutral measurable item · 3 pts nonresponse assessment
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Comparing two score distributions from summaries
Two sections of a course have the following summaries:
Section A: mean 14.2, SD 3.1, median 13.9, IQR 4.2
Section B: mean 16.0, SD 5.4, median 14.1, IQR 4.4
- Compare center using both resistant and nonresistant measures.
- Compare spread using both SD and IQR.
- What feature of Section B is suggested by the large mean–median gap and large SD relative to its IQR?
- Explain why summary statistics alone cannot prove a particular distribution shape.
Reveal worked answer + practice scoring guide
Worked answer
Section B has the higher mean by 1.8 points, but the medians are almost the same (14.1 vs 13.9). Section B also has much larger SD (5.4 vs 3.1), while the IQRs are similar (4.4 vs 4.2). That pattern is consistent with a high-end tail or influential high values in Section B, because nonresistant measures changed much more than resistant measures.
However, numerical summaries do not uniquely determine shape. A graph or the raw data would be needed to identify skewness, multimodality, gaps, or specific outliers conclusively.
10-point practice rubric
2 pts centers · 2 pts spreads · 3 pts defensible pattern interpretation · 3 pts limitation of summaries
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Two-sample t interval for mean difference
Independent random samples produce:
Group 1: n = 40, x̄ = 82.1, s = 9.2
Group 2: n = 35, x̄ = 77.4, s = 8.5
- Define μ1 − μ2 in context and state the appropriate procedure.
- Check the independence and shape/sample-size conditions.
- Construct a 95% confidence interval using a Welch two-sample t procedure.
- Interpret the interval and state whether 0 is a plausible value for the population mean difference.
Reveal worked answer + practice scoring guide
Worked answer
The observed difference is 82.1 − 77.4 = 4.7. The standard error is about 2.045; Welch df is about 72.8, giving t* ≈ 1.993. The 95% CI is approximately (0.62, 8.78).
Assuming independent random samples and no severe distribution problems for these sample sizes, we are 95% confident that Group 1’s population mean exceeds Group 2’s by about 0.6 to 8.8 units. Because 0 is not in the interval, equal population means are not supported by this interval at the corresponding two-sided 5% level.
10-point practice rubric
2 pts parameter/procedure · 2 pts conditions · 3 pts interval calculation · 3 pts contextual interpretation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Expected value of daily alerts
A device sends X alerts per day with probability distribution:
X: 0, 1, 2, 3 | P(X): 0.10, 0.25, 0.40, 0.25
- Verify this is a valid probability distribution.
- Find E(X) and SD(X).
- Interpret E(X) in repeated-use context.
- If a manager observes 18 alerts over 10 days, explain why comparing 18 directly with SD(X) would be a scale mistake.
Reveal worked answer + practice scoring guide
Worked answer
The probabilities are all between 0 and 1 and sum to 1. E(X) = 1.8 alerts/day. The standard deviation is about 0.927 alerts/day.
Over many days under the same process, the long-run average number of alerts per day approaches 1.8. Eighteen alerts is a 10-day total, while SD(X) describes one-day variability; they are different random quantities on different scales. A 10-day total would need its own distributional analysis rather than being compared directly with the one-day standard deviation.
10-point practice rubric
2 pts validity · 3 pts mean/SD · 2 pts expected-value interpretation · 3 pts scale/random-variable distinction
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 6
Blocked experiments, regression, chi-square homogeneity, and sampling distributions of proportions
Fertilizer trial blocked by field
An agricultural researcher compares fertilizers A and B on 60 plots located in three fields with noticeably different soil quality.
- Write an investigative question about average yield.
- Describe a randomized block design using field as the block.
- Explain what blocking is intended to accomplish.
- Distinguish the role of replication from the role of random assignment.
Reveal worked answer + practice scoring guide
Worked answer
Within each field, randomly assign approximately half the plots to fertilizer A and half to B, then compare yields within blocks before combining evidence. Blocking removes predictable field-to-field soil variation from the treatment comparison, improving precision.
Replication provides repeated observations under each treatment and lets random variation be assessed. Random assignment balances lurking variables in expectation and supports a causal interpretation of treatment differences; replication alone does not create causal validity.
10-point practice rubric
2 pts investigative question · 3 pts correct blocked randomization · 2 pts purpose of blocking · 3 pts replication vs randomization
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice sessions and error count
Six learners have practice sessions x and error count y:
x: 1, 2, 3, 4, 5, 6 | y: 18, 15, 13, 10, 9, 7
- Find the least-squares line, r, and r².
- Interpret the slope and r².
- At x = 4, calculate the predicted error count and residual for the observed y = 10.
- Explain why a negative slope does not imply the response variable itself is “negative.”
Reveal worked answer + practice scoring guide
Worked answer
The line is ŷ = 19.6 − 2.171x, r ≈ −0.991, and r² ≈ 0.982. Each additional practice session is associated with about 2.17 fewer predicted errors. About 98.2% of observed variation in error count is explained by the linear relationship with practice sessions.
At x = 4, ŷ ≈ 10.914, so the residual is 10 − 10.914 ≈ −0.914. A negative slope describes the direction of association between x and y; it does not mean y must take negative values within the observed range.
10-point practice rubric
3 pts line/r/r² · 3 pts interpretations · 2 pts prediction/residual · 2 pts slope-sign concept
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Chi-square test of homogeneity
Three independently sampled regions each contribute 50 respondents. Counts for Yes/No are:
Region 1: 32 / 18
Region 2: 28 / 22
Region 3: 20 / 30
- State hypotheses for a chi-square test of homogeneity.
- Find the expected counts and verify the large-count condition.
- Compute χ², df, and p-value.
- Conclude at α = 0.05.
Reveal worked answer + practice scoring guide
Worked answer
H0: the Yes/No population distribution is the same across the three regions. Ha: at least one region has a different distribution. Since 80 Yes and 70 No responses occur among 150 total, each 50-person row has expected counts 26.67 Yes and 23.33 No, all well above 5.
χ² = 6.000, df = 2, p ≈ 0.0498. At α = 0.05, reject H0; the data provide evidence that the Yes/No population distribution is not the same across all three regions.
10-point practice rubric
2 pts hypotheses · 2 pts expected counts/condition · 3 pts statistic/df/p · 3 pts contextual conclusion
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Sampling distribution of a survey proportion
Suppose the true population proportion supporting a proposal is p = 0.62. A simple random sample of 150 people is selected.
- Give the mean and standard deviation of p̂.
- Check whether a Normal approximation is reasonable.
- Approximate P(p̂ < 0.55).
- Explain why a sample result below 0.55 would be unusual but not impossible, and distinguish this probability model from a claim that p has changed.
Reveal worked answer + practice scoring guide
Worked answer
μp̂ = 0.62 and SDp̂ = √[0.62(0.38)/150] ≈ 0.03963. np = 93 and n(1−p) = 57, so the Normal approximation is strong.
For p̂ = 0.55, z ≈ −1.766, giving P(p̂ < 0.55) ≈ 0.0387. That is uncommon under p = 0.62, but random samples can produce tail outcomes. The calculation assumes p = 0.62; observing one unusual sample is evidence to investigate, not proof by itself that the parameter changed.
10-point practice rubric
2 pts sampling-distribution center/spread · 2 pts large counts · 3 pts probability · 3 pts evidence-vs-proof interpretation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 7
Confounding, discrete random variables, one-proportion testing, and sampling distributions of means
Exercise app observational study
Researchers compare weekly exercise time for people who voluntarily use a fitness app with people who do not.
- State a clear investigative question.
- Identify two plausible confounding variables.
- Explain why matching users and nonusers on age alone would not automatically remove all confounding.
- Describe a randomized experiment that could test one app feature causally without forcing anyone to own a particular phone.
Reveal worked answer + practice scoring guide
Worked answer
A suitable question asks whether app use is associated with mean weekly exercise time in a defined population. Baseline motivation, health status, access to exercise facilities, income, and prior activity are plausible confounders. Matching only on age balances one variable, not every systematic difference between self-selected users and nonusers.
For a causal feature test, recruit eligible participants who all have access to the study platform, then randomly assign one group to receive the feature (for example, goal reminders) and the other to a control version. Compare a prespecified exercise outcome using the same measurement window.
10-point practice rubric
2 pts investigative question · 2 pts confounders · 3 pts matching limitation · 3 pts randomized feature experiment
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Discrete reward distribution
A game awards X points with:
X: 0, 5, 10, 15 | P(X): 0.20, 0.40, 0.30, 0.10
- Find E(X) and SD(X).
- Interpret both values in context.
- Is E(X) required to be one of the possible outcomes? Explain.
- Explain how changing probability from the 0-point outcome to the 15-point outcome would affect the mean.
Reveal worked answer + practice scoring guide
Worked answer
E(X) = 6.5 points and SD(X) = 4.5 points. Over many plays, the average award approaches 6.5 points; the SD describes the typical scale of one-play deviation around that mean.
An expected value is a long-run weighted average and need not be an attainable single-play outcome. Shifting probability mass from 0 to 15 increases the expected value because more probability is placed on a larger outcome.
10-point practice rubric
3 pts mean/SD · 3 pts interpretations · 2 pts expected-value concept · 2 pts probability-shift effect
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.One-proportion significance test and power
A random sample of 150 customers finds 92 who prefer a redesigned checkout. Test H0: p = 0.55 versus Ha: p > 0.55 at α = 0.05.
- Check conditions using the null value.
- Compute the z statistic and p-value.
- Make a decision and conclude in context.
- Describe a Type I error and state one design change that generally increases power against a fixed alternative.
Reveal worked answer + practice scoring guide
Worked answer
Under H0, np0 = 82.5 and n(1−p0) = 67.5; with random sampling and the 10% condition if relevant, the one-proportion z test is appropriate. p̂ = 92/150 ≈ 0.6133, z ≈ 1.559, and the one-sided p-value ≈ 0.0595.
At α = 0.05, fail to reject H0; the sample does not provide quite enough evidence that the population preference proportion exceeds 0.55. A Type I error would be concluding p > 0.55 when in fact p = 0.55. Increasing sample size generally increases power for a fixed true alternative.
10-point practice rubric
2 pts conditions · 3 pts z/p · 2 pts decision/conclusion · 3 pts Type I + power
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Sampling distribution under two sample sizes
A process has individual measurements with μ = 500 and σ = 60.
- For samples of n = 36, state the sampling distribution center and standard error of x̄.
- Assuming an approximately Normal sampling distribution, find P(485 < x̄ < 515).
- Repeat the probability comparison conceptually for n = 144 and calculate the new probability.
- Explain why increasing n changes the sampling distribution of x̄ but does not change the population distribution of individual observations.
Reveal worked answer + practice scoring guide
Worked answer
For n = 36, μx̄ = 500 and SE = 60/6 = 10. The bounds 485 and 515 correspond to z = −1.5 and +1.5, so the probability is about 0.8664.
For n = 144, SE = 60/12 = 5; the same ±15 interval is now ±3 SE, giving probability about 0.9973. Larger n concentrates sample means more tightly around μ. It does not shrink the spread of the underlying individual measurements.
10-point practice rubric
2 pts center/SE · 3 pts n=36 probability · 3 pts n=144 comparison · 2 pts population-vs-sampling-distribution distinction
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 8
Multistage sampling, linear transformations, two-proportion intervals, and categorical-data interpretation
Statewide school survey with multistage sampling
A state education office wants to survey high-school students but cannot obtain one statewide student roster. It does have a complete list of schools and enrollment counts.
- Describe a multistage probability sample that begins with schools.
- Explain one advantage and one potential source of extra sampling variability compared with an SRS of individual students.
- State how school size should be considered so students at small and large schools are not accidentally given very different selection probabilities.
- Explain why convenience sampling the nearest schools would not be repaired merely by having a very large sample.
Reveal worked answer + practice scoring guide
Worked answer
Randomly select schools using a probability method, then obtain rosters within selected schools and randomly sample students from each selected school. School selection can use probability proportional to enrollment or the second-stage sample sizes can be adjusted so student inclusion probabilities are known and defensible.
Multistage sampling is operationally cheaper, but students within a school may be more similar than randomly selected students statewide, which can increase sampling variability. A huge convenience sample can still be systematically biased because size does not replace random selection.
10-point practice rubric
3 pts valid multistage design · 2 pts advantage/variance tradeoff · 3 pts unequal school-size handling · 2 pts convenience-sample bias
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Linear transformation of a distribution
A variable X has mean 52, SD 8, median 50, and IQR 10. Define Y = 1.8X + 32.
- Find the mean and SD of Y.
- Find the median and IQR of Y.
- Explain what happens to z-scores under this positive linear transformation.
- Would the shape of a histogram change, aside from the horizontal scale and labels? Explain.
Reveal worked answer + practice scoring guide
Worked answer
Mean(Y) = 1.8(52)+32 = 125.6; SD(Y) = 1.8(8) = 14.4. Median(Y) = 1.8(50)+32 = 122; IQR(Y) = 1.8(10) = 18.
For a positive linear transformation, each value and the center/scale transform together, so z-scores are unchanged. The distribution’s shape and ordering are preserved; only location, scale, and axis labels change.
10-point practice rubric
2 pts mean/SD · 2 pts median/IQR · 3 pts z-score invariance · 3 pts shape explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Two-proportion confidence interval
Independent random samples find 110 successes among 180 people in Group 1 and 84 successes among 160 in Group 2.
- Define p1 − p2 and compute the sample difference.
- Check large-count and independence conditions.
- Construct a 95% confidence interval for p1 − p2.
- Interpret the interval, including what it means that 0 is inside the interval.
Reveal worked answer + practice scoring guide
Worked answer
p̂1 = 110/180 ≈ 0.6111, p̂2 = 84/160 = 0.5250, so the difference is 0.0861. Both samples have at least 10 successes and failures, and the groups/samples must be independent; the 10% condition applies if sampling without replacement.
SE ≈ 0.05366. The 95% CI is 0.0861 ± 1.96(0.05366) = (−0.019, 0.191). The plausible values include 0, so the interval does not establish a nonzero population difference at the corresponding two-sided 5% level.
10-point practice rubric
2 pts parameter/estimate · 2 pts conditions · 3 pts interval · 3 pts interpretation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Neighborhood and policy support
A random sample of 150 residents gives:
Zone A: 45 support, 30 oppose
Zone B: 35 support, 40 oppose
- Compute the support proportion within each zone and the overall support proportion.
- Describe the association using a percentage-point difference.
- If residents were randomly sampled from each zone, what population conclusion is supported?
- Why would even a strong association not prove that moving from one zone to another changes a person’s opinion?
Reveal worked answer + practice scoring guide
Worked answer
Zone A support = 45/75 = 0.600; Zone B support = 35/75 ≈ 0.4667; overall support = 80/150 ≈ 0.5333. Support is about 13.3 percentage points higher in Zone A.
If each zone’s residents were randomly sampled, the association can be generalized to the corresponding zone populations, subject to the sampling design. Zone residence was not randomly assigned, so other characteristics can differ between zones; the association is not evidence that changing residence itself causes opinion to change.
10-point practice rubric
3 pts proportions · 2 pts association magnitude · 2 pts generalization scope · 3 pts causal limitation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 9
Blocked designs, simulation-based evidence, one-sample t testing, and exact binomial tail reasoning
Training method experiment across experience levels
A company compares two training methods among 84 employees. Previous experience is classified as low, medium, or high and strongly predicts task speed.
- Explain why experience is a useful blocking variable.
- Describe the random assignment when the three blocks contain different numbers of employees.
- State a primary response variable and one secondary response that could reveal a speed–accuracy tradeoff.
- If participants are all volunteers from one office, separate the causal conclusion from the generalization conclusion.
Reveal worked answer + practice scoring guide
Worked answer
Blocking on experience reduces known baseline variation and makes the treatment comparison more precise. Within each experience block, randomly assign employees to the two methods in roughly equal proportions; exact counts can differ by one when a block size is odd.
A primary response might be completion time; error rate could be secondary so a faster method is not judged without accuracy. Random assignment supports causal comparison among participating employees. Because volunteers from one office are not a random sample of all company employees, company-wide generalization is limited.
10-point practice rubric
2 pts blocking purpose · 3 pts within-block randomization with unequal sizes · 2 pts meaningful outcomes · 3 pts causal/generalization distinction
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Randomization simulation p-value
Under a null model of no treatment effect, a student randomly reassigns treatment labels 1,000 times. In 27 simulations, the difference in group means is at least as large as the observed difference in the direction specified by Ha.
- Estimate the simulation-based p-value.
- Interpret that p-value in context of the null model.
- At α = 0.05, what decision is suggested?
- Explain what would improve by increasing from 1,000 to 10,000 random reassignments.
Reveal worked answer + practice scoring guide
Worked answer
The estimated p-value is 27/1000 = 0.027. If the null model were true, a result at least this favorable to Ha would occur in about 2.7% of random reassignments under this simulation scheme.
At α = 0.05, the simulation suggests rejecting the null model. Increasing the number of random reassignments reduces Monte Carlo/simulation variability in the estimated p-value; it does not magically remove bias from a poor original study design.
10-point practice rubric
2 pts p-value · 3 pts null-model interpretation · 2 pts decision · 3 pts simulation-precision explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.One-sample t test for a mean
A random sample of 25 measurements has x̄ = 104.6 and s = 9.5. Test H0: μ = 100 versus Ha: μ > 100.
- Check the conditions that can be checked from the prompt and state what additional shape information matters.
- Compute the t statistic and df.
- Find the one-sided p-value and decide at α = 0.05.
- Write a contextual conclusion using “evidence” language rather than “the null is false.”
Reveal worked answer + practice scoring guide
Worked answer
Random sampling supports independence; if sampling without replacement, the population should be at least 250. With n = 25, a roughly symmetric distribution without strong outliers is important. t = (104.6−100)/(9.5/√25) ≈ 2.421, df = 24.
The one-sided p-value is about 0.0117. Reject H0 at α = 0.05. The sample provides convincing statistical evidence that the population mean exceeds 100.
10-point practice rubric
2 pts conditions · 3 pts t/df · 2 pts p-value/decision · 3 pts conclusion wording
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Unexpected defect count under a historical model
A process historically has defect probability 0.10. In a random sample of 40 independently produced items, 8 are defective.
- Under the historical model, calculate P(X ≥ 8).
- Interpret the probability as evidence about whether the historical rate still fits.
- State the expected defect count under the historical model and compare it with the observed 8.
- Explain why this calculation can indicate model tension but cannot identify the cause of any process change.
Reveal worked answer + practice scoring guide
Worked answer
With X ~ Binomial(40, 0.10), P(X ≥ 8) ≈ 0.0419. The expected count is np = 4, so 8 defects are twice the null-model expectation.
A tail probability around 4.2% is relatively unusual if the historical defect probability is still 0.10, so the result provides evidence that the model may no longer fit. The calculation identifies statistical inconsistency, not mechanism; machine wear, material changes, operator differences, or other causes require separate process investigation.
10-point practice rubric
3 pts exact tail probability · 2 pts evidence interpretation · 2 pts expected-vs-observed comparison · 3 pts model-vs-cause distinction
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 10
Sampling frames, regression diagnostics, chi-square independence, and paired-data reasoning
Sampling from an incomplete customer frame
A service company wants to estimate satisfaction among all customers from the past year, but its email database omits some phone-only customers.
- Name the bias risk created by the frame.
- Propose a combined-frame probability-sampling strategy.
- Explain why simply increasing the email sample size does not repair the missing phone-only customers.
- Describe one way duplicate customers appearing in both frames should be handled.
Reveal worked answer + practice scoring guide
Worked answer
The missing phone-only customers create undercoverage. A better design combines the email and phone/customer-account frames, deduplicates customers, assigns known selection probabilities, and randomly samples from the resulting coverage plan.
A larger sample from the same incomplete email frame reduces random sampling error around the wrong frame-based target but does not remove systematic undercoverage. Duplicates should be identified by customer ID or another defensible matching rule so one customer does not receive unintended extra selection probability.
10-point practice rubric
2 pts undercoverage · 3 pts combined-frame random design · 2 pts sample-size misconception · 3 pts duplicate handling
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Advertising and sign-ups regression
Advertising spend x (in hundreds of dollars) and sign-ups y are:
x: 10, 20, 30, 40, 50 | y: 18, 29, 33, 47, 54
- Find the least-squares line, r, and r².
- Predict sign-ups at x = 35.
- For x = 40 with observed y = 47, find the residual.
- Interpret the slope and explain why prediction at x = 100 is an extrapolation risk.
Reveal worked answer + practice scoring guide
Worked answer
The line is ŷ = 9.2 + 0.9x, r ≈ 0.990, and r² ≈ 0.980. At x = 35, ŷ = 40.7. At x = 40, ŷ = 45.2, so the residual is +1.8.
Within the observed range, each additional hundred dollars of advertising is associated with about 0.9 more predicted sign-ups. Predicting at x = 100 doubles the largest observed x and assumes the same linear relationship far beyond the data.
10-point practice rubric
3 pts line/r/r² · 2 pts prediction · 2 pts residual · 3 pts slope/extrapolation interpretation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Chi-square independence across three device categories
A random sample classifies device category and whether a task was completed:
Category 1: 45 complete, 15 not
Category 2: 30 complete, 30 not
Category 3: 25 complete, 35 not
- State hypotheses for independence.
- Find expected counts and check the large-count condition.
- Compute χ², df, and p-value.
- Interpret the result at α = 0.05.
Reveal worked answer + practice scoring guide
Worked answer
Total complete = 100, not complete = 80, and each row total is 60. Under independence, each row has expected counts 33.33 complete and 26.67 not complete, all above 5.
χ² ≈ 14.625, df = 2, p ≈ 0.00067. Reject independence. There is strong evidence of an association between device category and task-completion status in the population represented by the random sample.
10-point practice rubric
2 pts hypotheses · 2 pts expected counts · 3 pts χ²/df/p · 3 pts contextual conclusion
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Paired improvement scores
Ten matched participants have improvement scores (after − before):
2, 5, 1, 4, 3, 6, 2, 5, 4, 3
- Find the mean, median, and sample standard deviation of the differences.
- Interpret the sign of the differences and the mean.
- Explain why an independent two-sample analysis of the raw before and after measurements would discard useful design information.
- Describe what feature of the difference distribution you would inspect before using a small-sample paired t procedure.
Reveal worked answer + practice scoring guide
Worked answer
The differences have mean 3.5, median 3.5, and sample SD ≈ 1.581. Because d = after − before, positive values indicate improvement; the sample average improvement is 3.5 units.
The same participant contributes both measurements, so the observations are paired and within-person dependence is informative. Treating before and after as independent groups throws away that pairing. With only 10 differences, inspect the difference distribution for strong skewness and outliers before relying on a paired t procedure.
10-point practice rubric
3 pts summaries · 2 pts contextual sign/mean · 3 pts why pairing matters · 2 pts small-sample shape check
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 11
Blinding and placebo design, conditional probability, two-sample t testing, and difference-in-proportions sampling
Placebo-controlled notification study
A research team tests whether a new focus-mode notification reduces interruptions during a 45-minute task.
- Identify experimental units, treatment, control, and response.
- Explain how a placebo-like control could be created for a software notification study.
- Describe what can be blinded and what may be difficult to blind.
- Explain why random assignment matters even if the response is automatically recorded.
Reveal worked answer + practice scoring guide
Worked answer
Experimental units are participating users. The treatment is the new focus-mode behavior; a control could show a visually similar neutral interface or standard notification behavior so participants have comparable study interaction. The response might be interruption count or uninterrupted-task time.
Analysts can be blinded to treatment labels during data cleaning/scoring; participants may be partly blinded if both versions look similar, though complete blinding can be difficult if behavior is noticeable. Random assignment is still essential because automatic measurement does not balance motivation, baseline phone use, task difficulty, or other lurking variables.
10-point practice rubric
3 pts units/treatments/response · 2 pts credible control · 2 pts blinding limits · 3 pts random-assignment rationale
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Conditional probability in a Normal model
Scores X are modeled as Normal with mean 70 and SD 10.
- Find P(X > 80).
- Find P(X > 60).
- Find P(X > 80 | X > 60).
- Explain why the conditional probability is larger than the unconditional probability P(X > 80).
Reveal worked answer + practice scoring guide
Worked answer
P(X > 80) = P(Z > 1) ≈ 0.1587. P(X > 60) = P(Z > −1) ≈ 0.8413. Since X > 80 is a subset of X > 60, the conditional probability is 0.1587/0.8413 ≈ 0.1886.
Conditioning restricts attention to scores already above 60, removing the entire lower tail from the reference group. Therefore scores above 80 form a larger fraction of the restricted group than of the whole population.
10-point practice rubric
2 pts first probability · 2 pts conditioning event · 3 pts conditional calculation · 3 pts conceptual explanation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Two-sample t test for mean difference
Independent random samples give:
Group 1: n = 28, x̄ = 15.2, s = 4.1
Group 2: n = 30, x̄ = 12.8, s = 3.7
Test H0: μ1 − μ2 = 0 versus Ha: μ1 − μ2 > 0.
- Check conditions.
- Compute the Welch t statistic and approximate df.
- Find the one-sided p-value and make a decision at α = 0.05.
- Interpret the conclusion in context.
Reveal worked answer + practice scoring guide
Worked answer
Assuming independent random samples, the groups are independent; the 10% condition applies if sampled without replacement. With n around 30 in each group, the t procedure is fairly robust unless there is extreme skew or severe outliers.
SE ≈ 1.028, t ≈ 2.335, Welch df ≈ 54.4, and the one-sided p-value ≈ 0.0116. Reject H0 at α = 0.05. The sample provides convincing evidence that Group 1’s population mean exceeds Group 2’s.
10-point practice rubric
2 pts conditions · 3 pts SE/t/df · 2 pts p-value/decision · 3 pts contextual conclusion
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Difference in sample proportions under a model
Suppose two populations have true proportions p1 = 0.60 and p2 = 0.50. Independent samples of 100 are selected from each population.
- State the mean and standard deviation of p̂1 − p̂2.
- Check the Normal-approximation counts.
- Approximate P(p̂1 − p̂2 > 0.20).
- Explain how random sampling affects generalization and why this probability calculation does not imply a causal relationship between population membership and the outcome.
Reveal worked answer + practice scoring guide
Worked answer
The sampling-distribution mean is p1 − p2 = 0.10. The SD is √[0.60(0.40)/100 + 0.50(0.50)/100] = 0.070. All four expected success/failure counts are at least 40, so the Normal approximation is strong.
For a difference of 0.20, z = (0.20−0.10)/0.07 ≈ 1.429, giving tail probability ≈ 0.0766. Random sampling can support generalization to the sampled populations. The calculation compares population proportions; it does not create random assignment or prove that membership in a population causes the outcome.
10-point practice rubric
2 pts center/SD · 2 pts counts · 3 pts probability · 3 pts sampling-vs-causation interpretation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Practice Set 12
Cluster/stratified distinctions, binomial reasoning, paired t inference, and a regression capstone
Cluster versus stratified sampling in a district
A district has 24 schools and wants a student survey. An analyst suggests either (A) sample students from every school or (B) randomly choose six schools and survey randomly selected students within those six.
- Which plan is closer to stratified sampling and why?
- Which plan is closer to multistage cluster sampling and why?
- Give one circumstance in which plan B is cheaper but may have larger sampling variability.
- Explain why randomly choosing six schools is essential if the goal is district-wide inference.
Reveal worked answer + practice scoring guide
Worked answer
Plan A is stratified by school because every school contributes sampled students. Plan B is multistage/cluster-like because a random subset of schools is selected first and students are then sampled within selected schools.
Plan B can reduce travel and administration cost, but students within the same school may be more alike, so fewer schools can reduce effective information and increase sampling variability. Choosing convenient schools instead of random schools can systematically overrepresent certain neighborhoods, programs, or school types and undermine district-wide inference.
10-point practice rubric
2 pts identify stratified design · 2 pts identify multistage/cluster design · 3 pts cost/variance explanation · 3 pts random-school selection for inference
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Binomial tail and expected count
For 10 independent trials with constant success probability p = 0.35, let X be the number of successes.
- Find P(X ≥ 5).
- Find E(X) and SD(X).
- Explain why X is binomial but “the trial number of the first success” would be a different random variable.
- State why the latter topic should not be treated as required 2027 AP Statistics core practice.
Reveal worked answer + practice scoring guide
Worked answer
For X ~ Binomial(10, 0.35), P(X ≥ 5) ≈ 0.2485. E(X) = np = 3.5 and SD(X) = √[10(0.35)(0.65)] ≈ 1.508.
X counts successes in a fixed number of trials. A “trials until first success” variable has a different support and model structure. The revised 2026–27 AP Statistics course removed the geometric distribution from required content, so 2027 core practice should emphasize current probability and binomial reasoning rather than legacy geometric-distribution questions.
10-point practice rubric
3 pts binomial tail · 2 pts mean/SD · 2 pts random-variable distinction · 3 pts correct 2027 scope
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Paired t test and interval for a reduction
For 16 matched subjects, define d = after − before. The sample has d̄ = −2.8 and sd = 3.6. Test H0: μd = 0 versus Ha: μd < 0.
- Interpret μd under the chosen sign convention and check conditions.
- Calculate t, df, and the one-sided p-value.
- Make a decision at α = 0.05.
- Construct a 95% CI for μd and interpret the negative interval.
Reveal worked answer + practice scoring guide
Worked answer
μd is the population mean change after − before; negative values mean the response decreased. The matched pairs should be independent of other pairs and the difference distribution should have no severe skew/outliers for n = 16.
t = −2.8/(3.6/√16) ≈ −3.111, df = 15, and the one-sided p-value ≈ 0.0036. Reject H0. The 95% CI is approximately (−4.72, −0.88). The entire interval is below zero, supporting a population mean decrease of roughly 0.9 to 4.7 units under this difference definition.
10-point practice rubric
2 pts parameter/conditions · 3 pts t/df/p · 2 pts decision · 3 pts CI interpretation
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.Regression capstone with observational scope
Five days of a pilot study record exposure level x and response y:
x: 1, 2, 3, 4, 5 | y: 6, 9, 13, 15, 18
A separate random sample of 80 similar days finds 52 classified as “high response.”
- Find the least-squares line, r, and r² for the five paired observations.
- Predict y at x = 3.5 and find the residual for the observation (4, 15).
- Estimate the population proportion of “high response” days from the separate sample.
- Explain what can and cannot be concluded when the regression data are observational but the 80-day sample is random.
Reveal worked answer + practice scoring guide
Worked answer
The regression line is ŷ = 3.2 + 3.0x, r ≈ 0.996, and r² ≈ 0.991. At x = 3.5, ŷ = 13.7. At x = 4, ŷ = 15.2, so the residual is −0.2.
The separate sample proportion is 52/80 = 0.650. The observational regression supports a strong linear association within the observed x-range but not a causal effect of exposure. If the 80 days were randomly sampled from a well-defined population of days, 0.65 is a defensible estimator for that population proportion, subject to sampling error; random sampling still does not create causation.
10-point practice rubric
3 pts line/r/r² · 2 pts prediction/residual · 2 pts proportion estimate · 3 pts association/generalization/causation distinctions
This is a Salar Cafe practice rubric for diagnostic use, not an official College Board scoring guideline.How to write these FRQs like Bluebook responses
Beginning with the 2027 exam, there is no standard paper FRQ booklet. The strongest practice therefore trains concise typed statistical arguments rather than beautifully typeset notation.
Name the target
Identify the population parameter, variable, group order, or statistical question before calculation.
Show decision evidence
State the method or model and verify relevant conditions with numbers or design evidence instead of listing condition names mechanically.
Calculate with labels
Use calculator output when appropriate, but label the statistic, interval, residual, probability, expected count, or p-value.
Interpret in context
Name the population or variables, use the direction asked for, and separate association, generalization, and causation.
Eight FRQ mistakes that cost reasoning points
A number is not a statistical argument. Identify the quantity and why the method applies.
Use reject or fail to reject, then state the evidence conclusion about the population parameter.
A p-value is not P(H0 is true). It is a tail probability under the null model.
Random sampling supports generalization; random assignment supports causal inference.
For chi-square conditions, inspect expected cell counts from the null model.
Population, variable, direction, units, and group order should survive into the conclusion.
Do not spend core 2027 practice time on geometric distribution, goodness-of-fit, or slope inference.
Long prose does not earn credit if it never resolves the requested statistical decision.
Choose the next AP Statistics practice task
Use the focused resource that matches what you need next: a full timed exam, Section I MCQs, score planning, calculator workflows, or topic-by-topic practice.
AP Statistics FRQ Practice 2027 FAQs
How many FRQs are on the 2027 AP Statistics Exam?
There are 4 free-response questions in 90 minutes. Each is worth 10 raw points, and the FRQ section contributes 50% of the exam score.
What are the four 2027 FRQ types?
College Board describes one question focused on Practices 1–2, one on Practices 3–4, one inference question using Practices 2–4, and one question using Practices 2–4 across multiple content areas.
Are the 48 FRQs on this page released College Board questions?
No. They are original Salar Cafe practice questions written to the revised framework. That avoids copying secure or released wording and allows all 12 sets to follow the new four-question architecture.
Should I still use old released AP Statistics FRQs?
Yes, selectively. Older questions remain valuable for individual statistical ideas, but they use the old six-question architecture and can contain content removed from the 2027 required course. Map them to current skills rather than treating an old full section as a 2027 simulation.
Can I use a calculator on the FRQ section in 2027?
Yes. AP Statistics permits an approved handheld graphing calculator with statistical capability, and Bluebook includes Desmos. Statistical reasoning, setup, and interpretation still need to be communicated rather than replaced by calculator syntax.
Will I type my FRQ answers in 2027?
Yes. Beginning with the May 2027 administration, all AP Statistics free-response answers are submitted in Bluebook. Scratch paper remains available for planning and calculations.
Is chi-square still on AP Statistics in 2027?
Chi-square goodness-of-fit was removed from required content. Chi-square reasoning for homogeneity and independence remains relevant, which is why the practice bank includes those current forms but not goodness-of-fit.
How should I use the 12 sets?
Start with one untimed set to learn the four roles. Then use 90-minute full-set rehearsals. After each set, self-score every question, classify each lost point as design, method choice, conditions, calculation, or interpretation, and target the weakest category before taking another complete set.
Four questions. Forty raw points. Statistical reasoning in every response.
The best preparation is not memorizing six old FRQ templates. Practice the revised four-question roles, type complete statistical arguments, and diagnose exactly where each lost point comes from.
Track whether your 12 practice sets are actually improving your writing
Your self-scores are stored locally in this browser. The dashboard below summarizes fully scored sets and identifies the weakest of the four 2027 FRQ roles. It does not upload responses or require an account.
QWERTY notation habits for fully digital FRQs
p-hat or p̂Define the sample proportion in words if symbol entry slows you down.x-bar or x̄Plain-language notation is acceptable when the meaning is unambiguous.mu / μ, sigma / σUse the symbols menu or conventional keyboard wording.<=, >=, not equalPrioritize clear statistical meaning over decorative notation.