This guide explains the revised AP Statistics course and May 2027 exam using the current College Board framework. It distinguishes verified facts from the still-unpublished 2027 subject date, maps all five units, and shows how statistical practices become exam questions.
What is the AP Statistics Exam?
AP Statistics is College Board’s end-of-course assessment for a one-semester, introductory, non-calculus-based college statistics course. Beginning with the May 2027 administration, the revised exam is fully digital in Bluebook. It has 42 multiple-choice questions in 90 minutes and four 10-point free-response questions in 90 minutes. Each section contributes 50% of the composite score.
The revised course took effect in the 2026-27 school year. It organizes required content into five units and develops four statistical practices: formulate questions, collect data, analyze data, and interpret results. Calculators with statistical capabilities are expected, built-in Desmos is available, and students receive reference information.
The phrase “AP Statistics exam” can refer to the assessment, but it should not be confused with a school semester final or an unofficial online practice test. The operational AP Exam is created and scored through the AP Program and produces a score from 1 to 5. Individual colleges decide whether a particular score earns credit or placement.
2027 AP Statistics exam snapshot
| Feature | Revised exam | What it means for preparation |
|---|---|---|
| Delivery | Fully digital in Bluebook | Practice reading, calculating, and typing complete statistical responses in one workflow |
| Total testing time | 3 hours | Two equal 90-minute sections; administrative time is separate |
| Multiple choice | 42 questions, four choices each | Two three-question sets share prompts; all five units and four practices may appear |
| Free response | 4 questions, 10 points each | Responses emphasize design, analysis, inference, interpretation, and multi-topic reasoning |
| Weight | 50% MCQ and 50% FRQ | A balanced study plan must include both recognition and written justification |
| Calculator | Statistical calculator expected; Desmos built in | Know one dependable workflow and bring an approved handheld if preferred |
| Reference | Printed and digital reference information | Learn method selection and notation; do not depend on formula lookup alone |
| Specific 2027 date | Not yet shown on the public AP Statistics calendar as checked July 18, 2026 | Verify the official calendar and school instructions instead of shifting the 2026 date forward |
What changed from the previous AP Statistics exam
| Feature | Previous framework/exam | Revised 2026-27 framework and May 2027 exam |
|---|---|---|
| Course organization | Nine commonly taught units | Five units |
| Prerequisite | Second-year algebra recommendation | First-year algebra recommendation |
| MCQ count | 40 | 42 |
| Choices per MCQ | Five | Four |
| Shared-prompt MCQs | Not described as the revised two-set structure | Two sets of three questions, one probability and one regression |
| FRQ count | Six | Four |
| Points per FRQ | Four in the older design | Ten |
| Delivery | Hybrid digital in 2026, with paper FRQ booklets | Fully digital in Bluebook in May 2027 |
| Removed required topics | Included in older framework | Geometric distribution, combining random variables, chi-square goodness-of-fit, regression-slope inference, and named departures-from-linearity topic removed |
The revision does not make statistical reasoning optional. It changes the organization and delivery while keeping evidence-based decisions at the center. A student still needs to connect design to inference, conditions to methods, calculations to context, and association to the limits of causal claims.
The five revised AP Statistics units
The official percentage ranges describe exam weighting, not a rigid question-by-question allocation. A question can draw on content from one unit while assessing several practices. The following guide explains what each topic contributes to a complete statistical investigation.
Unit 1: Exploring One-Variable Data and Collecting Data (20%-30% of the score)
This official weighting is a range, not a promise that every form will devote exactly the midpoint to the unit. Questions can also combine content with several statistical practices.
Investigative questions
A statistical investigation begins with a question that anticipates variability and identifies the population, variables, and comparison of interest. A complete question makes the cases and variable roles visible before any graph or procedure is chosen.
Common failure: A vague request to ‘analyze the data’ is not an investigative question because it provides no target or scope. Useful application: Turn ‘Are school lunches good?’ into a measurable question about a defined population and response.
Categorical variables
The denominator must match the population or conditional group named in the claim. Categorical data place cases into groups, so counts, proportions, bar charts, and two-way tables carry the analysis.
Compare proportions rather than raw counts when group sizes differ. A frequent mistake is the opposite approach: a category code such as 1, 2, and 3 is not quantitative merely because numerals are used.
Quantitative distributions
At the heart of this topic, a quantitative distribution is described through shape, center, spread, and unusual features, always with units and context. Graphs and numerical summaries should agree: a skewed histogram should not be summarized as if it were symmetric.
Use it: Use median and IQR for a skewed distribution and mean and SD for a roughly symmetric distribution without outliers. Keep in mind that listing a mean and standard deviation without describing shape or outliers leaves the statistical description incomplete.
Graphical representations
Dotplots, stemplots, histograms, boxplots, and cumulative displays answer different questions about distribution structure. In a complete response, bin width, scale, and grouping can change what a reader sees, so graph construction is part of the analysis.
A boxplot does not display every mode or gap, while a histogram does not show individual values exactly. A better analysis follows this practical rule: choose a common scale when comparing groups so visual differences are not created by the axes.
Summary statistics
Mean, median, quartiles, IQR, range, variance, and standard deviation quantify different aspects of a distribution. A summary should be interpreted as a property of the observed distribution, not merely reported as calculator output.
Common failure: Standard deviation is not the average raw value and cannot be negative; it measures typical distance from the mean in the variable’s units. Useful application: Explain how adding a constant changes location but not spread, while multiplication changes both.
Random sampling
A simple random sample gives every sample of the stated size an equal chance; stratified, cluster, and systematic designs use different random mechanisms. Random sampling uses chance to select cases and supports generalization to the population represented by the sampling frame.
Name the sampling frame and identify who cannot be selected before claiming representativeness. A frequent mistake is the opposite approach: large convenience samples can remain biased because size does not repair systematic undercoverage or self-selection.
Experimental design
At the heart of this topic, an experiment imposes treatments and uses comparison, random assignment, control, and replication to isolate treatment effects. Blocking can reduce unexplained variability by comparing treatments within groups of similar experimental units.
Use it: Use matched pairs when each participant can receive both conditions or when close pairs can be formed. Keep in mind that random assignment and random sampling serve different purposes: causation versus population generalization.
Bias and data quality
Undercoverage, nonresponse, response bias, wording, measurement error, and implementation failure can distort results. In a complete response, a defensible report describes how data were obtained and which error sources remain plausible.
A small margin of error describes sampling variability under a model; it does not certify that the survey is unbiased. A better analysis follows this practical rule: separate the sampling-error calculation from a design audit of who was reached, who responded, and how variables were measured.
Unit 2: Probability, Random Variables, and Probability Distributions (15%-25% of the score)
This official weighting is a range, not a promise that every form will devote exactly the midpoint to the unit. Questions can also combine content with several statistical practices.
Simulation
Simulation represents chance behavior with an explicit random device, a correct outcome assignment, a stopping rule, and many repetitions. The simulated statistic must match the probability or long-run quantity in the question.
Common failure: Running many trials cannot rescue an invalid digit assignment or a model that ignores dependence. Useful application: For a 23% event, assign exactly 23 of 100 equally likely two-digit labels to success.
Probability rules
Use the general addition rule unless events are disjoint, and multiply conditional probabilities along a path. Complements, unions, intersections, and conditional probabilities describe events without double-counting outcomes.
A two-way table or tree diagram can expose the correct denominator before a formula is used. A frequent mistake is the opposite approach: mutually exclusive events with positive probability are not independent because occurrence of one prevents the other.
Conditional probability
At the heart of this topic, a conditional probability changes the reference group to cases for which the condition is true. The phrase ‘among those who’ signals that the denominator must be restricted.
Use it: State the condition in words with the numerical result so the denominator remains visible. Keep in mind that reversing P(A|B) and P(B|A) is a serious interpretation error, especially in screening or diagnostic contexts.
Independence
Events are independent when learning that one occurred does not change the probability of the other. In a complete response, check whether P(A|B)=P(A) or equivalently whether P(A and B)=P(A)P(B), when probabilities are nonzero.
A small observed difference is not proof of exact population independence; it may reflect sample variability. A better analysis follows this practical rule: distinguish an assumed independent-trials model from evidence about association in observed data.
Discrete random variables
A discrete random variable assigns numerical values to chance outcomes and has a probability distribution summing to one. Expected value is the long-run average of the variable, while standard deviation describes trial-to-trial spread.
Common failure: An expected value need not be a possible single outcome; it is a weighted mean across repetitions. Useful application: For a net-gain variable, subtract the cost before computing the expected value.
Binomial distributions
The model supports exact probabilities as well as mean np and standard deviation sqrt(np(1-p)). A binomial model requires a fixed number of trials, two outcomes per trial, independent trials, and constant success probability.
Use a complement for ‘at least one’ or ‘at least two’ when it shortens the calculation. A frequent mistake is the opposite approach: sampling without replacement from a small population may violate the independence requirement unless the sample fraction is small.
Normal distributions
At the heart of this topic, a normal model is symmetric and determined by its mean and standard deviation; standardized areas are found with z-scores. Probability is area under the density curve, and percentiles are obtained by reversing the standardization step.
Use it: Translate the final standardized result back to the original variable and units. Keep in mind that the empirical rule is approximate and applies only when a normal model is reasonable.
Linear transformations
For Y=a+bX, the mean changes to a+b mu_X and the standard deviation changes to |b| sigma_X. In a complete response, the constant shifts the distribution while the multiplier changes scale and may reverse order.
A negative multiplier never creates a negative standard deviation because spread is a magnitude. A better analysis follows this practical rule: unit conversions are a practical example: location and spread transform differently.
Unit 3: Inference for Categorical Data: Proportions (15%-25% of the score)
This official weighting is a range, not a promise that every form will devote exactly the midpoint to the unit. Questions can also combine content with several statistical practices.
Estimators and sampling variability
A sample proportion estimates a population proportion and varies from sample to sample. Under random independent sampling, p-hat is unbiased with standard deviation sqrt[p(1-p)/n].
Common failure: The observed p-hat is not the population parameter and should not replace p in a null standard error without justification. Useful application: Increasing n reduces standard error at the square-root rate rather than in direct proportion.
Sampling distribution of p-hat
Check np and n(1-p) for probability calculations; use sample counts for an interval and null counts for a test. The sampling distribution is centered at p and is approximately normal when expected successes and failures are sufficiently large.
Keep population proportion, sample proportion, and null value distinct in notation and language. A frequent mistake is the opposite approach: a sample size of 30 alone is not the condition for a proportion.
One-proportion confidence interval
At the heart of this topic, a one-proportion z interval combines p-hat with a margin based on its estimated standard error. Interpret the interval as a plausible range for the population proportion and confidence as long-run method performance.
Use it: Report the population and trait, not merely two decimal endpoints. Keep in mind that do not say that a fixed parameter has a post-data probability of lying in an already calculated frequentist interval.
Planning sample size
Required sample size follows from the requested confidence level, margin of error, and planning value for p. In a complete response, use p*=0.5 when no credible estimate exists because it maximizes p(1-p), then round upward.
Rounding down can violate the promised maximum margin of error. A better analysis follows this practical rule: precision planning does not address bias from a poor sampling frame or nonresponse.
One-proportion significance test
A one-proportion z test compares p-hat with a null value p0 using a standard error calculated under H0. Hypotheses refer to the population parameter and the alternative direction must match the research question.
Common failure: A p-value is not the probability the null is true; it is a tail probability computed assuming the null model. Useful application: State the decision at the chosen alpha and then give a contextual evidence statement.
Two-proportion confidence interval
Independent samples, large counts in both groups, and appropriate randomization or sampling support the method. The interval for p1-p2 uses separate sample estimates in the standard error because it estimates an unrestricted difference.
Define the subtraction order before interpreting positive and negative endpoints. A frequent mistake is the opposite approach: including zero means no difference remains plausible; it does not prove exact equality.
Two-proportion significance test
At the heart of this topic, under H0:p1=p2, successes are pooled to estimate the common null proportion. The pooled estimate is used only for the test’s null standard error, not for the ordinary two-proportion confidence interval.
Use it: Link the final conclusion to the design before claiming causation or broad generalization. Keep in mind that using an unpooled interval standard error inside the test mixes two different models.
Chi-square association
A chi-square procedure compares observed counts with counts expected under independence in a two-way table. In a complete response, expected counts come from row total times column total divided by the grand total, and degrees of freedom are (r-1)(c-1).
The test concerns categorical association, not linear correlation and not causal direction. A better analysis follows this practical rule: inspect cell residuals after significance to see which combinations contribute to the overall result.
Unit 4: Inference for Quantitative Data: Means (10%-20% of the score)
This official weighting is a range, not a promise that every form will devote exactly the midpoint to the unit. Questions can also combine content with several statistical practices.
Sampling distribution of x-bar
The sample mean is unbiased for mu and has standard deviation sigma/sqrt(n) under independent random sampling. A normal population gives an exactly normal sampling distribution; the central limit theorem gives an approximation for large n.
Common failure: The population’s skew does not disappear from individual observations merely because the sample mean is approximately normal. Useful application: Use the 10 percent condition when sampling without replacement from a finite population.
The t distribution
Degrees of freedom n-1 determine the appropriate t distribution for a one-sample procedure. When population sigma is unknown, standardizing x-bar with the sample SD produces a t statistic with heavier tails than z.
As degrees of freedom increase, t critical values approach the corresponding z values. A frequent mistake is the opposite approach: substituting s for sigma while continuing to use a z critical value understates uncertainty in small samples.
One-mean confidence interval
At the heart of this topic, a one-sample t interval estimates mu using x-bar plus or minus t* times s/sqrt(n). Randomness, independence, and a distribution shape suitable for the sample size must be justified.
Use it: Outliers matter strongly for small-sample t procedures because both mean and SD are nonresistant. Keep in mind that a calculator interval without a parameter definition, conditions, and contextual interpretation is incomplete.
One-mean significance test
A one-sample t test compares x-bar with mu0 using the sample SD and n-1 degrees of freedom. In a complete response, the alternative may be one-sided or two-sided, and the p-value follows that prespecified direction.
Failing to reject does not establish the null value as exactly true. A better analysis follows this practical rule: report the size and uncertainty of the estimated difference alongside statistical significance.
Paired data
Paired analysis converts linked measurements into one difference per pair and then applies a one-sample t procedure to those differences. The relevant sample size is the number of pairs, and conditions are checked on the difference distribution.
Common failure: Treating paired measurements as independent discards the within-pair structure and gives the wrong standard error. Useful application: Define the subtraction order so the sign of the mean difference has a clear meaning.
Two independent means
Independent groups, randomization or sampling, and shape checks for both groups support the method. Two-sample t procedures compare mu1-mu2 using separate sample means and standard deviations.
Interpret the interval or test in the defined subtraction order and respect the design’s scope. A frequent mistake is the opposite approach: the AP method does not require population variances to be equal unless a pooled procedure is explicitly chosen.
Power and errors
At the heart of this topic, type I error, Type II error, and power translate the consequences of decisions into the study’s context. Larger samples, larger true effects, less variability, and a larger alpha generally increase power.
Use it: Choose alpha based on consequences and pair significance with effect-size judgment. Keep in mind that changing alpha after inspecting the p-value invalidates the planned error-rate interpretation.
Practical importance
Statistical significance measures incompatibility with a null model, not whether an effect is large enough to matter. In a complete response, confidence intervals show effect magnitude and precision and should be compared with a contextual benchmark.
A huge sample can make a trivial effect statistically significant, while a small study can miss a meaningful effect. A better analysis follows this practical rule: discuss costs, benefits, units, and realistic decision thresholds after the inference calculation.
Unit 5: Regression Analysis (10%-20% of the score)
This official weighting is a range, not a promise that every form will devote exactly the midpoint to the unit. Questions can also combine content with several statistical practices.
Scatterplots
A scatterplot displays direction, form, strength, and unusual features in the relationship between two quantitative variables. The explanatory variable belongs on the horizontal axis and the response on the vertical axis.
Common failure: Correlation and regression summaries should not be used before checking whether a roughly linear description is sensible. Useful application: Describe clusters, gaps, and influential points instead of reducing the graph to one adjective.
Correlation
It is unchanged by positive linear unit conversions but changes sign if one variable’s direction is reversed. Correlation r measures the direction and strength of a linear association and is unitless between -1 and 1.
A restricted range of x-values can make the observed correlation differ from a broader population relationship. A frequent mistake is the opposite approach: correlation is not resistant, does not measure nonlinear strength well, and does not establish causation.
Least-squares line
At the heart of this topic, the least-squares line chooses intercept and slope to minimize the sum of squared vertical residuals. The slope is the predicted change in response per one-unit increase in the explanatory variable.
Use it: Write prediction units explicitly and limit use to the observed range unless extrapolation is defended. Keep in mind that interpreting the intercept outside a meaningful data range can produce a nonsensical contextual claim.
Residuals
A residual is observed response minus predicted response; its sign shows whether the model under- or overpredicted. In a complete response, residual plots assess whether systematic structure remains after fitting a line.
A pattern, changing spread, or unusual point warns that a simple linear summary is incomplete. A better analysis follows this practical rule: the residuals from a least-squares line with an intercept sum to approximately zero, but that alone does not prove a good fit.
Coefficient of determination
The coefficient r-squared is the proportion of sample variation in the response explained by the linear model with the predictor. Its interpretation names the response variable and the fitted linear relationship.
Common failure: It is not the percentage of points on the line, not a slope, and not a causal percentage. Useful application: Compare r-squared only alongside residual behavior, context, and the range of data.
Influential observations
Extreme x-values create leverage, while a residual measures vertical departure from the current line. A point can be influential when its removal substantially changes slope, intercept, correlation, or conclusions.
Verify the record and report sensitivity analyses rather than deleting a valid observation automatically. A frequent mistake is the opposite approach: a high-leverage point can have a small residual and still control the fitted line.
Prediction and extrapolation
At the heart of this topic, a regression prediction is conditional on an x-value and inherits uncertainty from individual variation and model estimation. Interpolation within the observed x-range is generally more defensible than extrapolation beyond it.
Use it: Round to a level supported by the data and state when the predictor lies outside the observed range. Keep in mind that a precise calculator output does not imply a precise prediction when residual spread is large.
Regression in the revised course
The revised unit retains scatterplots, correlation, least-squares models, residuals, prediction, and interpretation. In a complete response, inference for the population regression slope was removed from the required revised framework.
Older resources can overprepare slope-inference procedures while underpreparing the new statistical practices. A better analysis follows this practical rule: use current College Board unit descriptions to decide what belongs in May 2027 preparation.
Six integrated examples of AP Statistics reasoning
All six cases below are original hypothetical teaching examples. They illustrate methods and do not report real study findings.
Case 1: A district transportation survey
This original example starts with the question: among all students enrolled in the district, what proportion usually travels to school by bus, family vehicle, walking, cycling, or public transit, and how do those proportions differ by grade band? The population is the current district enrollment; the cases are students; travel mode is categorical; grade band is categorical. The question requires variation and a defined population, so it is statistical.
A proportional stratified random sample by school level can ensure representation from elementary, middle, and high school students. Within every stratum, the district should use a random mechanism on a current roster. A link sent to any student who chooses to respond would instead be a voluntary-response sample. Even with thousands of replies, students with unusual commutes or stronger opinions could be overrepresented. The report should record response rates by stratum and follow up consistently.
Because groups have different enrollment, compare conditional proportions rather than raw counts. A two-way table can show travel mode within each grade band. Side-by-side or segmented bar charts make distribution differences visible. If a chi-square test of association is later used, the hypotheses concern population association between grade band and travel mode, expected counts are calculated from marginal totals, and the conclusion remains noncausal. Grade is not randomly assigned and travel policies may confound the association.
The scope of inference comes from the design. Random sampling can support generalization to students represented by the roster if nonresponse and measurement problems are controlled. The survey cannot establish that changing grade level causes a different commute. A complete public report presents sample sizes and denominators, describes the random selection, explains missing responses, and avoids claiming that a narrow margin of error eliminates all bias.
Case 2: A probability model for support requests
Suppose a help desk classifies an incoming request as account access, payment, document upload, or other. To model the next request, the categories need probabilities that are nonnegative and sum to one. These probabilities could come from a defined historical period, but they should not be presented as universal rates. The random variable might be handling time in minutes, while the category itself remains categorical.
A simulation should copy the intended model. If account-access requests have probability 0.28, exactly 28 of 100 equally likely two-digit labels can represent that category. One random pair is generated per request, the assigned category is recorded, and the process repeats for the stated number of requests. If the real requests cluster by outage or time of day, an independent-trials simulation may be unrealistic. More repetitions reduce Monte Carlo noise but do not repair a false independence assumption.
Conditional probability answers questions such as the probability a request is resolved on the first contact among payment requests. Its denominator is payment requests, not all requests. Reversing the condition would answer a different question. A tree can multiply probabilities along paths and add disjoint paths, while a two-way table can compute the same quantities from counts. Agreement between the two representations is a useful arithmetic check.
If each of 20 independently selected requests has a constant 0.75 probability of first-contact resolution, a binomial model is possible. The mean number resolved is np and the standard deviation is the square root of np(1-p). The model is not justified merely because the response is yes/no; fixed trial count, independence, and constant probability are also required. The expected value describes a long-run average and need not be an integer outcome in a single batch.
Case 3: Estimating and testing proportions
Consider an original poll of a randomly selected set of students about access to a quiet study space. The target parameter p is the true proportion in the population defined by the roster and survey period. The observed sample proportion p-hat is an estimate, not the parameter itself. Random sampling, the 10 percent condition for sampling without replacement, and sufficiently large success and failure counts support a one-proportion z interval.
The interval combines the estimate with z-star times the estimated standard error. Its interpretation names the population and trait: the method yields a range of plausible values for the population proportion. A 95% confidence statement describes the long-run capture rate of intervals built this way. It does not mean that 95% of students are inside the interval, and it does not assign a changing probability to a fixed population parameter after the endpoints are calculated.
A test answers a different question. If a prior policy claim says p=0.60, the null standard error uses p0=0.60 because the p-value is calculated under that null model. The alternative direction must be chosen from the question before seeing the data. A small p-value says the observed result, or one more extreme in the alternative direction, would be unusual if the null were true. It is not the probability the policy claim is true.
If two independently sampled schools are compared, define p1-p2 before interpreting a sign. A confidence interval uses separate sample estimates in its standard error. A significance test of equality pools the successes because the null treats the two population proportions as one common value. That pooled-versus-unpooled distinction is a frequent method-selection test. Whether the conclusion generalizes or supports causation still depends on sampling and assignment.
Case 4: Comparing mean completion times
An original usability study can compare how long the same participants take to complete a form before and after a redesign. Because every person supplies two measurements, the analysis is paired. Define each difference, for example new time minus old time, and analyze the one-variable distribution of differences. The sign then has a stable meaning: a negative difference represents faster completion with the redesign.
The parameter is the population mean paired difference, not the difference between two unrelated sample means. A one-sample t interval or test on differences uses the number of pairs as n, the mean difference as x-bar, and the SD of differences as s. Randomness and independence concern the pairs, and the normal/outlier check concerns the difference distribution. Treating all before times as one independent group and all after times as another discards the pairing.
If participants were randomly sampled from a well-defined user population, the result may generalize to that population. If the order of old and new forms was randomized or counterbalanced, practice and fatigue effects are reduced. If every participant necessarily uses the old version first, an observed improvement can be confounded with familiarity. Random assignment of order addresses that threat; a larger sample alone does not.
After computing, interpretation must include magnitude. An interval entirely below zero supports a mean reduction under the stated subtraction order, but the endpoints should also be compared with a practical benchmark. A statistically clear reduction of a fraction of a second may not justify implementation cost, while a modest study can have an interval too wide to rule out an important benefit. Statistical and practical importance are separate judgments.
Case 5: Regression for prediction
Imagine an original data set relating the number of timed practice sets completed to final diagnostic score. The explanatory variable is set count and belongs on the horizontal axis; score is the response. A scatterplot should be examined for direction, form, strength, clusters, gaps, and unusual observations before a linear summary is interpreted. A curved pattern cannot be rescued by reporting a large correlation alone.
The least-squares slope is the predicted score change associated with one additional practice set, within the observed range. The word associated matters because students were not randomly assigned to set counts. Motivation, prior preparation, course attendance, and access to study time may confound the relationship. The intercept predicts score when set count is zero, but that interpretation is useful only if zero is in or near the data range and the linear model remains sensible there.
A residual equals observed minus predicted score. Positive residuals indicate scores above the fitted prediction; negative residuals indicate scores below it. A residual plot with no systematic pattern supports the linear form, while curvature or changing spread signals missing structure. The residual standard deviation describes typical vertical prediction error in score units. It is not the standard deviation of the explanatory variable.
R-squared is the proportion of sample response variation explained by the fitted linear relationship. It is not the percentage of students whose scores are correct and it does not establish causation. A point at an extreme set count can have high leverage and may be influential even if it lies close to the fitted line. Verify such a record and compare fits with and without it instead of deleting it automatically.
Case 6: A complete cross-unit investigation
A school wants to know whether a structured review schedule improves mastery. A defensible question names the eligible student population, the schedule being compared, a prespecified mastery measure, and the time at which mastery is measured. Students can be blocked by prior diagnostic performance and randomly assigned within blocks to the structured or existing schedule. Equal total study time and common content prevent the schedule comparison from being mixed with dosage.
During the study, the response could be a quantitative diagnostic score, while completion and attrition indicators are categorical. Quantitative distributions should be graphed by treatment and summarized with center, spread, and unusual features. A treatment comparison can use a two-sample t method if its design and distribution conditions hold. Attrition proportions can be compared separately, because unequal dropout can change the composition of observed groups.
Random assignment supports a causal conclusion for participants under the implemented conditions. Broad generalization depends on how participants and sites were selected. If all volunteers come from one advanced class, the result does not automatically apply to every AP Statistics student. Blocking improves precision but does not substitute for random sampling. Reporting both design features prevents causation and generalization from being conflated.
The final argument combines practices. Formulate the question and parameter; collect data through a controlled random assignment; analyze distributions and the treatment effect; interpret an interval or p-value in context; discuss practical magnitude; and state limitations. That connected reasoning is closer to the revised exam than memorizing isolated formulas. The calculator performs arithmetic, while the student supplies the design logic and evidence boundaries.
The four statistical practices
Practice 1: Formulate Questions. Students identify a statistical question, define the population and variables, and connect the requested conclusion to an answerable design. A good question anticipates variation. It is not enough to name a topic; the question must identify what will be measured and compared.
Practice 2: Collect Data. Students select or evaluate a sampling plan, experiment, or other data-producing process. This includes random mechanisms, controls, potential bias, and ethical or practical limitations. The data-collection step determines which population and causal claims remain defensible later.
Practice 3: Analyze Data. Students choose graphical, numerical, probabilistic, simulation, interval, test, or regression methods that match the variables and design. Technology can perform arithmetic, but the student must know why the procedure applies and whether its conditions hold.
Practice 4: Interpret Results. Students translate results into the original context, quantify uncertainty, state evidence at the right strength, and avoid causal or population claims the design cannot support. Interpretation includes recognizing practical importance and limitations, not merely repeating a p-value.
| FRQ position | Primary emphasis | A strong response demonstrates |
|---|---|---|
| Question 1 | Practices 1 and 2 | A precise question, population/variable definitions, and a justified data-collection plan |
| Question 2 | Practices 3 and 4 | Correct analysis and an interpretation tied to evidence |
| Question 3 | Inference through Practices 2, 3, and 4 | Conditions, procedure, calculation, and conclusion in context |
| Question 4 | Multiple content areas through Practices 2, 3, and 4 | A connected argument across several parts of the course |
How the exam measures statistical learning
The multiple-choice section tests more than vocabulary. Questions can ask students to read a graph, select a design, compute a probability, identify a parameter, diagnose an invalid inference, or interpret calculator output. The two shared-prompt sets reward careful reuse of a common context without treating each part as unrelated.
The free-response section rewards visible reasoning. A statistically correct calculation can earn less than a complete response if the parameter is undefined, conditions are absent, or the conclusion fails to name the population and direction. Conversely, lengthy prose cannot replace a correct method. The best responses are compact, explicit, and tied to the design.
Because the exam is fully digital, students should practice switching among the Bluebook prompt, scratch work, a calculator, the reference information, and the typed answer field. The interface changes the mechanics, not the standard of statistical reasoning. A complete typed response still needs notation that a scorer can follow and prose that says what the number means.
The built-in symbols menu and digital-response guidance reduce the need for specialized typesetting. Students can write conventional forms such as p-hat, x-bar, mu, H0, and Ha when used consistently and defined. The goal is unambiguous statistical communication, not decorative notation.
Scoring, AP scores, and college credit
The MCQ and FRQ sections each contribute half of the composite score. College Board combines component results and reports an AP score from 1 to 5. Operational score setting is not a simple permanent percentage table, so a website should not publish an invented 2027 raw-score cutoff.
A practice score is most useful when it diagnoses a skill. Separate mistakes into content, method selection, conditions, calculator entry, algebra, and interpretation. Two students with the same total may need completely different study plans.
| AP score | College Board recommendation language | Credit reality |
|---|---|---|
| 5 | Extremely well qualified | Often considered for credit or advanced placement, but the institution controls the policy |
| 4 | Very well qualified | Policies vary by college, department, and major |
| 3 | Qualified | Some institutions award credit or placement; others require a higher score |
| 2 | Possibly qualified | Credit is uncommon, but the score still documents exam performance |
| 1 | No recommendation | Does not normally earn credit |
Before registering for a later course, check the current credit policy for the exact college and program. A general university policy can differ from a major-specific prerequisite. Students should also ask whether credit, placement, or both are granted.
A preparation sequence that matches the exam
First, map the course. Use the revised five-unit description, not an old nine-unit checklist, and identify which topics have been removed. Mark each topic as secure, developing, or not yet learned based on actual problem performance.
Second, build method selection. Before calculating, write the variable types, parameter, design, and requested conclusion. This short routine prevents choosing a proportion procedure for a mean, treating paired data as independent, or using a test standard error in an interval.
Third, integrate technology. Practice lists, one- and two-variable summaries, normal probabilities, binomial calculations, intervals, tests, chi-square tables, and regression on the calculator that will be used. Repeat essential work in Desmos so the built-in option is familiar.
Fourth, write complete conclusions. Use the population and parameter, the alternative direction, the p-value or interval evidence, and the design’s limits. Remove claims that a p-value is the probability of the null or that observational association proves causation.
Fifth, rehearse the clock. Complete a 42-question set in 90 minutes and a four-question FRQ set in 90 minutes. Review every error afterward. Timing without review measures speed; review without timing may hide pacing problems.
| Study evidence | What it reveals | Response |
|---|---|---|
| Repeated errors on the same method | Concept or selection gap | Return to a small set of contrasting examples and explain why each method differs |
| Correct method, wrong arithmetic | Calculator or algebra gap | Record exact keystrokes and estimate the expected magnitude before accepting output |
| Correct numbers, weak conclusion | Interpretation gap | Use a sentence frame naming parameter, population, direction, and uncertainty |
| Strong untimed work, incomplete timed sets | Pacing gap | Practice checkpoints and skip-return decisions under a visible timer |
| Strong old-format practice only | Format gap | Use revised 42/4 Bluebook-aligned sets and typed responses |
Fourteen-week AP Statistics study plan
Count backward from the official date once College Board and the school publish it. If fewer than fourteen weeks remain, combine adjacent rows while preserving the order: diagnose, rebuild, interleave, time, and repair.
| Checkpoint | Focus | Work to complete | Evidence before moving on |
|---|---|---|---|
| 14 weeks out | Baseline and revised map | Complete a mixed untimed diagnostic; tag every miss by the five revised units and four practices. | A written list of secure, developing, and missing skills; old removed topics are separated from current requirements. |
| 13 weeks out | Unit 1 distributions | Rebuild graph choice, shape-center-spread descriptions, transformations, z-scores, and resistant-summary decisions. | Two comparative descriptions that name variables, groups, units, and unusual features without relying on calculator output alone. |
| 12 weeks out | Unit 1 data collection | Contrast SRS, stratified, cluster, and systematic samples; then design randomized and blocked experiments. | For ten scenarios, correctly state whether population generalization, causation, both, or neither is supported. |
| 11 weeks out | Unit 2 probability | Use tables and trees for unions, intersections, complements, conditions, and independence; simulate at least two chance processes. | Every simulation states an outcome assignment, repetition rule, statistic, and interpretation of the estimated probability. |
| 10 weeks out | Unit 2 random variables | Practice expected value, standard deviation, binomial conditions, binomial probabilities, and normal-model areas and percentiles. | Calculator results are paired with a sketch, formula or parameter definition, and an answer in the original units. |
| 9 weeks out | Unit 3 proportion intervals | Build one- and two-proportion intervals, sample-size plans, condition checks, and long-run confidence interpretations. | No interval conclusion assigns probability to a fixed parameter or claims that margin of error includes nonresponse bias. |
| 8 weeks out | Unit 3 proportion tests | Write hypotheses, distinguish pooled test SE from unpooled interval SE, and connect p-values to contextual evidence. | Each response defines the population parameter before H0 and Ha and uses the design to limit its final claim. |
| 7 weeks out | Chi-square association | Compute expected counts, contributions, and degrees of freedom; interpret significance and inspect signed residual direction. | A complete conclusion says associated or not enough evidence of association, never correlated or caused solely from the test. |
| 6 weeks out | Unit 4 one-mean and paired t | Review sampling distributions of means, t distributions, one-sample intervals/tests, and paired differences. | Conditions are checked on the correct data object, especially the difference distribution in paired work. |
| 5 weeks out | Unit 4 two means and power | Compare independent means, contextualize Type I and II errors, and explain how design choices affect power. | Calculations are followed by effect-size and practical-importance discussion rather than significance alone. |
| 4 weeks out | Unit 5 regression | Interpret scatterplots, r, slope, intercept, residuals, r-squared, leverage, influence, and extrapolation. | A residual-plot analysis accompanies every fitted model and observational associations are not described causally. |
| 3 weeks out | Mixed method selection | Complete interleaved sets in which the procedure is not named; write variable type, parameter, and design first. | At least 90% of selected procedures match the parameter and sampling/experimental structure before arithmetic begins. |
| 2 weeks out | Revised digital format | Complete 42 MCQs in 90 minutes and four FRQs in 90 minutes using Bluebook-style typed notation, scratch work, and the intended calculator. | The set is finished within section limits without omitting condition checks or final contextual sentences. |
| 1 week out | Repair, not cramming | Redo missed questions cold, verify Bluebook and calculator readiness, review the supplied reference information, and protect sleep. | The error log shows corrected reasoning in the student's own words and no unresolved device or account problem remains. |
| Final day | Light retrieval and logistics | Review definitions, decision cues, and a few representative solutions; follow the school's arrival and materials instructions. | The student can state the two 90-minute section structures, calculator plan, and skip-return pacing rule without adding new content. |
Method selection map
Procedure selection begins with the variable type, parameter, design, and requested conclusion. The table is a navigation aid, not a substitute for checking the full conditions of a chosen method.
| Question signal | Target | Likely method | What must accompany it |
|---|---|---|---|
| Describe one quantitative variable | Distribution of observed values | Graph plus shape, center, spread, unusual features | Units, skew, clusters, gaps, outliers, and resistant summaries when needed |
| Compare one quantitative variable across groups | Group distributions | Parallel displays and comparative descriptions | Use common scales and comparison words; do not write isolated group summaries |
| Estimate one population proportion | p | One-proportion z interval | Randomness, independence/10%, at least 10 observed successes and failures |
| Test a claim about one proportion | p | One-proportion z test | Use p0 in the null standard error and state a p-value conclusion in context |
| Estimate a difference in independent proportions | p1-p2 | Two-proportion z interval | Separate sample estimates in the unpooled standard error; define subtraction order |
| Test equality of two proportions | p1-p2 | Two-proportion z test | Pool successes only under H0:p1=p2; check all four null counts |
| Study association between two categorical variables | Population association | Chi-square test of independence or homogeneity as design indicates | Use counts, expected-count conditions, (r-1)(c-1) df, and noncausal language |
| Estimate one population mean | mu | One-sample t interval | Random/independent data and shape suitable for n; use s and n-1 df |
| Test a claim about one mean | mu | One-sample t test | Write H0 and Ha about mu, inspect outliers/skew, and interpret the p-value |
| Estimate a mean paired change | mu_d | One-sample t interval on differences | One difference per pair, conditions on differences, and a defined subtraction order |
| Test a paired before-after effect | mu_d | Paired t test | Do not treat the two columns as independent groups; n is the number of complete pairs |
| Estimate a difference in independent means | mu1-mu2 | Two-sample t interval | Independent groups, separate means/SDs, shape checks, and contextual subtraction order |
| Test equality of independent means | mu1-mu2 | Two-sample t test | Match the alternative to the question and connect causation/generalization to design |
| Calculate a binomial probability | Distribution of X, the number of successes | Binomial formula or calculator | Fixed n, two outcomes, independent trials, constant p; use complements strategically |
| Find a normal probability or percentile | A continuous normal model | Standardize with z or use normal CDF/inverse CDF | Shade the requested area, use the correct tail, and return to original units |
| Predict a quantitative response | Conditional mean response under a fitted line | Least-squares regression prediction | Check linear form, stay within the observed range, and acknowledge residual variability |
| Interpret regression fit | Slope, intercept, residuals, r, and r-squared | Contextual interpretation plus residual diagnostics | Avoid causal claims, percent-on-the-line claims, and meaningless extrapolated intercepts |
| Evaluate a sample design | Population quantities represented by selected cases | SRS, stratified, cluster, systematic, or multistage reasoning | Name the frame, random mechanism, undercoverage, nonresponse, and realistic scope |
| Evaluate an experiment | Treatment effects for experimental units | Comparison, random assignment, control, replication, and blocking | Separate assignment from sampling and identify uncontrolled implementation differences |
| Judge a statistical conclusion | Claim supported by design and uncertainty | Synthesize estimate, interval/test evidence, practical size, and limitations | Do not turn association into causation or statistical significance into practical importance |
Core AP Statistics vocabulary
| Term | Meaning in AP Statistics | How it appears in reasoning | Do not confuse it with |
|---|---|---|---|
| Case | The object described by one row or observational unit in a data set. | Identify what one observation represents before naming variables or independence conditions. | Do not count a row twice merely because it contains several variables. |
| Variable | A recorded characteristic that can differ from case to case. | Classify each variable as categorical or quantitative and assign explanatory or response roles when relevant. | Numeric category labels remain categorical and should not be averaged as measurements. |
| Parameter | A numerical feature of a population, usually unknown and fixed for the defined population. | Define it in context before writing an interval, hypotheses, or a conclusion. | A sample estimate is not the population parameter. |
| Statistic | A numerical feature computed from sample data and used to estimate or test a parameter. | Distinguish the observed estimate from the unknown population target in notation and prose. | A statistic can vary; the defined finite-population parameter does not vary across repeated samples. |
| Sampling variability | The natural change in a statistic across repeated random samples from the same population. | Use it to explain why two random samples from the same population need not give identical answers. | Sampling variability is not the same as bias or measurement error. |
| Bias | A systematic tendency of a method to miss the target in a particular direction. | Audit selection and measurement before treating a narrow confidence interval as trustworthy. | Low variability does not imply low bias. |
| Confounding | A mixing of effects that prevents the separate influence of an explanatory variable from being identified. | Name plausible common causes whenever an observational association is mistaken for a treatment effect. | Confounding is not removed by a larger observational sample alone. |
| Random sample | A sample chosen by a known chance mechanism from a defined sampling frame. | Describe the exact chance mechanism and the population represented by its sampling frame. | Random sampling is not random assignment. |
| Random assignment | Use of chance to allocate experimental units to treatments, supporting causal comparison. | Connect it to comparable treatment groups and the justification for a cause-and-effect conclusion. | Random assignment does not automatically make participants representative of a wider population. |
| Stratum | A subgroup sampled separately in a stratified design, usually chosen for internal similarity or guaranteed representation. | Explain why every subgroup is sampled and how proportional or equal allocation is carried out. | A stratum is sampled from every group; a cluster design selects whole groups. |
| Block | A group of similar experimental units within which treatments are randomly assigned. | Use it to reduce within-group variability before comparing experimental treatments. | Blocking organizes an experiment, while stratification organizes a sample. |
| Distribution | The pattern of values of a variable, including frequencies, shape, center, spread, and unusual features. | Describe shape, center, spread, and unusual features with the variable's units and context. | A list of values is not a full distribution description. |
| Standard deviation | A nonnegative measure of spread around the mean in the variable's original units. | Pair it with the mean for symmetric distributions and explain its sensitivity to outliers. | Standard deviation is not resistant and is not the average observation. |
| Standard error | The standard deviation of an estimator's sampling distribution, estimated when necessary from sample data. | Select the correct formula for a mean, proportion, difference, interval, or test. | Standard error measures estimator variability, not variation among individual observations. |
| z-score | A standardized location measured in standard deviations above or below a mean. | Compare relative standing across scales, then translate the result back into context. | A z-score is not a raw-point or percentage difference. |
| Expected value | The probability-weighted long-run average of a random variable. | Interpret it as a long-run average rather than a guaranteed outcome on one trial. | Expected value need not be a possible outcome in a single play. |
| Binomial model | A fixed-trial model with two outcomes, independent trials, and constant success probability. | Check all four conditions before using its probability, mean, or standard-deviation formulas. | A binary outcome alone does not guarantee a binomial model. |
| Confidence level | The long-run proportion of intervals from a method that capture the target parameter. | Interpret the interval-producing method, not a changing probability assigned to a fixed parameter. | Confidence level is not the proportion of sample observations inside the interval. |
| Margin of error | The critical value multiplied by standard error; half the width of a symmetric confidence interval. | Relate precision to confidence level, variability, and sample size. | Margin of error does not include every source of survey bias. |
| Null hypothesis | A precise population claim used to calculate the reference distribution for a test. | Write equality in the null and calculate the test under that stated model. | Failing to reject the null is not proof that it is exactly true. |
| Alternative hypothesis | The population direction or difference for which the study seeks evidence. | Choose its direction from the research question before inspecting the sample result. | A two-sided alternative is not selected after seeing which direction the sample moved. |
| p-value | A null-model probability of the observed statistic or one more extreme in the alternative's direction. | Use it as an evidence measure under the null and never as the probability that the null is true. | A small p-value does not measure effect size or practical importance. |
| Type I error | Rejecting a null hypothesis that is in fact true. | Translate the false-alarm consequence into the setting before choosing alpha. | The Type I error rate is controlled by the prespecified alpha under the null model. |
| Type II error | Failing to reject a null hypothesis when a specified alternative is true. | Translate the missed-effect consequence and explain its relation to sample size and power. | A Type II error can only be described relative to a particular true alternative. |
| Power | The probability that a test rejects the null for a specified true alternative. | Evaluate the chance of detecting a specified meaningful effect under a planned design. | High power does not certify an unbiased design or a practically important effect. |
| Degrees of freedom | A count that determines the reference distribution after constraints or estimated quantities are considered. | Use n-1 for one-sample t work and (r-1)(c-1) for chi-square association tables. | Degrees of freedom are not always equal to sample size. |
| Expected count | A table cell count predicted by the null model, used in a chi-square calculation. | Check null-model adequacy with expected rather than observed cell counts. | Expected counts are model-based and can be nonintegers. |
| Residual | Observed response minus the response predicted by a fitted model. | Calculate its sign and size, then inspect patterns across the entire residual plot. | Residuals are vertical response differences, not horizontal distances. |
| Leverage | The potential of an observation to affect a regression fit because its explanatory value is far from the center. | Recognize why an extreme explanatory value deserves a sensitivity check even when its residual is small. | Leverage describes predictor location; it is not identical to influence. |
| Influence | The actual change in a fitted analysis or conclusion when an observation is removed. | Compare fitted results with and without a point and report the reason for any major change. | Influence is an empirical change in the fit, not a reason to delete valid data automatically. |
Questions students ask about the AP Statistics Exam
Is AP Statistics a math course?
Yes, but its central work is statistical reasoning rather than long symbolic manipulation. Students calculate, use technology, evaluate study design, and explain what evidence does and does not support.
Is calculus required?
No. College Board describes AP Statistics as a non-calculus-based introductory college course. The revised recommended prerequisite is a first-year algebra course.
How many units are in the revised course?
Five units are used beginning in the 2026-27 school year. Older nine-unit lists describe the previous framework and should not be used as the only 2027 study map.
How long is the exam?
The testing time is three hours: 90 minutes for 42 MCQs and 90 minutes for four FRQs. Administrative time and any scheduled break can make the room commitment longer; see the timing guide.
Is the exam fully digital?
Beginning in May 2027, both sections are completed in Bluebook. The digital testing guide covers device preparation, notation, scratch paper, Desmos, and the handheld-calculator option.
What is the exact 2027 exam date?
As checked July 18, 2026, College Board’s public student calendar still displayed the 2026 subject date rather than a confirmed 2027 AP Statistics date. The date-status page should be checked before making travel or school plans.
How many answer choices does each MCQ have?
The revised multiple-choice questions have four answer choices. This is a change from the older five-choice format.
Are all MCQs independent?
No. The revised section includes two sets of three questions that share a prompt: one set involving probability, random variables, and distributions, and one involving regression analysis.
How are the FRQs organized?
Question 1 emphasizes formulating questions and collecting data; Question 2 emphasizes analyzing data and interpreting results; Question 3 focuses on inference; Question 4 combines multiple content areas.
Is a calculator allowed?
Yes. College Board expects a calculator with statistical capabilities and also supplies built-in Desmos in Bluebook. Check the current official calculator policy before exam day.
Can a student bring two calculators?
The current College Board policy permits up to two approved handheld calculators. Batteries, modes, and stored programs should comply with the policy and local proctor directions.
Is a formula sheet supplied?
Reference information is provided in printed form and digitally in Bluebook for the fully digital exam. The formula sheet guide explains what is supplied and what understanding must still be learned.
Does the formula sheet replace memorization?
It reduces formula recall but does not select a procedure, check conditions, define a parameter, or write a contextual conclusion. Those decisions are where many points are earned or lost.
Does a score of 3 always earn college credit?
No. AP scores are reported from 1 to 5, but credit and placement policies are set by each college or university and may differ by program.
Can an online score calculator predict the 2027 cut score?
No public calculator can know a future operational score conversion. Use practice scores diagnostically and avoid presenting estimated boundaries as official.
Are older released questions still useful?
Yes for durable reasoning, but students must map them to the revised five-unit framework and recognize removed topics and the old six-FRQ paper format.
Which topics were removed?
The revision removed departures from linearity as a named topic, combining random variables, the geometric distribution, chi-square goodness-of-fit, and inference for regression slopes.
What does a complete inference response contain?
It identifies the parameter and hypotheses or interval target, names the procedure, checks relevant conditions, calculates, and concludes in context with the correct strength of evidence.
What is the best way to study?
Alternate retrieval practice, mixed problem sets, timed work, calculator fluency, and written error analysis. Reading notes alone does not test whether a student can select and justify a method.
Where should students get authentic AP questions?
Use College Board’s released-question page and AP Classroom through an authorized teacher. Do not rely on sites distributing secure exams.
How to verify a current AP Statistics fact
Start with the College Board page that owns the fact. The course page controls unit names and weighting; the assessment and revisions pages control exam structure; the calculator policy controls device rules; the student calendar controls the scheduled subject date. A search-result snippet, school handout, or older article can help locate a source but should not overrule a dated first-party page.
Match the administration year before carrying a fact into a plan. The public AP Statistics assessment page currently combines the revised fully digital structure with a displayed May 7, 2026 date. That does not make May 7 the 2027 date. The correct response is to report the format as revised for May 2027 while leaving the exact 2027 subject date unconfirmed until the official schedule itself changes.
Keep scope visible. A College Board statement that a handheld calculator is allowed does not mean every model is allowed; the approved-calculator policy supplies the model rules. A statement that reference information is provided does not mean every formula or interpretation is printed. A statement that recent FRQs are public does not authorize redistribution of secure practice exams from AP Classroom.
Record the page title, URL, claim, administration year, and date checked whenever an annual fact is used. Then revisit annual facts near registration, when the exam schedule is released, and before publication. This prevents a correct 2026 sentence from becoming a false 2027 sentence merely because a title was updated.
When two official pages appear out of sync, preserve both facts and describe the mismatch instead of silently choosing the convenient one. For example, a page can be updated with the revised 42-question format while its date field still shows the prior administration. The defensible article says exactly what each field verifies, includes a checked date, and links readers to the calendar for the unresolved annual detail. This approach protects students from planning around an invented date and keeps stable course information useful while the schedule is pending.
Apply the same rule to school logistics: College Board defines the exam, while the local AP coordinator supplies the student’s reporting location, arrival instruction, approved device plan, and any individually authorized accommodation. A general web page cannot replace those local directions, deadlines, or test-day communications accurately.
- Unit count and weights agree with the revised five-unit course page.
- Question counts, choices, section times, and weights agree with the revised assessment.
- The digital statement explicitly applies beginning in May 2027.
- The exact subject date is taken from the 2027 schedule, not inferred from 2026.
- Calculator language points to the current policy rather than an undated model list.
- Released-question links remain on College Board and do not expose secure files.
Official sources
| Official source | What it verifies |
|---|---|
| AP Statistics course page | Five units, unit weight ranges, prerequisites, and course description |
| AP Statistics revisions | 2026-27 changes, removed topics, fully digital May 2027 delivery, 42 MCQs, and four FRQs |
| AP Statistics assessment | Three-hour duration, section timing and weighting, and revised FRQ roles |
| AP calculator policy | Calculator expectations and approved-device rules |
| Reference information for AP exams | Reference materials and subject-specific exam details |
| Released AP Statistics questions | Public released materials and AP Classroom access for secure older resources |