P-Value, Significance Level, and Test Statistic
Separate three linked but different ideas: the test statistic measures how far the sample result is from the null model, the p-value measures tail evidence under that model, and alpha supplies the decision threshold.
P-Value, Significance Level, and Test Statistic: direct answer
In a p value significance test statistic problem, begin by identifying the population target and the inferential role of the data. Separate three linked but different ideas: the test statistic measures how far the sample result is from the null model, the p-value measures tail evidence under that model, and alpha supplies the decision threshold.
This page uses p value significance test statistic as its single primary search focus. The lesson, numerical cases, and retained questions are restricted to that intent so the page does not function as a generic inference question bank.
Quick reference for P-Value, Significance Level, and Test Statistic
| Test statistic | Standardized discrepancy from H₀ |
|---|---|
| P-value | Tail probability assuming H₀ |
| α | Prespecified significance threshold |
| Reject H₀ | p≤α |
| Fail to reject H₀ | p>α |
| Not a p-value interpretation | Probability that H₀ is true |
Concept mastery: p value significance test statistic
The test statistic standardizes discrepancy
A test statistic compares the observed estimate with the null benchmark in standard-error units. Its reference distribution depends on the procedure. Large absolute values often indicate stronger disagreement with H0, but direction matters for one-sided alternatives.
The p-value is conditional on H0
A p-value is the probability, assuming H0 and the model are true, of obtaining a test statistic at least as extreme as the observed one in the direction or directions specified by Ha. It is not P(H0 is true | data).
Alpha is chosen before the evidence is examined
The significance level alpha sets the long-run Type I error rate for the decision rule under H0. Common choices such as .05 are conventions, not natural laws, and should not be changed after seeing the p-value.
Tail direction controls the p-value
A positive z can produce a small upper-tail p-value, a large lower-tail p-value, or a two-sided p-value twice the smaller symmetric tail. Always read Ha before converting a statistic into evidence.
Statistical significance is a comparison, not a magnitude
Calling a result significant means p<=alpha. It does not mean the effect is large, important, or practically useful. Effect size and uncertainty need separate interpretation.
A p-value near alpha is not a scientific cliff
Results just below .05 and just above .05 can contain very similar evidence. The decision rule may differ, but interpretation should still consider estimate size, uncertainty, design quality, and the broader context.
Fail to reject is not accept
When p>alpha, the data do not provide enough evidence for Ha at that threshold. This is not the same as proving H0 true. Low power or high variability can also produce a large p-value.
Smaller p-values indicate greater incompatibility with H0
Within the specified model and test, a smaller p-value means the observed result lies farther into the relevant tail region. It does not measure the probability that the result will replicate or the probability that Ha is correct.
Test statistics from different procedures are not interchangeable
A z statistic, t statistic, chi-square statistic, and other test statistics have different reference distributions and sometimes different sign behavior. A value of 2 has meaning only after the procedure and degrees of freedom are known.
Report the chain from statistic to conclusion
Strong reporting names the statistic, p-value, alpha if a formal decision is requested, and contextual conclusion. Keeping these pieces distinct makes it clear what was calculated, what evidence means, and what decision was made.
A p-value of .049 and a p-value of .051 are not different kinds of evidence
With α=.05, one result crosses the formal rejection threshold and the other does not, but the numerical evidence is nearly the same. Treating .049 as a discovery and .051 as no evidence exaggerates what a decision rule can say. Report the actual p-value, the estimated effect and its uncertainty, and the study design. The threshold is useful for a prespecified decision process, but it should not erase the continuity of statistical evidence.
Large samples can make tiny effects statistically significant
A test statistic divides an observed departure by its standard error. As sample size grows, standard errors often shrink, so a small departure can produce a large standardized statistic and a small p-value. This is why statistical significance cannot stand in for practical importance. A policy change of one tenth of a percentage point may be precisely estimated and highly significant while still being too small to matter operationally.
One-sided and two-sided p-values answer different alternatives
The same observed z statistic can generate different p-values depending on Hₐ. A greater-than alternative uses the upper tail; a less-than alternative uses the lower tail; a two-sided alternative counts extremeness in both directions. The alternative must therefore be chosen from the research question before the data are examined. Selecting the smaller tail after seeing the sign of the statistic would make the reported p-value too favorable.
Reference distributions give test statistics their meaning
A statistic of 2.3 is incomplete information unless the reference distribution is known. z=2.3 is interpreted through the standard normal curve; t=2.3 depends on degrees of freedom; a chi-square statistic is nonnegative and uses a right-tail chi-square distribution. The p-value is created by placing the statistic in the correct reference distribution and measuring the tail region specified by the alternative.
Confidence intervals and p-values provide complementary evidence
A p-value focuses on compatibility with a particular null value, whereas a confidence interval displays a range of plausible parameter values under its procedure. For a corresponding two-sided test and interval, exclusion of the null value is closely connected to significance at the matching level. The interval usually adds information about effect magnitude and precision that a binary reject/fail-to-reject decision cannot provide by itself.
Report significance without turning it into certainty
A complete report can state the estimate, test statistic, p-value, and decision threshold when one was prespecified, followed by a contextual conclusion. Avoid phrases such as “the null is definitely false,” “there is a 3% chance the null is true,” or “the result happened by chance.” Those statements assign meanings that the frequentist p-value does not contain. The evidence is conditional on the null model and on the assumptions that justify the reference distribution.
Rounding should not manufacture or erase significance
When a calculator reports a p-value close to the chosen alpha, keep enough digits to make the comparison honestly. Rounding p=.0496 to .05 and then treating it as exactly equal to the threshold can obscure the original evidence; rounding p=.0504 to .05 can do the same in the opposite direction. Report a sensible number of significant digits, compare using the unrounded value when available, and avoid writing p=0 when software displays a value smaller than its normal decimal precision. A very small p-value is still positive unless the mathematical tail probability is exactly zero.
A small p-value cannot rescue a biased or confounded study
The reference distribution answers a model-based question: how unusual would a statistic this extreme be if H₀ and the stated assumptions were true? It does not certify that the sample represents the intended population, that treatment groups are comparable, or that a lurking variable has been controlled. A badly biased survey can produce an extremely small p-value for a precisely estimated but systematically distorted effect. Likewise, an observational association can be statistically significant without establishing causation. Design validity and statistical evidence are separate checkpoints, and both must survive before the conclusion is trusted.
Simulation-based p-values use the same evidence logic
When a null distribution is generated by randomization or simulation rather than a closed-form z, t, or chi-square curve, the p-value still measures the proportion of null outcomes at least as extreme as the observed statistic in the direction specified by Hₐ. The mechanics change, but the interpretation does not become the probability that H₀ is true. This connection is useful because it shows that the core idea of a p-value is tail extremeness under a null model, not dependence on one particular formula or calculator command.
Worked analysis for p value significance test statistic
Evidence case 1: Municipal Water Office
A valid test concerning households reporting no service interruption produces z=-2.65 with a less alternative. Converting that statistic through the standard normal reference distribution gives p=0.0040. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.01, 0.0040 ≤ 0.01, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0040, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 2: State Park
A valid test concerning visitors who used the marked trail system produces z=-1.85 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.0643. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.05, 0.0643 > 0.05, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0643, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 3: Food Cooperative
A valid test concerning members who renewed before the deadline produces z=-1.32 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.9066. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.10, 0.9066 > 0.10, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.9066, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 4: University Advising Center
A valid test concerning appointments starting within ten minutes produces z=0.74 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.4593. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.01, 0.4593 > 0.01, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.4593, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 5: Regional Manufacturer
A valid test concerning parts meeting the diameter specification produces z=1.62 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.0526. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.05, 0.0526 > 0.05, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0526, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 6: Public Health Clinic
A valid test concerning clients returning for the scheduled checkup produces z=2.41 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.0160. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.10, 0.0160 ≤ 0.10, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0160, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 7: Urban Recreation Program
A valid test concerning participants completing the eight-week session produces z=2.88 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.0020. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.01, 0.0020 ≤ 0.01, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0020, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 8: School District
A valid test concerning families responding to the annual survey produces z=-2.12 with a less alternative. Converting that statistic through the standard normal reference distribution gives p=0.0170. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.05, 0.0170 ≤ 0.05, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0170, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 9: Local Election Office
A valid test concerning mailed ballots returned before election day produces z=-2.65 with a less alternative. Converting that statistic through the standard normal reference distribution gives p=0.0040. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.10, 0.0040 ≤ 0.10, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0040, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 10: Energy-Efficiency Pilot
A valid test concerning homes meeting the target reduction produces z=-1.85 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.0643. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.01, 0.0643 > 0.01, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0643, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 11: Community Broadband Project
A valid test concerning households achieving the advertised speed produces z=-1.32 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.9066. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.05, 0.9066 > 0.05, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.9066, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 12: Regional Bus Network
A valid test concerning trips arriving within the on-time window produces z=0.74 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.4593. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.10, 0.4593 > 0.10, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.4593, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 13: Campus Dining Service
A valid test concerning transactions using reusable containers produces z=1.62 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.0526. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.01, 0.0526 > 0.01, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0526, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 14: Workforce Training Program
A valid test concerning participants earning the credential produces z=2.41 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.0160. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.05, 0.0160 ≤ 0.05, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0160, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 15: County Recycling Audit
A valid test concerning sampled loads meeting contamination limits produces z=2.88 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.0020. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.10, 0.0020 ≤ 0.10, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0020, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 16: University Residence Halls
A valid test concerning rooms passing the first safety inspection produces z=-2.12 with a less alternative. Converting that statistic through the standard normal reference distribution gives p=0.0170. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.01, 0.0170 > 0.01, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0170, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 17: Telehealth Pilot
A valid test concerning appointments completed without rescheduling produces z=-2.65 with a less alternative. Converting that statistic through the standard normal reference distribution gives p=0.0040. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.05, 0.0040 ≤ 0.05, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0040, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 18: Public Museum
A valid test concerning visitors using the audio guide produces z=-1.85 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.0643. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.10, 0.0643 ≤ 0.10, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0643, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 19: Youth Sports League
A valid test concerning players completing concussion training produces z=-1.32 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.9066. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.01, 0.9066 > 0.01, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.9066, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 20: Rural Pharmacy Network
A valid test concerning prescriptions filled within the service target produces z=0.74 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.4593. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.05, 0.4593 > 0.05, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.4593, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 21: City Tree Program
A valid test concerning new plantings surviving the first year produces z=1.62 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.0526. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.10, 0.0526 ≤ 0.10, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0526, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 22: District Tutoring Initiative
A valid test concerning students attending at least six sessions produces z=2.41 with a two-sided alternative. Converting that statistic through the standard normal reference distribution gives p=0.0160. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.01, 0.0160 > 0.01, so the decision is to fail to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0160, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 23: Regional Call Center
A valid test concerning calls resolved during the first contact produces z=2.88 with a greater alternative. Converting that statistic through the standard normal reference distribution gives p=0.0020. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.05, 0.0020 ≤ 0.05, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0020, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
Evidence case 24: City Transit Survey
A valid test concerning riders who used the mobile ticket option produces z=-2.12 with a less alternative. Converting that statistic through the standard normal reference distribution gives p=0.0170. The statistic reports standardized distance from the null benchmark; the p-value reports how much tail area is at least as extreme, assuming H₀ and the test model are correct.
Using α=0.10, 0.0170 ≤ 0.10, so the decision is to reject H₀. Statistical significance here is a threshold comparison, not an effect-size label. The p-value does not mean that H₀ has probability 0.0170, and a result near α should still be interpreted with the estimate, design quality, uncertainty, and practical stakes in view.
P Value Significance Test Statistic multiple-choice practice
Question 1. Statistic, tail area, and alpha
A valid test for follow-up completion rate at community health network produces z=-2.45 with a less alternative and α=0.01. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: B
z=-2.45 records standardized distance from the null benchmark; it is not itself a probability. The less reference-tail calculation gives p=0.0071. Since 0.0071 ≤ α=0.01, the formal significance decision is reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 2. Statistic, tail area, and alpha
A valid test for mean weekly study time at regional college produces z=-1.78 with a two-sided alternative and α=0.05. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: A
z=-1.78 records standardized distance from the null benchmark; it is not itself a probability. The two-sided reference-tail calculation gives p=0.0751. Since 0.0751 > α=0.05, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 3. Statistic, tail area, and alpha
A valid test for on-time arrival proportion at city transit agency produces z=-1.12 with a greater alternative and α=0.10. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: D
z=-1.12 records standardized distance from the null benchmark; it is not itself a probability. The greater reference-tail calculation gives p=0.8686. Since 0.8686 > α=0.10, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 4. Statistic, tail area, and alpha
A valid test for program participation rate at county library system produces z=0.86 with a two-sided alternative and α=0.01. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: C
z=0.86 records standardized distance from the null benchmark; it is not itself a probability. The two-sided reference-tail calculation gives p=0.3898. Since 0.3898 > α=0.01, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 5. Statistic, tail area, and alpha
A valid test for mean component strength at manufacturing plant produces z=1.44 with a less alternative and α=0.05. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: B
z=1.44 records standardized distance from the null benchmark; it is not itself a probability. The less reference-tail calculation gives p=0.9251. Since 0.9251 > α=0.05, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 6. Statistic, tail area, and alpha
A valid test for graduation-plan completion proportion at school district produces z=1.97 with a two-sided alternative and α=0.10. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: A
z=1.97 records standardized distance from the null benchmark; it is not itself a probability. The two-sided reference-tail calculation gives p=0.0488. Since 0.0488 ≤ α=0.10, the formal significance decision is reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 7. Statistic, tail area, and alpha
A valid test for visitor satisfaction proportion at public parks department produces z=2.31 with a greater alternative and α=0.01. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: D
z=2.31 records standardized distance from the null benchmark; it is not itself a probability. The greater reference-tail calculation gives p=0.0104. Since 0.0104 > α=0.01, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 8. Statistic, tail area, and alpha
A valid test for mean household savings at energy pilot produces z=2.76 with a two-sided alternative and α=0.05. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: C
z=2.76 records standardized distance from the null benchmark; it is not itself a probability. The two-sided reference-tail calculation gives p=0.0058. Since 0.0058 ≤ α=0.05, the formal significance decision is reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 9. Statistic, tail area, and alpha
A valid test for 30-day readmission proportion at hospital system produces z=-2.45 with a less alternative and α=0.10. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: B
z=-2.45 records standardized distance from the null benchmark; it is not itself a probability. The less reference-tail calculation gives p=0.0071. Since 0.0071 ≤ α=0.10, the formal significance decision is reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 10. Statistic, tail area, and alpha
A valid test for credential completion rate at workforce program produces z=-1.78 with a two-sided alternative and α=0.01. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: A
z=-1.78 records standardized distance from the null benchmark; it is not itself a probability. The two-sided reference-tail calculation gives p=0.0751. Since 0.0751 > α=0.01, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 11. Statistic, tail area, and alpha
A valid test for mean resolution time at municipal call center produces z=-1.12 with a greater alternative and α=0.05. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: D
z=-1.12 records standardized distance from the null benchmark; it is not itself a probability. The greater reference-tail calculation gives p=0.8686. Since 0.8686 > α=0.05, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 12. Statistic, tail area, and alpha
A valid test for household reliability proportion at broadband project produces z=0.86 with a two-sided alternative and α=0.10. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: C
z=0.86 records standardized distance from the null benchmark; it is not itself a probability. The two-sided reference-tail calculation gives p=0.3898. Since 0.3898 > α=0.10, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 13. Statistic, tail area, and alpha
A valid test for mean repair turnaround time at university housing produces z=1.44 with a less alternative and α=0.01. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: B
z=1.44 records standardized distance from the null benchmark; it is not itself a probability. The less reference-tail calculation gives p=0.9251. Since 0.9251 > α=0.01, the formal significance decision is fail to reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 14. Statistic, tail area, and alpha
A valid test for member renewal proportion at food cooperative produces z=1.97 with a two-sided alternative and α=0.05. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: A
z=1.97 records standardized distance from the null benchmark; it is not itself a probability. The two-sided reference-tail calculation gives p=0.0488. Since 0.0488 ≤ α=0.05, the formal significance decision is reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 15. Statistic, tail area, and alpha
A valid test for mean weekly attendance at recreation department produces z=2.31 with a greater alternative and α=0.10. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: D
z=2.31 records standardized distance from the null benchmark; it is not itself a probability. The greater reference-tail calculation gives p=0.0104. Since 0.0104 ≤ α=0.10, the formal significance decision is reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
Question 16. Statistic, tail area, and alpha
A valid test for compliance proportion at water authority produces z=2.76 with a two-sided alternative and α=0.01. Which interpretation correctly distinguishes the test statistic, p-value, and significance decision?
Answer: C
z=2.76 records standardized distance from the null benchmark; it is not itself a probability. The two-sided reference-tail calculation gives p=0.0058. Since 0.0058 ≤ α=0.01, the formal significance decision is reject H₀. The p-value remains conditional on H₀ and does not assign a posterior probability to H₀.
P Value Significance Test Statistic free-response practice
FRQ 1. Interpret the evidence chain
For a test concerning mean component strength at manufacturing plant, the standardized test statistic is z=-2.08 and Hₐ determines a less tail calculation. Compute or identify the p-value, compare it with α=0.05, explain the formal decision, and state two interpretations that would be incorrect.
Model response
The less tail area gives p≈0.0188. Because 0.0188 ≤ 0.05, the decision is to reject H₀. The p-value is the probability, under H₀ and the model, of obtaining a statistic at least as extreme in the direction specified by Hₐ. It is incorrect to call p the probability H₀ is true, and it is also incorrect to treat statistical significance as a measure of practical importance. The statistic z=-2.08, p-value, and alpha play different roles in the reasoning chain.
FRQ 2. Interpret the evidence chain
For a test concerning 30-day readmission proportion at hospital system, the standardized test statistic is z=1.63 and Hₐ determines a greater tail calculation. Compute or identify the p-value, compare it with α=0.05, explain the formal decision, and state two interpretations that would be incorrect.
Model response
The greater tail area gives p≈0.0516. Because 0.0516 > 0.05, the decision is to fail to reject H₀. The p-value is the probability, under H₀ and the model, of obtaining a statistic at least as extreme in the direction specified by Hₐ. It is incorrect to call p the probability H₀ is true, and it is also incorrect to treat statistical significance as a measure of practical importance. The statistic z=1.63, p-value, and alpha play different roles in the reasoning chain.
FRQ 3. Interpret the evidence chain
For a test concerning mean repair turnaround time at university housing, the standardized test statistic is z=2.34 and Hₐ determines a two-sided tail calculation. Compute or identify the p-value, compare it with α=0.01, explain the formal decision, and state two interpretations that would be incorrect.
Model response
The two-sided tail area gives p≈0.0193. Because 0.0193 > 0.01, the decision is to fail to reject H₀. The p-value is the probability, under H₀ and the model, of obtaining a statistic at least as extreme in the direction specified by Hₐ. It is incorrect to call p the probability H₀ is true, and it is also incorrect to treat statistical significance as a measure of practical importance. The statistic z=2.34, p-value, and alpha play different roles in the reasoning chain.
FRQ 4. Interpret the evidence chain
For a test concerning follow-up completion rate at community health network, the standardized test statistic is z=-1.41 and Hₐ determines a two-sided tail calculation. Compute or identify the p-value, compare it with α=0.10, explain the formal decision, and state two interpretations that would be incorrect.
Model response
The two-sided tail area gives p≈0.1585. Because 0.1585 > 0.10, the decision is to fail to reject H₀. The p-value is the probability, under H₀ and the model, of obtaining a statistic at least as extreme in the direction specified by Hₐ. It is incorrect to call p the probability H₀ is true, and it is also incorrect to treat statistical significance as a measure of practical importance. The statistic z=-1.41, p-value, and alpha play different roles in the reasoning chain.
FRQ 5. Interpret the evidence chain
For a test concerning mean component strength at manufacturing plant, the standardized test statistic is z=2.71 and Hₐ determines a greater tail calculation. Compute or identify the p-value, compare it with α=0.01, explain the formal decision, and state two interpretations that would be incorrect.
Model response
The greater tail area gives p≈0.0034. Because 0.0034 ≤ 0.01, the decision is to reject H₀. The p-value is the probability, under H₀ and the model, of obtaining a statistic at least as extreme in the direction specified by Hₐ. It is incorrect to call p the probability H₀ is true, and it is also incorrect to treat statistical significance as a measure of practical importance. The statistic z=2.71, p-value, and alpha play different roles in the reasoning chain.
FRQ 6. Interpret the evidence chain
For a test concerning 30-day readmission proportion at hospital system, the standardized test statistic is z=-2.52 and Hₐ determines a two-sided tail calculation. Compute or identify the p-value, compare it with α=0.05, explain the formal decision, and state two interpretations that would be incorrect.
Model response
The two-sided tail area gives p≈0.0117. Because 0.0117 ≤ 0.05, the decision is to reject H₀. The p-value is the probability, under H₀ and the model, of obtaining a statistic at least as extreme in the direction specified by Hₐ. It is incorrect to call p the probability H₀ is true, and it is also incorrect to treat statistical significance as a measure of practical importance. The statistic z=-2.52, p-value, and alpha play different roles in the reasoning chain.
Next steps after mastering p value significance test statistic
After this page, take several reported test results and identify which number is the standardized statistic, which is the p-value, and which is alpha. Then write one sentence for each number that states what it does and one sentence that states what it does not mean. This prevents the three quantities from collapsing into a single vague idea of “significance.”
For threshold sensitivity, compare the same p-value with alpha=.10, .05, and .01 while keeping the estimate and test statistic unchanged. The formal decision can change even though the data do not. That exercise is useful for understanding why effect size, interval estimates, design quality, and context should accompany a binary significance label.