Quick Answer: When Do You Reject H₀?
That rule assumes three things are already in place: you chose the right statistical test for your question and data, you verified the test's assumptions are reasonably satisfied, and you set α before seeing the data. Skip any of those and the decision loses its meaning. The sections below build out each part of that process.
The p-value and critical-value methods are two routes to the same decision. When applied consistently — same test, same α, same tail specification — they agree every time.
H₀, H₁, α, and the p-Value: The Four Pieces
Before making any decision about H₀, you need four things defined. Here is what each one is and what it is not.
The Null Hypothesis (H₀)
The null hypothesis is the baseline statistical claim you are testing. It is written as an equality. Common forms:
| Test Type | Example H₀ |
|---|---|
| One-sample mean | H₀: μ = 100 |
| Two-sample mean | H₀: μ₁ = μ₂ |
| Proportion | H₀: p = 0.50 |
| Independence (chi-square) | H₀: variables are independent |
| ANOVA | H₀: μ₁ = μ₂ = ... = μk |
| Correlation | H₀: ρ = 0 |
One note worth keeping: H₀ does not always represent "no effect." Sometimes it specifies a particular value (H₀: μ = 50) or a constraint (H₀: μ₁ = μ₂). What it always does is give the test a specific model to compute probabilities against. For more null hypothesis examples, the dedicated page covers many test types.
The Significance Level (α)
Alpha (α) is a number you choose before collecting data. It is the maximum proportion of the time you are willing to wrongly reject a true H₀ — that is, commit a Type I error. Common levels are 0.05, 0.01, and 0.10. The choice depends on your field's conventions, the consequences of a false positive, and whether you're following a preregistered protocol.
α = 0.05 is the default in many courses, but it is a convention rather than a universal law. In pharmaceutical trials, α = 0.01 or lower is typical. In exploratory research, α = 0.10 is sometimes used. For more on this, the significance level page goes deeper. The key point here: always set α before you see p.
The p-Value
The p-value is calculated from your data after assuming H₀ is true. It answers the question: if H₀ were the correct model, what is the probability of getting a test result this extreme, or more extreme, just by chance?
A p-value of 0.03 does not mean there is a 3% probability that H₀ is true. The p-value is computed assuming H₀ is true. It measures how compatible your data are with that assumption, not the probability that the assumption itself is correct.
A small p-value means the observed data would be unusual under H₀. A large p-value means the data are consistent with H₀, though not proof of it. See the p-value examples page for more on reading these correctly.
The P-Value Decision Rule
The rule is the same for every standard test:
p = p-value from your test
α = significance level set before the test
If p ≤ α: Reject H₀
When p ≤ α, the observed result is sufficiently inconsistent with H₀ to warrant rejection at your chosen threshold. The result is called statistically significant at level α.
Here is a concrete run of examples with α = 0.05:
| p-value | α | Compare | Decision |
|---|---|---|---|
| 0.03 | 0.05 | 0.03 < 0.05 | Reject H₀ |
| 0.01 | 0.05 | 0.01 < 0.05 | Reject H₀ |
| 0.049 | 0.05 | 0.049 < 0.05 | Reject H₀ |
| 0.05 | 0.05 | 0.05 = 0.05 | Reject H₀ (p = α) |
If p > α: Fail to Reject H₀
When p > α, the data did not provide enough evidence, at the chosen threshold, to reject the null model. Use the phrase "fail to reject H₀" — not "accept H₀." The next section explains why that distinction matters.
| p-value | α | Compare | Decision |
|---|---|---|---|
| 0.18 | 0.05 | 0.18 > 0.05 | Fail to reject H₀ |
| 0.07 | 0.05 | 0.07 > 0.05 | Fail to reject H₀ |
| 0.051 | 0.05 | 0.051 > 0.05 | Fail to reject H₀ |
What If p = α Exactly?
By the conventional rule p ≤ α → reject, a p-value equal to α leads to rejection. In practice, this boundary case is rare with continuous test statistics, since continuous probability distributions produce p-values of any specific value with probability zero. If it does come up, reject H₀ per the standard rule and note the exact p-value in your report.
Same p-Value, Different α — Different Decision
This is worth seeing explicitly. Suppose p = 0.03:
| α chosen | p = 0.03 vs α | Decision |
|---|---|---|
| 0.10 | 0.03 < 0.10 | Reject H₀ |
| 0.05 | 0.03 < 0.05 | Reject H₀ |
| 0.01 | 0.03 > 0.01 | Fail to reject H₀ |
The data did not change. The decision changed because the threshold changed. This is precisely why α must be set before looking at p. Choosing α to match a convenient p-value after the fact is a form of inflating the Type I error rate.
The Critical-Value Approach
Instead of computing a p-value and comparing it to α, you can compare the test statistic directly to a precomputed cutoff called the critical value. Reject H₀ when the test statistic falls into the rejection region. The two approaches are mathematically equivalent when used consistently.
The rejection region depends on three things: the test distribution, α, and whether the test is one-tailed or two-tailed. See the full discussion of one-tailed vs two-tailed tests and the hypothesis-testing decision rule page for details. Here are the three standard cases for a Z-test at α = 0.05:
Two-Tailed Tests (H₁: μ ≠ μ₀)
The 1.96 comes from the standard normal distribution: 2.5% of the distribution sits in each tail (α/2 = 0.025 each), totalling α = 0.05.
Right-Tailed Tests (H₁: μ > μ₀)
Left-Tailed Tests (H₁: μ < μ₀)
Critical values depend on the test distribution (Z, t, χ², F), the degrees of freedom, α, and the tail direction. A chi-square statistic should never be compared against ±1.96. Use the statistical test selector to confirm which test applies, then look up the correct critical value.
Confidence Intervals and Null-Hypothesis Decisions
For many standard two-sided tests, there is a direct correspondence between the hypothesis test and the confidence interval: if the null value lies outside the matching confidence interval, the test rejects H₀ at the corresponding α.
Illustration
H₀: μ = 100, α = 0.05, 95% CI = [103.2, 109.8]
The null value 100 lies outside the interval. Under the corresponding two-sided t-test at α = 0.05, you would also reject H₀. The p-value in that test falls below 0.05. The confidence interval and the test point to the same conclusion.
If μ₀ = 106 in the same example, 106 falls inside [103.2, 109.8], and the test would fail to reject H₀. The correspondence holds when the test and interval are constructed from the same model with matching α and assumptions. One-sided tests require appropriately matched one-sided bounds, not standard two-sided intervals.
P-value: Reject H₀ when p ≤ α. | Critical value: Reject when the test statistic enters the rejection region. | CI: Reject when the null value falls outside the matching confidence interval.
Step-by-Step Decision Process
Decision Flowchart — When to Reject H₀
Result is statistically significant at level α
Insufficient evidence against H₀ at level α
State H₀ and H₁
Write both hypotheses before touching the data. H₀ is an equality. H₁ specifies the direction (≠, >, <). The direction of H₁ determines whether the test is two-tailed, right-tailed, or left-tailed, which in turn determines the rejection region.
Choose the Right Test
The test must match your question, data type, and structure. The statistical test selector walks through this. Using the wrong test produces a meaningless p-value regardless of what it says.
Set α Before Seeing the Data
Choose your significance level based on your field's conventions and the stakes involved. Common choices are 0.05, 0.01, and 0.10. Write it down. Do not revisit it after computing p.
Check Assumptions
Every test requires certain conditions: independence, approximate normality, equal variances, adequate sample size, and so on. A tiny p-value from a violated-assumption test does not justify a valid conclusion. See the assumptions page for test-specific guidance.
Compute the Test Statistic and p-Value
The test statistic converts your sample data into a single number on the test's distribution. The p-value is the tail probability of that statistic (or more extreme) under H₀. Use the t-test calculator, z-test calculator, or chi-square calculator for the arithmetic.
Make the Decision and Write the Conclusion
Compare p to α. Apply the rule. Then state the conclusion in plain language that connects the statistical decision back to the original question. The conclusion section below has exact wording templates for both outcomes.
Seven Worked Examples
Each example below uses a different scenario. The step structure matches the process above. For a broader library of test types, see the hypothesis testing examples page.
Example 1: p < α — Reject H₀
A manufacturer claims the mean weight of its product is μ₀ = 500g. A quality inspector samples 40 units and finds x̄ = 496g with σ = 10g (known). Test at α = 0.05, two-tailed.
Hypotheses: H₀: μ = 500 | H₁: μ ≠ 500 (two-tailed)
α = 0.05. Critical values: z = ±1.96.
Test: One-sample Z-test (σ known, n = 40).
Test statistic: SE = 10/√40 = 1.581. z = (496 − 500)/1.581 = −2.53.
p-value: P(Z < −2.53) ≈ 0.0057 (one tail). Two-tailed p = 2 × 0.0057 = 0.0114.
Decision: p = 0.0114 < α = 0.05 → Reject H₀. Also |z| = 2.53 > 1.96.
✅ Reject H₀. At the 5% significance level, the data provide evidence that the population mean differs from 500g. The sample mean of 496g is statistically significantly different from the claimed value.
Example 2: p > α — Fail to Reject H₀
A researcher tests whether a coin is fair. Out of 80 flips, 44 heads are observed. Test H₀: p = 0.50 at α = 0.05, two-tailed.
Hypotheses: H₀: p = 0.50 | H₁: p ≠ 0.50
α = 0.05. Critical values: z = ±1.96.
Test: One-proportion Z-test. p̂ = 44/80 = 0.55.
Test statistic: SE = √(0.50 × 0.50/80) = 0.0559. z = (0.55 − 0.50)/0.0559 = 0.89.
p-value: Two-tailed p ≈ 2 × P(Z > 0.89) ≈ 2 × 0.1867 = 0.373.
Decision: p = 0.373 > α = 0.05 → Fail to reject H₀.
✗ Fail to reject H₀. The data do not provide sufficient evidence at α = 0.05 to conclude the coin is biased. This does not prove the coin is fair — it means the observed departure from 0.50 is within normal random variation for a sample of this size.
Example 3: Right-Tailed Test
A university claims its graduates earn more than the national median of $50,000. A sample of 50 graduates has x̄ = $52,400, σ = $9,000. Test at α = 0.05, right-tailed.
Hypotheses: H₀: μ = 50,000 | H₁: μ > 50,000
α = 0.05, right-tailed. Critical value: z = 1.645.
Test: One-sample Z-test (σ known).
Test statistic: SE = 9000/√50 = 1272.8. z = (52400 − 50000)/1272.8 = 1.886.
p-value: P(Z > 1.886) ≈ 0.030.
Decision: z = 1.886 > 1.645 (rejection region). Also p = 0.030 < 0.05 → Reject H₀.
✅ Reject H₀. At α = 0.05, there is sufficient evidence to conclude that graduate earnings exceed $50,000. Both methods — p-value and critical value — confirm the same rejection.
Example 4: Left-Tailed Test — Fail to Reject
A factory claims its defect rate has dropped below 20%. An auditor samples 100 items and finds 18 defective. Test H₀: p = 0.20 at α = 0.05, left-tailed (testing whether the true rate is lower than 0.20).
Hypotheses: H₀: p = 0.20 | H₁: p < 0.20
α = 0.05, left-tailed. Critical value: z = −1.645.
Test: One-proportion Z-test. p̂ = 18/100 = 0.18.
Test statistic: SE = √(0.20 × 0.80/100) = 0.04. z = (0.18 − 0.20)/0.04 = −0.50.
p-value: P(Z < −0.50) ≈ 0.308.
Decision: z = −0.50 > −1.645 (not in rejection region). p = 0.308 > 0.05 → Fail to reject H₀.
✗ Fail to reject H₀. The data do not support the claim that the defect rate has dropped below 20% at α = 0.05. The observed rate of 18% is not statistically significantly lower.
Example 5: Two-Tailed Z-Test, Equivalence of Both Methods
H₀: μ = 75 vs H₁: μ ≠ 75. Observed z = −2.40, α = 0.05 (two-tailed). Demonstrate both decision methods give the same result.
Critical-value method: Two-tailed critical values are ±1.96. |z| = |−2.40| = 2.40 > 1.96 → test statistic is in the rejection region → Reject H₀.
P-value method: Two-tailed p = 2 × P(Z < −2.40) = 2 × 0.0082 = 0.0164. Since 0.0164 < 0.05 → Reject H₀.
✅ Both methods agree: Reject H₀. At the 5% level, there is sufficient evidence that the population mean differs from 75. When test, α, and tail direction are applied consistently, the two approaches always produce the same reject/fail-to-reject decision.
Example 6: Confidence Interval Method
H₀: μ = 100. The 95% confidence interval for μ is [101.4, 107.8]. Make the rejection decision without computing a p-value directly.
The test: Two-sided, α = 0.05 (matching the 95% CI).
Check: Does the null value 100 fall inside the interval [101.4, 107.8]? No. 100 < 101.4.
Decision: When the null value lies outside the matching confidence interval, the corresponding two-sided test rejects H₀ → Reject H₀.
✅ Reject H₀. The CI tells us the data support population means in the range [101.4, 107.8], and 100 is not among them. The p-value for the corresponding test falls below 0.05.
Example 7: The Borderline Case — p = 0.049 vs p = 0.051
Two studies, identical in design, produce p = 0.049 and p = 0.051 respectively. α = 0.05. What is the decision in each case, and what should we make of the difference?
Study A (p = 0.049): 0.049 < 0.05 → Reject H₀. Result is statistically significant at α = 0.05.
Study B (p = 0.051): 0.051 > 0.05 → Fail to reject H₀. Result is not statistically significant at α = 0.05.
The important point: The binary decision differs, but the underlying evidence is almost identical. A difference of 0.002 in p-values does not mean one study found an important effect and the other found nothing. Both provide nearly the same degree of incompatibility with H₀.
✅ Interpretation: Report exact p-values. Accompany them with effect sizes and confidence intervals. The threshold at α = 0.05 is a useful convention, not a scientific boundary that turns 0.049 into proof and 0.051 into nothing.
What "Reject" and "Fail to Reject" Actually Mean
What Rejecting H₀ Does and Does Not Mean
Rejecting H₀ means the observed evidence crossed the rejection threshold defined by your procedure. It does not mean:
- H₀ is logically impossible or proven false
- H₁ has been established with certainty
- The effect is large, important, or causal
- There is zero probability of a Type I error in this instance
Think of it this way: a jury verdict of "guilty" does not mean the defendant is certainly guilty — it means the evidence met the legal burden. Similarly, rejecting H₀ means the statistical evidence met the pre-set threshold, nothing more. The result should still be reported with an effect size and a confidence interval so readers can judge the practical significance themselves.
Why "Fail to Reject" Is Not the Same as "Accept H₀"
This distinction shows up on every statistics exam for good reason. When you fail to reject H₀, you are not concluding H₀ is true. You are concluding that the data, under the conditions of your test, did not meet the threshold for rejection. That non-rejection might occur because:
- H₀ is a reasonable approximation of reality
- The sample size was too small to detect the true effect
- The statistical power of the test was low
- The effect exists but is small relative to measurement variability
The correct wording after a non-significant result:
- Reject H₀: "At the α = 0.05 significance level, the data provide sufficient evidence to conclude that [H₁ in plain language]."
- Fail to reject H₀: "The data do not provide sufficient evidence, at α = 0.05, to reject the null hypothesis that [H₀ in plain language]."
- Never write: "Accept H₀" or "H₀ is proven" or "There is no effect."
Statistical Significance vs Practical Significance
A result can be statistically significant without being practically meaningful. Large samples make this easy to produce. With n = 1,000,000, a difference of 0.01 points on a scale from 0 to 100 might return p < 0.001 — perfectly real, statistically, but almost certainly irrelevant to any decision someone needs to make.
The reverse is also true: a meaningful effect may not reach significance if the sample is small, the measurement is noisy, or the test lacks power. Statistical significance answers "did the result cross the threshold?" — not "does this matter?"
After making a rejection decision, always ask:
- What is the effect size (Cohen's d, Pearson's r, odds ratio, or similar)?
- What does the confidence interval say about the plausible range of the effect?
- Is the effect size meaningful given the units and context?
Type I and Type II Errors
| H₀ Is True | H₀ Is False (Alternative Is True) | |
|---|---|---|
| Reject H₀ | Type I Error (α) False positive — wrongly rejected a true H₀ |
Correct Rejection Power = 1 − β |
| Fail to Reject H₀ | Correct Non-Rejection Probability = 1 − α |
Type II Error (β) False negative — missed a real effect |
The Type I error rate is exactly α under the null model and its assumptions — this is what α controls. The Type II error rate (β) depends on the true effect size, sample size, variability, and test power. Reducing α (requiring stronger evidence to reject) reduces Type I errors but typically increases Type II errors, holding everything else fixed.
Common Mistakes
| Mistake | Why It's Wrong | Correct Approach |
|---|---|---|
| Always comparing p to 0.05 | p should be compared to the prespecified α, which may be 0.01, 0.05, 0.10, or another value | Compare p to the α set before seeing the data |
| Writing "accept H₀" | Failing to reject H₀ is not proof that H₀ is true | Write "fail to reject H₀" |
| Interpreting p as P(H₀ is true) | p is computed assuming H₀ is true; it is not a posterior probability about H₀ | Describe p as the probability of observing this result (or more extreme) if H₀ were true |
| Changing α after seeing p | Post-hoc threshold adjustment inflates the true Type I error rate | Fix α before data collection; report the exact p-value |
| Using 1.96 as a universal critical value | 1.96 applies to two-tailed Z-tests at α = 0.05 only; t, χ², and F tests use different distributions | Look up the critical value for the specific test, degrees of freedom, and α |
| Equating statistical significance with practical importance | Large samples make trivial effects significant; small samples may miss meaningful ones | Report effect size and confidence interval alongside the p-value |
| Treating p > α as proof of no effect | Non-significance may reflect low power, not absence of an effect | Consider power, sample size, and the width of the confidence interval |
| Running multiple tests at the same α without adjustment | With 20 true-null tests at α = 0.05, the chance of at least one spurious rejection is about 64% | Use a Bonferroni correction, false discovery rate control, or other multiple-comparison procedure |
| Switching from two-tailed to one-tailed after seeing the direction of the result | Tail direction must be specified from the research question before the test | Pre-specify one-tailed or two-tailed based on the scientific question, not the observed direction |
| Ignoring test assumptions | A p-value from a violated-assumption test does not support valid inference | Check normality, independence, variance equality, and sample size requirements before running the test |
P-Value & Rejection Decision Calculator
Enter your test values below. The calculator computes the test statistic, p-value, and makes the reject/fail-to-reject decision automatically.
Hypothesis Test Decision Calculator
Practice Problems
Frequently Asked Questions
FAQ
Do you reject the null hypothesis when p < 0.05?
Only if you set α = 0.05 before the test. The rule is p ≤ α, not p < 0.05 specifically. With α = 0.01, a p-value of 0.03 does not justify rejection.
FAQ
What does it mean to fail to reject the null hypothesis?
It means the test result did not exceed your rejection threshold. It is not proof that H₀ is true. A non-significant test result may reflect limited power, a small sample, or measurement noise rather than a genuine absence of effect.
FAQ
What is the rejection region in hypothesis testing?
The rejection region is the set of test-statistic values that lead to rejection of H₀. Its boundaries — the critical values — depend on the test distribution, α, and the number of tails. For a two-tailed Z-test at α = 0.05, the rejection region is |z| > 1.96.
FAQ
What if p exactly equals alpha?
By the standard rule p ≤ α → reject, you would reject H₀. This exact boundary case is uncommon in practice with continuous distributions. Report the exact p-value so readers can assess the marginal result themselves.
FAQ
Are the p-value and critical-value approaches always equivalent?
Yes — when the same test, same α, and same tail specification are used consistently. A z-statistic of 2.30 with α = 0.05 two-tailed gives p ≈ 0.021 (< 0.05) and |z| = 2.30 > 1.96. Both approaches reject H₀.
FAQ
Does rejecting the null hypothesis prove causation?
No. Rejecting H₀ indicates a statistically significant association or difference under your test model. Causation requires a strong study design — typically a randomized controlled experiment — not just a small p-value.
Key Takeaways
- The rule: Reject H₀ when p ≤ α. Fail to reject H₀ when p > α.
- Set α first: Choose the significance level before looking at the data. Do not change it to match a convenient p-value.
- Both methods agree: The p-value and critical-value approaches reach the same decision when applied consistently.
- Reject ≠ prove false: Rejecting H₀ means the evidence crossed the threshold — not that H₀ is impossible or H₁ is certain.
- Fail to reject ≠ accept: A non-significant result may reflect low power or a small sample, not the absence of any effect.
- Significance ≠ importance: Always check the effect size and confidence interval alongside the p-value.
- Assumptions matter: A valid decision requires that the right test was chosen and its assumptions are reasonably satisfied.
- 1.96 is not universal: Use the critical value for your specific test, degrees of freedom, α, and tail direction.