Type I & Type II Error Visualizer
The four outcomes of any hypothesis test. Values update when you change controls in the Visualizer tab.
| Reality ↓ / Decision → | Fail to reject H₀ | Reject H₀ |
|---|---|---|
| H₀ is true | Correct non-rejection (1−α) = — |
Type I error (False positive) (α) = — |
| Specified H₁ is true | Type II error (False negative) (β) = — |
Correct detection (Power) (1−β) = — |
Power = 1 − β · Correct non-rejection = 1 − α
Row 1 sums to 1 · Row 2 sums to 1
β is not one universal number. It changes with the chosen alternative μ₁, sample size n, variability σ, and significance level α. The values above apply only to the currently selected scenario.
What Are Type I and Type II Errors?
Every statistical hypothesis test can produce one of four outcomes, and two of those outcomes are mistakes. A Type I error is rejecting the null hypothesis when it is actually true — a false positive. A Type II error is failing to reject the null hypothesis when a specified alternative is actually true — a false negative.
What makes these two errors worth studying together is the tension between them. Holding the study design fixed, making the rejection rule stricter (lower α) tends to reduce false positives but push up false negatives. The visualizer above makes this tradeoff visible rather than abstract.
How to Use This Visualizer
The chart shows two sampling distributions: the grey-blue curve centered at μ₀ is what the sample mean distribution looks like if H₀ is true, and the orange curve centered at μ₁ is what it looks like if the specified alternative is true. A vertical line marks the critical cutoff — the sample mean value at which the test switches from "fail to reject" to "reject."
The red shaded area under the H₀ curve to the right of the cutoff (for a right-tailed test) is the Type I error region. The yellow shaded area under the H₁ curve to the left of the cutoff is the Type II error region. The green area under H₁ to the right of the cutoff is power. These three things — what the test rejects, what it misses, and at what cost — are all on the same picture.
Type I Error and Alpha (α)
The Type I error rate is the probability that the test rejects H₀ when H₀ is true. In this Z-test model, that probability equals the selected significance level α exactly, because the critical value is chosen to make that so.
A common choice is α = 0.05, meaning that in repeated experiments conducted under these model assumptions when H₀ is true, the test produces a false rejection about 5% of the time. This is not the probability that H₀ is false — it is a long-run frequency property of the test procedure.
Type II Error and Beta (β)
The Type II error rate is the probability that the test fails to reject H₀ when a specific alternative is true. Its symbol is β. Unlike α, which the researcher sets directly, β is a consequence of multiple design choices together.
There is no single β for a given test. If the true mean is μ₁ = 101, β will be much larger than if the true mean is μ₁ = 115, because a smaller deviation from H₀ is harder to detect. This is why the visualizer has a "true alternative μ₁" control — power and β are only meaningful against a stated alternative.
Statistical Power
Power is the probability that the test correctly detects the specified effect. Power = 1 − β. A test with 80% power against a given alternative will, in the long run, detect that effect in 8 out of 10 repeated experiments. 80% is a common planning convention, not a universal requirement — the appropriate target depends on the context, the cost of missing an effect, and the field's standards.
Power is not the probability that H₁ is true. It is a conditional probability, calculated assuming a particular alternative is the true state of the world.
How Sample Size Changes the Picture
The standard error SE = σ/√n shrinks as n grows. Narrower sampling distributions mean less overlap between the H₀ and H₁ curves, which means less probability of missing the effect — β decreases and power increases. Crucially, α does not change when n increases: the significance level remains at whatever you set it to. Increasing sample size is the main way to improve power without accepting more false positives.
The Alpha–Beta Tradeoff
With n, σ, and the true alternative fixed, lowering α moves the critical boundary farther into the tail of H₀. This makes false positives less likely, but it also makes the test harder to pass, so more genuine effects go undetected — β rises and power falls. Raising α has the opposite effect. This is the classic error tradeoff the visualizer demonstrates when you drag the α slider while holding everything else constant.
Effect Size and Detection
A larger gap between μ₀ and μ₁ means the two sampling distributions are more separated. With greater separation, the area under H₁ that falls in the rejection zone grows — power increases. Effect size does not change α, and it does not move the rejection boundary; it moves the alternative distribution relative to that fixed boundary.
One-Sided vs. Two-Sided Tests
A one-sided test concentrates all of α in a single tail. A two-sided test splits α across both tails, placing half in each. For a given positive true effect, a right-tailed test will almost always have more power than a two-sided test with the same total α, because the two-sided test spends rejection probability on the opposite direction where the effect does not lie. The visualizer lets you switch between these modes to see the difference directly.
Type I vs. Type II Error — Summary Table
| Concept | Type I Error | Type II Error |
|---|---|---|
| Symbol | α | β |
| What happens | Reject H₀ when H₀ is true | Fail to reject H₀ when H₁ is true |
| Plain language | False positive | False negative |
| Set directly by researcher? | Yes — chosen as significance level | No — depends on design and effect |
| Related quantity | 1 − α (correct non-rejection) | Power = 1 − β |
| Fixed per test? | Yes, at the chosen α | No — changes with the true alternative |
Common Misconceptions
- α is not the probability H₀ is false. It is P(reject H₀ | H₀ true).
- β is not universal. It depends on which specific alternative is actually true.
- Power is not P(H₁ is true). It is the conditional probability of detection given that a specific alternative holds.
- Larger n does not lower α. With α fixed, n only affects β and power.
- "Fail to reject" is not "accept H₀." The test simply doesn't have enough evidence to reject.
- Neither error type is universally worse. Which matters more depends on the decision's costs and context.
Related Pages
Further reading:
- Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum.
- Neyman, J., & Pearson, E. S. (1933). On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society A, 231, 289–337.
- NIST Engineering Statistics Handbook — Type I and Type II Errors
Frequently Asked Questions
A Type I error is a false alarm. The test concludes that something is happening, such as an effect, difference, or relationship, when in reality nothing is. In the courtroom analogy, it corresponds to convicting an innocent person. Its probability is the significance level α that you choose before running the test.
A Type II error is a missed detection. The test fails to find an effect that is genuinely present. The analogy is a medical test that comes back negative even though the patient has the condition. Its probability β depends on how large the true effect is, how many observations are in the sample, and how strict the significance threshold is.
No. When α is held fixed, the nominal Type I error rate stays at α regardless of n. What larger n does is narrow the sampling distributions, which reduces β and raises power. You get better detection without paying more in false positives. That is the main statistical argument for larger samples.
Lowering α moves the critical cutoff deeper into the tail of H₀, which makes the rejection rule stricter. With n, σ, and the true alternative unchanged, this means more of the alternative distribution falls outside the rejection zone, so β increases and power decreases. This direct tradeoff is one of the main things the visualizer demonstrates. The only way to lower α and maintain power is to also increase n or target a larger effect.
No. 80% is a widely cited convention, introduced by Cohen (1988) as a reasonable starting point in the behavioral sciences. The right power target depends on the cost of missing the effect, the resources available, and the field's norms. Some contexts demand 90% or 95% power. Others may accept 70%. The visualizer lets you explore what n and effect size you need for any target power level.
Not simply by adjusting the rejection threshold. Moving the cutoff in one direction helps one error and hurts the other. But improving the study design can reduce β without changing α. The main tool for this is increasing sample size: narrower sampling distributions mean less overlap, lower β, and higher power, all at the same α. Reducing measurement noise, or lowering σ, has the same effect.
No, and confusing these two things is one of the most common errors in applied statistics. The p-value is calculated from an observed sample; α is set before data collection. A p-value below α means the test rejects H₀ in this particular result. α is the long-run frequency with which the test makes that rejection when H₀ is true. For an interactive look at p-values specifically, see the P-Value Visualizer.
Neither is universally worse. In drug safety testing, a Type I error, such as approving an ineffective or harmful drug, may be catastrophic. In cancer screening, a Type II error, or missing a genuine cancer, may be the greater concern. The relative cost depends on the decision context, its consequences, and who bears the risk. Researchers and regulators set α to reflect those tradeoffs, not because 0.05 is inherently the right threshold.