Statistical Power and Sample Size Calculator
What Is Statistical Power?
Statistical power is the probability that a testing procedure rejects the null hypothesis when a specified alternative effect is true, given the assumptions used in the calculation. If power is 0.80, that does not mean there is an 80% probability that the research hypothesis is correct. It describes the long-run behavior of the test under one particular alternative.
Power is not a property of sample size by itself. The same sample size can have very different power under a different effect size, alpha level, test, allocation ratio, variance assumption, or one-sided versus two-sided alternative.
Type I error
α = prespecified false-positive rate under H₀Type II error
β = P(fail to reject H₀ | specified alternative)Power
1 − βPlanning question
What N gives the desired power for this test and effect?How Power Analysis Works
A useful a priori power analysis starts with the planned statistical test and an effect that would matter scientifically or practically. It then combines that effect with alpha, desired power, tails, and design assumptions. The calculator solves the corresponding power function for the missing quantity.
Sample Size, Effect Size, Alpha and Power
| Quantity | What it controls | Typical planning implication |
|---|---|---|
| Sample size | Amount of information available to the test. | For fixed design assumptions, larger N generally increases power. |
| Effect size | The alternative departure from the null that the calculation is designed around. | Larger effects are generally easier to detect than smaller effects. |
| Alpha | The Type I error criterion under the null. | A smaller alpha generally requires more information for the same effect and power. |
| Desired power | The target probability of rejecting H₀ under the specified alternative. | A higher target generally increases the required sample size. |
| Tails | Whether rejection can occur in one prespecified direction or both directions. | A one-sided test is not a sample-size shortcut; its direction should be justified before seeing the data. |
How to Choose an Effect Size
Base the planning effect on the smallest difference or association that would change a scientific, clinical, operational, or policy conclusion. Prior high-quality studies, meta-analyses, measurement literature, and well-designed pilot data can help. Conventional labels such as “small,” “medium,” and “large” are rough historical conventions, not substitutes for a meaningful effect.
For raw mean-difference modes, the assumed standard deviation matters just as much as the expected difference. Use a plausible SD from earlier studies or other defensible evidence rather than deriving it from the effect you hope to find.
One-Sided vs Two-Sided Tests
A two-sided test allows rejection for effects in either direction. A one-sided test places the rejection region in a prespecified direction. The direction should follow the research question and analysis plan, not a desire to reduce the required sample size after looking at the data.
Power Analysis by Statistical Test
One-Sample Mean
The standardized effect is d = (μ₁ − μ₀) / σ. The calculator uses the noncentral t distribution with df = n − 1. In raw mode, the expected mean difference is divided by the assumed standard deviation before power is evaluated.
Two Independent Means
This mode uses a pooled-variance two-sample t-test planning model. Cohen-style d is the expected mean difference divided by the common within-group standard deviation. The allocation ratio is Group 2 : Group 1. If unequal variances are central to the design, use a method specifically built for Welch-type planning rather than treating a pooled model as universal.
Paired Means
Paired power depends on the distribution of within-pair differences. The standardized effect used here is the mean paired difference divided by the standard deviation of those paired differences. Pre- and post-measurements are not treated as independent samples.
One-Way ANOVA
The ANOVA mode uses Cohen’s f, equal allocation across groups, and the noncentral F distribution. For k groups, numerator df = k − 1, denominator df = N − k, and the noncentrality parameter is λ = Nf². Unequal group sizes are not modeled in this version.
Correlation
The correlation mode uses Fisher’s z transformation for the difference between the alternative population correlation ρ₁ and null correlation ρ₀. This is an approximation rather than the exact distribution of the sample correlation. The approximation is usually more suitable once n is not extremely small.
Proportions
The one- and two-proportion modes use normal approximations without continuity correction. Critical boundaries are set under the null model and power is evaluated under the entered alternative proportions. Exact binomial procedures and continuity-corrected methods can produce different results, especially with small samples or proportions near 0 or 1.
Worked Examples
Two independent means
For a two-sided pooled two-sample t-test with d = 0.50, α = 0.05, desired power = 0.80, and equal allocation, the implemented noncentral-t method gives 64 observations per group, 128 total. The power after integer rounding is about 80.15%. The purpose of the example is to show the method; d = 0.50 is not a default recommendation.
Correlation sensitivity
Using the Fisher-z approximation with ρ₀ = 0, ρ₁ = 0.30, α = 0.05, 80% desired power, and a two-sided test, the minimum sample under this method is 85. A calculator using the exact correlation distribution can differ slightly.
A Priori, Sensitivity and “Post Hoc” Power
A priori analysis chooses a sample size before data collection for a specified meaningful alternative. Sensitivity analysis starts with a fixed sample size and asks what effect could be detected at a chosen alpha and power. Both are direct planning questions.
Power calculated from the observed effect using the same data is often of limited interpretive value. After a study is complete, effect estimates, confidence intervals, model diagnostics, and the study design usually provide more useful information than a retrospective “observed power” number. The power mode on this page therefore requires a specified alternative effect.
Common Power Analysis Mistakes
- Choosing a sample size before specifying the statistical test.
- Using a conventional “medium” effect without a scientific reason.
- Interpreting 80% power as an 80% probability that a hypothesis is true.
- Ignoring one-sided versus two-sided alternatives.
- Ignoring group allocation in two-group designs.
- Rounding a required sample size down.
- Treating paired observations as independent.
- Replacing post-study estimation with observed post hoc power.
- Assuming a larger N fixes bias, confounding, or measurement problems.
- Using simple independent-observation formulas for clustered, longitudinal, survival, noninferiority, or multilevel designs.
Related Statistics Tools and Guides
Method Notes and Verification
The t-test modes evaluate power with central and noncentral t distributions. Balanced one-way ANOVA uses a central critical F value and a noncentral F distribution under the alternative. The sample-size and effect solvers use bounded deterministic searches and round required integer sample sizes upward until the requested power is met.
Frequently Asked Questions
A power analysis examines the operating characteristics of a specified statistical test. In planning, it is commonly used to solve for sample size given an effect, alpha, desired power, tails, and other design assumptions. It can also solve for power at a fixed sample or for a detectable effect at a fixed sample.
Under the statistical model and alternative effect used in the calculation, 80% power means the procedure has an 80% probability of rejecting the null hypothesis if that specified alternative is true. It is not an 80% probability that the hypothesis is correct.
No universal target applies to every study. Targets such as 80% and 90% are common examples, but an appropriate target depends on the scientific setting, consequences of a Type II error, feasibility, regulation, and the role of the study.
Select the planned analysis, define the effect in the units required by that test, set alpha and desired power, specify the alternative direction, and enter allocation assumptions where relevant. The calculator then finds the smallest integer sample configuration that reaches the target under its stated method.
Check the effect-size definition, test family, tails, variance model, allocation ratio, exact versus approximate distribution, continuity correction, and rounding convention. A difference does not necessarily mean one calculator is wrong; the methods may not be identical.
No. A larger sample usually increases power for a fixed model, but it cannot repair selection bias, confounding, poor measurement, a weak intervention, incorrect randomization, or model misspecification. Very large samples can also make scientifically trivial effects statistically detectable.
If dropout is expected, a recruitment target can be increased so the planned analyzable sample remains available. This page uses recruitment target = analyzable N / (1 − attrition proportion). That arithmetic does not remove bias when dropout is related to outcomes or treatment.
Use design-specific methods for clustered or multilevel data, repeated measures beyond a simple paired comparison, time-to-event outcomes, noninferiority or equivalence tests, sequential monitoring, complex regression models, or other designs where the independence and distribution assumptions of the basic modes do not match the planned analysis.