Quick answer: how to choose a statistical test
Ask what you want to compare or associate, identify the outcome and explanatory variables, count the groups, decide whether observations are independent or paired, and then check the assumptions of the methods that remain. A decision tree can narrow the field, but it cannot decide what your scientific question means.
- Two independent group means: consider a two-sample t procedure, often Welch's t-test.
- Two paired measurements: consider a paired t-test when inference about the mean difference is appropriate.
- Three or more independent group means: one-way ANOVA is a common starting point.
- Two categorical variables: chi-square test of independence, or Fisher's exact test for small tables when appropriate.
- Two quantitative variables: Pearson for linear association; Spearman for monotonic rank association.
- Rank-based group comparisons: Mann-Whitney U, Wilcoxon signed-rank, or Kruskal-Wallis, depending on the design.
Statistical tests at a glance
The table is a screening tool, not a substitute for reading the assumptions. The same dataset can require different methods for different questions.
| Test | Main question | Typical data/design | What to check |
|---|---|---|---|
| Independent-samples t-test | Do two independent group means differ? | Continuous outcome, 2 independent groups | Independence, model for means, outliers; variance assumption depends on version |
| Paired-samples t-test | Is the mean paired difference zero? | Continuous paired differences | Correct pairing, independence across pairs, distribution of differences |
| One-way ANOVA | Are all group means equal? | Continuous outcome, 3+ independent groups | Independence, residual behavior, variance structure |
| Chi-square test of independence | Are two categorical variables associated? | Counts in a contingency table | Independent observations, expected-count adequacy for approximation |
| Fisher's exact test | Are two categorical variables associated in a small table? | Usually a 2 by 2 count table | Sampling design and fixed-margin interpretation |
| Pearson correlation test | Is there a linear association? | Two quantitative variables | Linearity, independence, influential points, assumptions for inference |
| Spearman rank correlation | Is there a monotonic rank association? | Ordinal or quantitative variables | Independent pairs, meaningful ordering, monotonic pattern |
| Mann-Whitney U test | Do two independent groups differ in rank distribution? | Ordinal/continuous outcome, 2 independent groups | Independence; interpretation depends on distributional relationship |
| Wilcoxon signed-rank test | Are paired differences centered around zero in the signed-rank sense? | Paired ordinal/continuous measurements | Paired structure, independent pairs, symmetry for common location-shift interpretation |
| Kruskal-Wallis test | Do 3+ independent groups differ in rank distribution? | Ordinal/continuous outcome, 3+ independent groups | Independent groups; shape assumptions matter for location interpretation |
What is a statistical test?
Null and alternative hypotheses
The null hypothesis is the reference claim used to calculate the test statistic and p-value. The alternative describes the competing values or patterns considered by the test. The exact null is test-specific. It might state equality of means, independence of categorical variables, zero correlation, or a specified parameter value.
Test statistic, p-value, and significance level
A test statistic summarizes how the observed data compare with what the null model would lead us to expect. A p-value is then calculated under that null model. It describes how incompatible the observed test result, or a result at least as extreme according to the chosen statistic, is with the null. It is not the probability that the null hypothesis is true.
The significance level, usually written as α, is a decision threshold chosen before interpreting the test. Values such as 0.05 or 0.01 are conventions, not universal scientific rules. A Type I error means rejecting a true null hypothesis; a Type II error means failing to reject a false null hypothesis. Statistical power is the probability of rejecting the null when a specified alternative is true.
How to choose the right statistical test
Test selection works better as a sequence of questions than as a lookup based on one variable. Before the steps, identify what each variable represents.
Identify your variables before choosing a test
| Variable type | Examples | Why it matters |
|---|---|---|
| Continuous | Height, income, blood pressure, response time | Can support mean-based models, correlation, regression, and rank-based analyses depending on the question. |
| Ordinal | Satisfaction ratings, ordered severity categories | The categories have an order, but equal numerical spacing may not be defensible. |
| Nominal/categorical | Treatment group, product choice, disease category | Often analyzed through counts, proportions, contingency tables, or categorical regression models. |
| Binary | Yes/no, success/failure, event/no event | A two-category outcome often calls for proportion methods or binary-response models rather than a mean test. |
The outcome variable is the response you want to analyze. The explanatory or grouping variable defines groups or helps explain variation. For the question "Do mean exam scores differ between teaching methods?", exam score is the outcome and teaching method is the grouping variable.
Write the research question in statistical terms
Are you comparing means, comparing distributions, testing association between categories, measuring correlation, or estimating a parameter against a reference value?
Identify the outcome and explanatory variables
Continuous, ordinal, nominal, and binary variables support different summaries and models. Measurement scale matters, but it is only one part of the decision.
Count the groups or measurement occasions
Two groups and four groups are different analysis problems. Repeatedly running pairwise t-tests across many groups inflates the chance of false positives unless multiplicity is handled.
Decide whether observations are independent or paired
Before-and-after observations from the same person are paired. Measurements from separate treatment and control participants are usually independent. The analysis must respect that dependence structure.
Check assumptions in the context of the chosen model
Look at outliers, distributional shape, variance structure, sample size, dependence, and the exact inferential target. Do not use a preliminary normality test as an automatic switch between parametric and nonparametric procedures.
Plan interpretation before running the test
Decide what estimate, effect size, confidence interval, and diagnostic information you will report with the p-value.
Parametric vs nonparametric tests
Parametric methods specify a statistical model through parameters and structural or distributional assumptions. A t-test, for example, is built around a model for means and standard errors. Nonparametric and rank-based methods use a different set of assumptions and targets. They are not assumption-free, and they are not automatically the right answer whenever data look non-normal.
| Feature | Parametric methods | Rank-based/nonparametric methods |
|---|---|---|
| Typical target | Parameters such as means, regression coefficients, correlations | Ranks, distributions, stochastic ordering, or other nonparametric features |
| Assumptions | Model-specific distributional and structural assumptions | Different or fewer distributional assumptions, but still assumptions |
| Common examples | t-tests, ANOVA, Pearson correlation | Mann-Whitney, Wilcoxon signed-rank, Kruskal-Wallis, Spearman |
| Selection rule | Choose based on the question, design, estimand, data behavior, and assumptions. Do not use a single normality p-value as the rule. | |
Read the full parametric vs nonparametric comparison.
1. Independent-samples t-test
Use a two-sample t procedure when the scientific question is about a difference in means between two independent groups. In many routine analyses, Welch's t-test is a safer default than the pooled equal-variance version because it does not require equal population variances.
Example: compare mean exam scores between students taught with method A and a different group taught with method B.
Do not use it for: repeated measurements on the same participants. Those observations are paired, not independent.
2. Paired-samples t-test
The paired t-test turns each pair into a difference and analyzes those differences. Pairing can remove person-to-person variation from the comparison, but only when the pairing is real and correctly recorded.
Example: blood pressure measured before treatment and again after treatment in the same patients.
Do not use it for: two unrelated samples just because they happen to have the same sample size.
3. One-way ANOVA
One-way ANOVA tests a global null hypothesis about the group means. A significant F test tells you that the equal-means model is not adequate, but it does not tell you which groups differ. That requires follow-up comparisons suited to the analysis plan.
Example: compare mean reaction times across four interface designs assigned to separate participant groups.
Do not use it for: repeated measurements from the same people unless you use a method that models the repeated structure.
4. Chi-square test of independence
The chi-square test of independence compares observed cell counts with the counts expected if the row and column variables were independent. The usual chi-square reference distribution is an approximation, so sparse expected counts need attention.
Example: test whether preferred payment method is associated with age group in a survey sample.
Do not use it for: paired binary responses. A method such as McNemar's test is designed for matched binary data.
5. Fisher's exact test
Fisher's exact test is often taught as the small-sample counterpart to the chi-square test for a 2 by 2 table. The word "exact" refers to how the p-value is calculated under the conditional null distribution. It does not remove the need to think about how the data were sampled.
Example: compare treatment success and failure between two small treatment groups.
6. Pearson correlation test
Pearson correlation describes the direction and strength of a linear relationship. A value near zero can occur even when a strong nonlinear relationship exists, so a scatterplot should come before interpretation.
Example: measure the linear association between study hours and exam score.
Do not infer: that a significant correlation proves one variable causes the other.
7. Spearman rank correlation
Spearman correlation works with ranks. It can capture a steadily increasing or decreasing relationship even when the curve is not linear. It should not be treated as a universal replacement for Pearson just because one variable fails a normality test.
Example: assess whether higher satisfaction ranks tend to accompany higher loyalty ranks.
8. Mann-Whitney U test
The Mann-Whitney U test is useful for two independent groups when a rank-based comparison matches the scientific question. If the group distributions have the same shape and differ mainly by a location shift, a location or median interpretation may be reasonable. Without that condition, the test can respond to differences in spread or shape as well.
Example: compare an ordinal symptom score between two independent treatment groups.
9. Wilcoxon signed-rank test
The Wilcoxon signed-rank test is the paired rank-based method in this list. It uses more information than a sign test because the magnitudes of the differences contribute through their ranks. It is not simply "the paired t-test for non-normal data." The null and assumptions should match the question you want to answer.
Example: compare before-and-after pain scores from the same patients when a signed-rank analysis is appropriate.
10. Kruskal-Wallis test
Kruskal-Wallis extends the independent rank comparison to more than two groups. A significant result indicates that the group distributions are not all alike in the way assessed by the ranks. Calling it a test of medians requires extra shape assumptions that are often left unstated.
Example: compare an ordinal service-quality score across four independent store formats.
Which statistical test should I use?
Use the table below to narrow the candidates. Then check the linked test guide before analyzing your data.
| Research situation | Candidate test | Why |
|---|---|---|
| Compare means from 2 independent groups | Welch two-sample t-test | Mean-based comparison without assuming equal variances |
| Compare 2 paired mean measurements | Paired t-test | Analyzes within-pair differences |
| Compare means across 3+ independent groups | One-way ANOVA | Global mean comparison for one factor |
| Test association between 2 categorical variables | Chi-square independence | Compares observed and expected contingency-table counts |
| Small or sparse 2 by 2 categorical table | Fisher's exact test | Exact conditional inference rather than large-sample chi-square approximation |
| Measure linear association between 2 quantitative variables | Pearson correlation | Targets linear correlation |
| Measure monotonic rank association | Spearman correlation | Works with ranks and monotonic relationships |
| Compare 2 independent groups using ranks | Mann-Whitney U | Rank-based distributional comparison |
| Compare 2 paired measurements using signed ranks | Wilcoxon signed-rank | Uses signs and ranks of paired differences |
| Compare 3+ independent groups using ranks | Kruskal-Wallis | Rank-based multi-group comparison |
Text decision tree for common test-selection problems
This tree covers common introductory situations. Regression, repeated-measures models, mixed models, survival analysis, count models, clustered data, complex surveys, missing-data methods, and causal analyses require a wider framework.
Important test comparisons
| Comparison | First method | Second method |
|---|---|---|
| Independent vs paired t-test | Different people or units in the two groups | Linked measurements; analyzes one difference per pair |
| T-test vs ANOVA | Usually 2-group mean comparison | Global mean comparison across 3+ groups |
| Pearson vs Spearman | Linear association in original quantitative values | Monotonic association in ranks |
| Chi-square vs Fisher exact | Large-sample reference approximation | Exact conditional calculation, often useful for small 2 by 2 tables |
| Independent t-test vs Mann-Whitney | Mean-based inferential target | Rank-distribution target; not automatically a median test |
| Paired t-test vs Wilcoxon | Mean of paired differences | Signed ranks of paired differences |
| ANOVA vs Kruskal-Wallis | Mean-based multi-group model | Rank-based multi-group distributional comparison |
How to interpret a test result without stopping at p < 0.05
A significance test answers a narrow question under a model. It does not tell you whether the effect is large, useful, clinically important, economically important, or causal. A small p-value can accompany a trivial effect in a large dataset, while an imprecise study can miss an effect that matters.
Report more than a threshold
Examples of useful effect measures include a raw mean difference, standardized mean difference, correlation coefficient, odds ratio, risk difference, Cramér's V, and rank-based effect measures. The right choice depends on the question and design.
Common statistical test selection mistakes
Choosing the test before defining the question
A dataset can support many tests. The test must match the specific inferential target.
Using a normality test as an automatic switch
Distribution shape, outliers, sample size, robustness, study design, and estimand matter too.
Ignoring pairing
Treating repeated or matched observations as independent can give the wrong standard error and p-value.
Running many pairwise t-tests across several groups
Use a planned multi-group strategy and address multiple comparisons.
Calling every rank test a median test
Mann-Whitney and Kruskal-Wallis can respond to broader distributional differences.
Treating p < .05 as proof of importance
Report effect magnitude, uncertainty, and practical context.
Interpreting correlation as causation
Causal claims require design and assumptions beyond a correlation test.
Deleting outliers automatically
Investigate data quality and influence first. Removal needs a defensible reason.
See more common statistics mistakes.
When a simple statistical test is not enough
The tests in this guide are useful teaching tools and cover many introductory designs. Real studies can be more complicated. You may need regression or generalized linear models when several predictors matter at once, mixed models for clustered or repeated observations, survival methods for time-to-event outcomes, or methods designed for counts and rates.
The analysis also depends on how data were collected. Randomization, sampling, measurement quality, missing data, confounding, and dependence can matter more than the name of the test. A significant result from a simple test does not repair a weak design.
Frequently asked questions
Common tests include t-tests, ANOVA, chi-square tests, correlation tests, Fisher's exact test, and rank-based procedures such as Mann-Whitney U, Wilcoxon signed-rank, and Kruskal-Wallis. The useful classification depends on the question being asked, the variable types, the study design, and the assumptions behind the method.
Start with the research question. Then identify the outcome and explanatory variables, their measurement scales, the number of groups, whether observations are independent or paired, and the assumptions relevant to the candidate test. A decision chart can narrow the options, but it cannot replace checking the study design and inferential target.
For a continuous outcome and two independent groups, a two-sample t procedure is often considered when mean-based inference is appropriate. Welch's t-test is commonly preferred when equal variances are not well justified. Mann-Whitney U addresses a different rank-based distributional question and should not be chosen only because a normality test is significant.
Before-and-after measurements on the same people are paired. A paired t-test is used for inference about the mean paired difference under its assumptions. A Wilcoxon signed-rank test is a rank-based option when its assumptions and inferential target fit the paired differences.
A two-sample t-test compares two group means. One-way ANOVA tests whether a set of three or more group means are all equal under the model. A significant ANOVA result does not identify which groups differ, so planned contrasts or post hoc comparisons may be needed.
A chi-square test of independence uses a large-sample approximation for contingency tables. Fisher's exact test computes an exact conditional p-value and is especially useful for small 2 by 2 tables or when expected counts make the chi-square approximation questionable. The sampling design still matters.
No. Rank-based and other nonparametric methods still have assumptions about issues such as independence, pairing, measurement scale, exchangeability, or distributional shape, depending on the procedure and the interpretation you want to make.
No. A small p-value is not a measure of effect size or practical importance. Interpret the estimate, an appropriate effect size, its confidence interval, the study design, and the domain context alongside the test result.
Key takeaways
Define the question first. Then identify the variables, groups, pairing, and study design. Check the assumptions of the methods that fit that structure. Run the test only after you know what its null hypothesis and estimand mean. Finally, interpret the estimate, effect size, confidence interval, and study limitations alongside the p-value.
Sources and further reading
These references were used to check the methodological statements in this guide. They are listed for readers who want the formal details behind the introductory explanations.
- American Statistical Association: Statement on Statistical Significance and P-Values
- Penn State STAT 200: Elementary Statistics
- NIST/SEMATECH e-Handbook: Two-Sample t-Test for Equal Means
- Delacre, Lakens & Leys (2017): Why use Welch's t-test instead of Student's t-test by default
- Penn State STAT 504: Fisher's Exact Test
- BMJ: Mann-Whitney test is not just a test of medians
- Penn State STAT 415: The Wilcoxon Tests
- NIST Dataplot: Kruskal-Wallis Test