Study Tips Hypothesis Testing Beginner Guide 24 min read October 1, 2026
BY: Statistics Fundamentals Team
Reviewed By: Minsa A (Senior Statistics Editor)

10 Types of Statistical Tests and When to Use Each

Choosing a statistical test is easier when you stop treating test names as a memorization exercise. Start with the question you want the data to answer. Then look at the outcome variable, the grouping or explanatory variable, the number of groups, whether observations are independent or paired, and the assumptions behind the candidate method.

This guide covers 10 common statistical tests and shows the kind of problem each one is designed to address. It also explains the cases where a familiar shortcut can send you in the wrong direction. Use the tables to narrow your options, then read the relevant test section before running the analysis.

What You'll Learn
  • ✓ How the research question changes test selection
  • ✓ Why variable type, group count, and pairing matter
  • ✓ When each of 10 common statistical tests is useful
  • ✓ Which assumptions deserve attention before interpreting a p-value
  • ✓ How t-tests, ANOVA, correlation tests, and rank-based tests differ
  • ✓ Why effect sizes and confidence intervals belong beside significance tests

Quick answer: how to choose a statistical test

Start with the question, not the software menu

Ask what you want to compare or associate, identify the outcome and explanatory variables, count the groups, decide whether observations are independent or paired, and then check the assumptions of the methods that remain. A decision tree can narrow the field, but it cannot decide what your scientific question means.

Fast orientation
  • Two independent group means: consider a two-sample t procedure, often Welch's t-test.
  • Two paired measurements: consider a paired t-test when inference about the mean difference is appropriate.
  • Three or more independent group means: one-way ANOVA is a common starting point.
  • Two categorical variables: chi-square test of independence, or Fisher's exact test for small tables when appropriate.
  • Two quantitative variables: Pearson for linear association; Spearman for monotonic rank association.
  • Rank-based group comparisons: Mann-Whitney U, Wilcoxon signed-rank, or Kruskal-Wallis, depending on the design.

Statistical tests at a glance

The table is a screening tool, not a substitute for reading the assumptions. The same dataset can require different methods for different questions.

TestMain questionTypical data/designWhat to check
Independent-samples t-testDo two independent group means differ?Continuous outcome, 2 independent groupsIndependence, model for means, outliers; variance assumption depends on version
Paired-samples t-testIs the mean paired difference zero?Continuous paired differencesCorrect pairing, independence across pairs, distribution of differences
One-way ANOVAAre all group means equal?Continuous outcome, 3+ independent groupsIndependence, residual behavior, variance structure
Chi-square test of independenceAre two categorical variables associated?Counts in a contingency tableIndependent observations, expected-count adequacy for approximation
Fisher's exact testAre two categorical variables associated in a small table?Usually a 2 by 2 count tableSampling design and fixed-margin interpretation
Pearson correlation testIs there a linear association?Two quantitative variablesLinearity, independence, influential points, assumptions for inference
Spearman rank correlationIs there a monotonic rank association?Ordinal or quantitative variablesIndependent pairs, meaningful ordering, monotonic pattern
Mann-Whitney U testDo two independent groups differ in rank distribution?Ordinal/continuous outcome, 2 independent groupsIndependence; interpretation depends on distributional relationship
Wilcoxon signed-rank testAre paired differences centered around zero in the signed-rank sense?Paired ordinal/continuous measurementsPaired structure, independent pairs, symmetry for common location-shift interpretation
Kruskal-Wallis testDo 3+ independent groups differ in rank distribution?Ordinal/continuous outcome, 3+ independent groupsIndependent groups; shape assumptions matter for location interpretation

What is a statistical test?

Definition
A statistical test is a formal procedure for evaluating how compatible sample data are with a specified null hypothesis or statistical model.

Null and alternative hypotheses

The null hypothesis is the reference claim used to calculate the test statistic and p-value. The alternative describes the competing values or patterns considered by the test. The exact null is test-specific. It might state equality of means, independence of categorical variables, zero correlation, or a specified parameter value.

Test statistic, p-value, and significance level

A test statistic summarizes how the observed data compare with what the null model would lead us to expect. A p-value is then calculated under that null model. It describes how incompatible the observed test result, or a result at least as extreme according to the chosen statistic, is with the null. It is not the probability that the null hypothesis is true.

The significance level, usually written as α, is a decision threshold chosen before interpreting the test. Values such as 0.05 or 0.01 are conventions, not universal scientific rules. A Type I error means rejecting a true null hypothesis; a Type II error means failing to reject a false null hypothesis. Statistical power is the probability of rejecting the null when a specified alternative is true.

How to choose the right statistical test

Test selection works better as a sequence of questions than as a lookup based on one variable. Before the steps, identify what each variable represents.

Identify your variables before choosing a test

Variable typeExamplesWhy it matters
ContinuousHeight, income, blood pressure, response timeCan support mean-based models, correlation, regression, and rank-based analyses depending on the question.
OrdinalSatisfaction ratings, ordered severity categoriesThe categories have an order, but equal numerical spacing may not be defensible.
Nominal/categoricalTreatment group, product choice, disease categoryOften analyzed through counts, proportions, contingency tables, or categorical regression models.
BinaryYes/no, success/failure, event/no eventA two-category outcome often calls for proportion methods or binary-response models rather than a mean test.

The outcome variable is the response you want to analyze. The explanatory or grouping variable defines groups or helps explain variation. For the question "Do mean exam scores differ between teaching methods?", exam score is the outcome and teaching method is the grouping variable.

1

Write the research question in statistical terms

Are you comparing means, comparing distributions, testing association between categories, measuring correlation, or estimating a parameter against a reference value?

2

Identify the outcome and explanatory variables

Continuous, ordinal, nominal, and binary variables support different summaries and models. Measurement scale matters, but it is only one part of the decision.

3

Count the groups or measurement occasions

Two groups and four groups are different analysis problems. Repeatedly running pairwise t-tests across many groups inflates the chance of false positives unless multiplicity is handled.

4

Decide whether observations are independent or paired

Before-and-after observations from the same person are paired. Measurements from separate treatment and control participants are usually independent. The analysis must respect that dependence structure.

5

Check assumptions in the context of the chosen model

Look at outliers, distributional shape, variance structure, sample size, dependence, and the exact inferential target. Do not use a preliminary normality test as an automatic switch between parametric and nonparametric procedures.

6

Plan interpretation before running the test

Decide what estimate, effect size, confidence interval, and diagnostic information you will report with the p-value.

Parametric vs nonparametric tests

Parametric methods specify a statistical model through parameters and structural or distributional assumptions. A t-test, for example, is built around a model for means and standard errors. Nonparametric and rank-based methods use a different set of assumptions and targets. They are not assumption-free, and they are not automatically the right answer whenever data look non-normal.

FeatureParametric methodsRank-based/nonparametric methods
Typical targetParameters such as means, regression coefficients, correlationsRanks, distributions, stochastic ordering, or other nonparametric features
AssumptionsModel-specific distributional and structural assumptionsDifferent or fewer distributional assumptions, but still assumptions
Common examplest-tests, ANOVA, Pearson correlationMann-Whitney, Wilcoxon signed-rank, Kruskal-Wallis, Spearman
Selection ruleChoose based on the question, design, estimand, data behavior, and assumptions. Do not use a single normality p-value as the rule.

Read the full parametric vs nonparametric comparison.

1. Independent-samples t-test

Two independent groups
Question answered: Is the difference between two population means compatible with the null value, usually zero?
OutcomeContinuous quantitative variable
DesignTwo independent groups
Typical statistict statistic based on the mean difference and its standard error
Useful effect estimateMean difference with a confidence interval; standardized difference when useful

Use a two-sample t procedure when the scientific question is about a difference in means between two independent groups. In many routine analyses, Welch's t-test is a safer default than the pooled equal-variance version because it does not require equal population variances.

Example: compare mean exam scores between students taught with method A and a different group taught with method B.

Do not use it for: repeated measurements on the same participants. Those observations are paired, not independent.

2. Paired-samples t-test

Two linked measurements
Question answered: Is the population mean of the paired differences equal to a specified value, usually zero?
OutcomeContinuous difference scores
DesignSame people twice, matched pairs, or other one-to-one pairing
What is analyzedOne difference value per pair
Key checkThe distributional assumption concerns the paired differences, not each raw measurement separately

The paired t-test turns each pair into a difference and analyzes those differences. Pairing can remove person-to-person variation from the comparison, but only when the pairing is real and correctly recorded.

Example: blood pressure measured before treatment and again after treatment in the same patients.

Do not use it for: two unrelated samples just because they happen to have the same sample size.

3. One-way ANOVA

Three or more independent groups
Question answered: Are all population means equal across the groups in a one-factor design?
OutcomeContinuous quantitative variable
Grouping variableOne categorical factor with 3 or more independent groups
Test statisticF statistic comparing between-group to within-group variation
After a significant resultUse planned contrasts or post hoc procedures to locate differences

One-way ANOVA tests a global null hypothesis about the group means. A significant F test tells you that the equal-means model is not adequate, but it does not tell you which groups differ. That requires follow-up comparisons suited to the analysis plan.

Example: compare mean reaction times across four interface designs assigned to separate participant groups.

Do not use it for: repeated measurements from the same people unless you use a method that models the repeated structure.

4. Chi-square test of independence

Categorical association
Question answered: Are two categorical variables independent in the population?
DataCounts in a contingency table
Null modelExpected counts are calculated under independence
Statisticχ² = Σ (O - E)² / E
Effect sizeCramér's V is one common association measure

The chi-square test of independence compares observed cell counts with the counts expected if the row and column variables were independent. The usual chi-square reference distribution is an approximation, so sparse expected counts need attention.

Example: test whether preferred payment method is associated with age group in a survey sample.

Do not use it for: paired binary responses. A method such as McNemar's test is designed for matched binary data.

5. Fisher's exact test

Small categorical tables
Question answered: For a contingency table under an exact conditional framework, how surprising is the observed association under the null?
DataCategorical counts, most commonly a 2 by 2 table
Why use itIt avoids relying on the large-sample chi-square approximation
Common situationSmall samples or sparse expected cell counts
Report with itA measure such as an odds ratio and its interval when appropriate

Fisher's exact test is often taught as the small-sample counterpart to the chi-square test for a 2 by 2 table. The word "exact" refers to how the p-value is calculated under the conditional null distribution. It does not remove the need to think about how the data were sampled.

Example: compare treatment success and failure between two small treatment groups.

6. Pearson correlation test

Linear association
Question answered: Is the population linear correlation between two quantitative variables equal to a specified value, often zero?
VariablesTwo quantitative variables measured on the same observational units
CoefficientPearson's r ranges from -1 to +1
Pattern measuredLinear association
Important diagnosticInspect a scatterplot for curvature and influential points

Pearson correlation describes the direction and strength of a linear relationship. A value near zero can occur even when a strong nonlinear relationship exists, so a scatterplot should come before interpretation.

Example: measure the linear association between study hours and exam score.

Do not infer: that a significant correlation proves one variable causes the other.

7. Spearman rank correlation

Monotonic rank association
Question answered: Is there a monotonic association between the ranks of two variables?
VariablesOrdinal or quantitative variables that can be meaningfully ranked
CoefficientSpearman's rank correlation, often written ρs or rs
Pattern measuredMonotonic rather than specifically linear association
Useful whenRanks are natural, or the relationship is monotonic but not well described by a straight line

Spearman correlation works with ranks. It can capture a steadily increasing or decreasing relationship even when the curve is not linear. It should not be treated as a universal replacement for Pearson just because one variable fails a normality test.

Example: assess whether higher satisfaction ranks tend to accompany higher loyalty ranks.

8. Mann-Whitney U test

Two independent rank distributions
Question answered: Do observations from two independent groups show the same rank distribution under the null model?
OutcomeOrdinal or continuous values that can be ranked
DesignTwo independent groups
What it usesRanks of pooled observations
Interpretation warningIt is not automatically a test of medians

The Mann-Whitney U test is useful for two independent groups when a rank-based comparison matches the scientific question. If the group distributions have the same shape and differ mainly by a location shift, a location or median interpretation may be reasonable. Without that condition, the test can respond to differences in spread or shape as well.

Example: compare an ordinal symptom score between two independent treatment groups.

9. Wilcoxon signed-rank test

Paired rank comparison
Question answered: Are paired differences centered around zero according to their signed ranks?
OutcomePaired ordinal or continuous measurements with rankable differences
DesignMatched pairs or repeated measurements
What it usesSigns and ranks of the absolute paired differences
Common location interpretationUsually relies on symmetry of the difference distribution

The Wilcoxon signed-rank test is the paired rank-based method in this list. It uses more information than a sign test because the magnitudes of the differences contribute through their ranks. It is not simply "the paired t-test for non-normal data." The null and assumptions should match the question you want to answer.

Example: compare before-and-after pain scores from the same patients when a signed-rank analysis is appropriate.

10. Kruskal-Wallis test

Three or more independent rank distributions
Question answered: Do three or more independent groups have the same rank distribution under the null model?
OutcomeOrdinal or continuous values that can be ranked
DesignThree or more independent groups
StatisticH statistic based on group rank sums or mean ranks
After rejectionUse suitable post hoc rank comparisons with multiplicity control

Kruskal-Wallis extends the independent rank comparison to more than two groups. A significant result indicates that the group distributions are not all alike in the way assessed by the ranks. Calling it a test of medians requires extra shape assumptions that are often left unstated.

Example: compare an ordinal service-quality score across four independent store formats.

Which statistical test should I use?

Use the table below to narrow the candidates. Then check the linked test guide before analyzing your data.

Research situationCandidate testWhy
Compare means from 2 independent groupsWelch two-sample t-testMean-based comparison without assuming equal variances
Compare 2 paired mean measurementsPaired t-testAnalyzes within-pair differences
Compare means across 3+ independent groupsOne-way ANOVAGlobal mean comparison for one factor
Test association between 2 categorical variablesChi-square independenceCompares observed and expected contingency-table counts
Small or sparse 2 by 2 categorical tableFisher's exact testExact conditional inference rather than large-sample chi-square approximation
Measure linear association between 2 quantitative variablesPearson correlationTargets linear correlation
Measure monotonic rank associationSpearman correlationWorks with ranks and monotonic relationships
Compare 2 independent groups using ranksMann-Whitney URank-based distributional comparison
Compare 2 paired measurements using signed ranksWilcoxon signed-rankUses signs and ranks of paired differences
Compare 3+ independent groups using ranksKruskal-WallisRank-based multi-group comparison

Text decision tree for common test-selection problems

Comparing a continuous outcome across groups?
→
2 independent groups: t procedure; 2 paired: paired t; 3+ independent: ANOVA
Comparing rankable outcomes with a rank-based target?
→
2 independent: Mann-Whitney; 2 paired: Wilcoxon signed-rank; 3+ independent: Kruskal-Wallis
Testing association between categorical variables?
→
Chi-square; consider Fisher's exact for small or sparse tables
Testing association between two quantitative or ordered variables?
→
Pearson for linear association; Spearman for monotonic rank association
Decision-tree limit

This tree covers common introductory situations. Regression, repeated-measures models, mixed models, survival analysis, count models, clustered data, complex surveys, missing-data methods, and causal analyses require a wider framework.

Important test comparisons

ComparisonFirst methodSecond method
Independent vs paired t-testDifferent people or units in the two groupsLinked measurements; analyzes one difference per pair
T-test vs ANOVAUsually 2-group mean comparisonGlobal mean comparison across 3+ groups
Pearson vs SpearmanLinear association in original quantitative valuesMonotonic association in ranks
Chi-square vs Fisher exactLarge-sample reference approximationExact conditional calculation, often useful for small 2 by 2 tables
Independent t-test vs Mann-WhitneyMean-based inferential targetRank-distribution target; not automatically a median test
Paired t-test vs WilcoxonMean of paired differencesSigned ranks of paired differences
ANOVA vs Kruskal-WallisMean-based multi-group modelRank-based multi-group distributional comparison

How to interpret a test result without stopping at p < 0.05

A significance test answers a narrow question under a model. It does not tell you whether the effect is large, useful, clinically important, economically important, or causal. A small p-value can accompany a trivial effect in a large dataset, while an imprecise study can miss an effect that matters.

Report more than a threshold

The estimate that answers the scientific question
An effect size suited to the design
A confidence interval or other uncertainty interval
The test statistic, degrees of freedom where relevant, and p-value
Assumption checks and important diagnostics
Study-design limits, including confounding or non-random sampling

Examples of useful effect measures include a raw mean difference, standardized mean difference, correlation coefficient, odds ratio, risk difference, Cramér's V, and rank-based effect measures. The right choice depends on the question and design.

Common statistical test selection mistakes

Mistake 1

Choosing the test before defining the question

A dataset can support many tests. The test must match the specific inferential target.

Mistake 2

Using a normality test as an automatic switch

Distribution shape, outliers, sample size, robustness, study design, and estimand matter too.

Mistake 3

Ignoring pairing

Treating repeated or matched observations as independent can give the wrong standard error and p-value.

Mistake 4

Running many pairwise t-tests across several groups

Use a planned multi-group strategy and address multiple comparisons.

Mistake 5

Calling every rank test a median test

Mann-Whitney and Kruskal-Wallis can respond to broader distributional differences.

Mistake 6

Treating p < .05 as proof of importance

Report effect magnitude, uncertainty, and practical context.

Mistake 7

Interpreting correlation as causation

Causal claims require design and assumptions beyond a correlation test.

Mistake 8

Deleting outliers automatically

Investigate data quality and influence first. Removal needs a defensible reason.

See more common statistics mistakes.

When a simple statistical test is not enough

The tests in this guide are useful teaching tools and cover many introductory designs. Real studies can be more complicated. You may need regression or generalized linear models when several predictors matter at once, mixed models for clustered or repeated observations, survival methods for time-to-event outcomes, or methods designed for counts and rates.

The analysis also depends on how data were collected. Randomization, sampling, measurement quality, missing data, confounding, and dependence can matter more than the name of the test. A significant result from a simple test does not repair a weak design.

Frequently asked questions

Common tests include t-tests, ANOVA, chi-square tests, correlation tests, Fisher's exact test, and rank-based procedures such as Mann-Whitney U, Wilcoxon signed-rank, and Kruskal-Wallis. The useful classification depends on the question being asked, the variable types, the study design, and the assumptions behind the method.

Start with the research question. Then identify the outcome and explanatory variables, their measurement scales, the number of groups, whether observations are independent or paired, and the assumptions relevant to the candidate test. A decision chart can narrow the options, but it cannot replace checking the study design and inferential target.

For a continuous outcome and two independent groups, a two-sample t procedure is often considered when mean-based inference is appropriate. Welch's t-test is commonly preferred when equal variances are not well justified. Mann-Whitney U addresses a different rank-based distributional question and should not be chosen only because a normality test is significant.

Before-and-after measurements on the same people are paired. A paired t-test is used for inference about the mean paired difference under its assumptions. A Wilcoxon signed-rank test is a rank-based option when its assumptions and inferential target fit the paired differences.

A two-sample t-test compares two group means. One-way ANOVA tests whether a set of three or more group means are all equal under the model. A significant ANOVA result does not identify which groups differ, so planned contrasts or post hoc comparisons may be needed.

A chi-square test of independence uses a large-sample approximation for contingency tables. Fisher's exact test computes an exact conditional p-value and is especially useful for small 2 by 2 tables or when expected counts make the chi-square approximation questionable. The sampling design still matters.

No. Rank-based and other nonparametric methods still have assumptions about issues such as independence, pairing, measurement scale, exchangeability, or distributional shape, depending on the procedure and the interpretation you want to make.

No. A small p-value is not a measure of effect size or practical importance. Interpret the estimate, an appropriate effect size, its confidence interval, the study design, and the domain context alongside the test result.

Key takeaways

A practical way to remember the page

Define the question first. Then identify the variables, groups, pairing, and study design. Check the assumptions of the methods that fit that structure. Run the test only after you know what its null hypothesis and estimand mean. Finally, interpret the estimate, effect size, confidence interval, and study limitations alongside the p-value.

Sources and further reading

These references were used to check the methodological statements in this guide. They are listed for readers who want the formal details behind the introductory explanations.