What Is a Chi-Square Goodness-of-Fit Test?

Definition — Chi-Square Goodness-of-Fit Test

A chi-square goodness-of-fit test is a hypothesis test used to determine whether observed frequencies across the categories of one categorical variable are consistent with a specified set of expected proportions. It asks whether the discrepancies between what was observed and what was expected under a null hypothesis are larger than random sampling variation alone can reasonably explain.

χ² = Σ[(Oᵢ − Eᵢ)² / Eᵢ]

The key phrase is one categorical variable. The test applies when you have a single variable with two or more mutually exclusive categories, a set of observed counts, and a hypothesized distribution that specifies the probability of each category. You might be asking whether four product preferences are equally popular, whether a die produces each face with probability 1/6, or whether customer channel choices match last year's distribution.

What the test does not ask is whether the observed counts are exactly equal to expected counts. They almost never will be, because sampling naturally produces some scatter. The statistical question is whether the size of the discrepancies is too large to attribute to chance under the null distribution.

💡
The Central Principle

The chi-square goodness-of-fit test measures how far observed category counts depart from a specified expected distribution, relative to the counts expected under the null hypothesis, and uses that discrepancy to evaluate whether the data are compatible with the proposed distribution.

When Should You Use a Chi-Square Goodness-of-Fit Test?

The goodness-of-fit test fits situations where you have one categorical variable and a specific null distribution to test against. Common scenarios include:

  • Testing whether survey responses are equally distributed across categories
  • Checking whether a die or spinner appears fair
  • Comparing observed genetic ratios against a Mendelian expected ratio (e.g., 3:1)
  • Verifying whether defect types follow a historically specified distribution
  • Testing whether website traffic sources match last year's proportions
  • Evaluating whether customer preference data fits a theoretically specified distribution

A key distinction: goodness of fit involves one categorical variable. If you want to test whether two categorical variables are related to each other, that calls for a chi-square test of independence, which uses a different data structure and different formulas. If you are unsure which test applies to your situation, the statistical test selector can help.

The goodness-of-fit test also does not require that the expected proportions be equal across categories. You can test against any completely specified null distribution, including unequal proportions derived from theory, prior data, or regulatory standards.

Chi-Square Goodness-of-Fit Hypotheses

Null Hypothesis

For a variable with k categories and hypothesized probabilities p₁₀, p₂₀, ..., pₖ₀, the null hypothesis states:

Null Hypothesis (H₀)
H₀: p₁ = p₁₀, p₂ = p₂₀, ..., pₖ = pₖ₀
pᵢ = true population probability for category i pᵢ₀ = hypothesized probability from the null model

In plain terms: the population distribution of the categorical variable follows the specified null distribution exactly. For the equal-proportions case with four categories, this becomes H₀: p₁ = p₂ = p₃ = p₄ = 0.25.

Notice that the null probabilities must collectively describe a complete probability distribution. Their sum must equal 1. If you are working from a theory that does not provide all category probabilities, the null hypothesis is incompletely specified and cannot be tested as written.

Alternative Hypothesis

The alternative hypothesis states that the population distribution does not match the specified null distribution:

Alternative Hypothesis (H₁)
H₁: at least one pᵢ ≠ pᵢ₀

This matters: the alternative does not claim that every category differs from its expected proportion. It claims only that the probability vector does not match the null specification in at least one respect. A significant omnibus result tells you the observed distribution is inconsistent with the null model; it does not automatically identify which specific categories drove that inconsistency. That requires follow-up analysis using residuals.

Chi-Square Goodness-of-Fit Formula

Pearson Chi-Square Goodness-of-Fit Formula
χ² = Σ[(Oᵢ − Eᵢ)² / Eᵢ]
χ² = chi-square test statistic Oᵢ = observed frequency for category i Eᵢ = expected frequency for category i Σ = sum across all categories

Each category contributes one term, (Oᵢ − Eᵢ)² / Eᵢ, to the total statistic. These contributions are summed across all categories to produce the single chi-square value that is then compared against a reference distribution.

Why Are Differences Squared?

If you simply summed the raw differences, positive and negative deviations would cancel out. A category with 5 more observations than expected and a category with 5 fewer would contribute nothing to the total, even though both represent genuine departures from the null model.

Squaring the differences solves this: all contributions become non-negative, so the statistic accumulates evidence of discrepancy regardless of direction. As a further consequence, larger departures contribute disproportionately more than smaller ones, which makes sense because a category that is 10 counts off from expected provides stronger evidence against the null than one that is 2 counts off.

Division by the expected frequency Eᵢ then scales each squared difference relative to the magnitude of counts expected in that category. A difference of 10 from an expected count of 200 is less remarkable than a difference of 10 from an expected count of 20. The division handles that rescaling automatically.

How to Calculate Expected Frequencies

When the null hypothesis fully specifies the probability of each category, the expected frequency for category i is:

Expected Frequency Formula
Eᵢ = n × pᵢ₀
n = total sample size pᵢ₀ = null probability for category i

Two arithmetic checks confirm your expected counts are set up correctly. First, the null probabilities must sum to 1: Σpᵢ₀ = 1. If they do not, the null distribution is invalid. Second, because the formula multiplies each probability by the same total n, the expected counts must also sum to n: ΣEᵢ = n. Verifying both before computing χ² catches setup errors before they propagate into the result.

Expected proportions need not be equal. Goodness of fit applies equally well to a null distribution that assigns 40% to category A, 30% to B, 20% to C, and 10% to D. The requirement is only that the proportions are specified in advance from theory, prior data, or some other justified source, not derived from the observed data you are about to test.

⚠️
Critical Point About Expected Counts

Expected counts must come from the null hypothesis model, not from the observed sample. If you calculate expected proportions to match the data exactly, the test becomes circular and meaningless. The null distribution must be independently justified before you collect or examine the data.

How to Perform the Test Step by Step

1

State H₀ and H₁

Write out the null hypothesis specifying all k category probabilities (p₁₀, p₂₀, ..., pₖ₀) and confirm they sum to 1. State the alternative: the population distribution differs from H₀ in at least one category.

2

Record observed category counts

List the observed count Oᵢ for each category. Verify ΣOᵢ = n. The test requires actual frequency counts, not percentages or proportions by themselves.

3

Calculate expected counts

Apply Eᵢ = n × pᵢ₀ to each category. Verify ΣEᵢ = n. This confirms both that your probabilities sum to 1 and that your arithmetic is consistent.

4

Check approximation conditions

Before computing, inspect the expected counts. If any are very small, the chi-square approximation may be unreliable. See the assumptions section for guidance on what to do in that situation.

5

Compute each category's chi-square contribution

For each category, calculate (Oᵢ − Eᵢ)² / Eᵢ. Build a contribution table to keep the arithmetic organized and to make it easy to identify which categories depart most from the null model.

6

Sum contributions to obtain χ²

Add the contributions from all categories: χ² = Σ[(Oᵢ − Eᵢ)² / Eᵢ]. This single number is the test statistic that carries all the evidence about distributional fit.

7

Determine degrees of freedom

For a standard test with k categories and all null probabilities specified externally: df = k − 1. If you estimated parameters from the same data, adjust accordingly (see the df section).

8

Find the p-value or critical value

Look up the p-value from a chi-square distribution with df degrees of freedom, using the upper tail. Alternatively, compare the observed χ² against the critical value from a chi-square table for your chosen significance level.

9

Make the statistical decision

If p ≤ α, reject H₀. If p > α, fail to reject H₀. Use a pre-specified significance level (commonly α = 0.05) chosen before examining the data.

10

Interpret the result in context

State what the result means for the specific research question. A significant χ² means the observed distribution is inconsistent with the null model; it does not tell you which categories differ or by how much. For that you need residual analysis.

Worked Example: Equal Expected Proportions

This is the foundational example from which the test's mechanics are easiest to see. Four response categories, equal null probabilities, n = 100.

Worked Example 1 — Equal Proportions

Scenario: 100 participants each chose one of four product options (A, B, C, D). If preferences are equally distributed, each option should attract 25 responses. The observed counts were 18, 22, 27, and 33. Is there evidence that preferences are not equally distributed? Use α = 0.05.

Step 1 — State the Hypotheses

H₀

H₀: p_A = p_B = p_C = p_D = 0.25. Product preferences are equally distributed across the four options.

H₁

H₁: At least one pᵢ ≠ 0.25. Preferences are not equally distributed.

Step 2 — Calculate Expected Counts

Eᵢ = n × pᵢ₀ = 100 × 0.25 = 25 for each category. Check: ΣEᵢ = 100 = n ✓

Step 3 — Contribution Table

CategoryObserved (O)Expected (E)O − E(O − E)²(O − E)² / E
A1825−7491.96
B2225−390.36
C2725+240.16
D3325+8642.56
Total1001000—χ² = 5.04
4

Total χ²: 1.96 + 0.36 + 0.16 + 2.56 = 5.04

5

Degrees of freedom: df = k − 1 = 4 − 1 = 3

6

p-value: P(χ²₃ ≥ 5.04) = 0.169. The chi-square critical value at α = 0.05 with df = 3 is 7.815. Because 5.04 < 7.815, the statistic does not fall in the rejection region. You can verify this with the chi-square table or by learning how to read a chi-square table.

✅ Decision: Fail to reject H₀ (p = 0.169 > 0.05). The data do not provide sufficient evidence to reject the hypothesis that the four options are equally preferred. This does not prove preferences are exactly equal — it means the observed pattern is consistent with what equal preferences would produce.

Example With Unequal Expected Proportions

This example addresses one of the most common misconceptions about the test: that expected proportions must be equal. They do not. The null distribution can assign any probabilities to the categories as long as they sum to 1 and are specified from a justified source.

Worked Example 2 — Unequal Proportions (Significant Result)

Scenario: A retailer's historical sales split across four product lines is A = 40%, B = 30%, C = 20%, D = 10%. A new sample of n = 200 transactions shows counts of 95, 48, 32, 25. Has the sales mix shifted? Use α = 0.05.

H₀

H₀: p_A = 0.40, p_B = 0.30, p_C = 0.20, p_D = 0.10. The current sales distribution matches historical proportions.

LineNull pObserved (O)Expected (E)O − E(O − E)²(O − E)² / E
A0.409580+152252.8125
B0.304860−121442.4000
C0.203240−8641.6000
D0.102520+5251.2500
Total2002000—χ² = 8.0625
→

df = k − 1 = 3, χ² = 8.063, p = 0.045 — just below the 0.05 threshold.

✅ Decision: Reject H₀ (p = 0.045 ≤ 0.05). There is statistically significant evidence that the current sales distribution differs from the historical proportions. Product line A accounts for the largest contribution (2.81), followed by B (2.40) — both categories show the most departure from their expected counts. However, these contributions are not individually adjusted for multiple comparisons; further analysis using residuals is appropriate before drawing category-level conclusions.

Interpreting a Non-Significant Result

When p > α, the correct conclusion is that the data do not provide sufficient evidence to reject the specified null distribution at the chosen significance level. This is not the same as proving the null distribution is correct.

A non-significant result means the observed discrepancies are small enough to be plausibly explained by random sampling variation under H₀. It does not rule out the possibility that a real but smaller distributional difference exists that the study lacked the power to detect. As sample sizes grow, smaller and smaller departures from the null distribution become detectable, so a non-significant result in a small study carries less information than the same result in a large one.

⚠️
Never Say "We Proved the Distribution Is Correct"

Failing to reject H₀ means the data are consistent with the null distribution. Consistency is not proof of truth. A different model might also be consistent with the same data. Absence of evidence against H₀ is not evidence of absence of a real difference.

How to Calculate Degrees of Freedom

For a standard goodness-of-fit test where all k null category probabilities are specified in advance and no parameters are estimated from the data being analyzed:

Degrees of Freedom
df = k − 1
k = number of categories

The intuition: once the total n is fixed and you know the counts in k − 1 categories, the count in the final category is determined. That constraint costs one degree of freedom.

When df Is Not Simply k − 1

The k − 1 rule applies when the null probabilities are given externally without using the analyzed data to estimate any parameters. If expected probabilities require estimating distribution parameters from the same dataset, degrees of freedom are reduced by one for each independently estimated parameter. In that setting, a common framework is:

Adjusted Degrees of Freedom (Parameter Estimation Case)
df = k − 1 − m
m = number of parameters estimated from the analyzed data

For example, if you fit a Poisson distribution to binned counts using the sample mean to estimate λ, you estimate one parameter from the data, so df = k − 1 − 1 = k − 2. This adjustment is important for accuracy; using the unadjusted df when parameters have been estimated produces a test that is too liberal. The scope for this adjustment varies by statistical context, so apply it only when it genuinely applies rather than mechanically to every problem.

Assumptions and Conditions

✓ Assumptions Checklist

  • Frequency data: The test requires actual counts, not percentages, proportions, or means. Percentages must be converted to counts (which requires knowing n) before computing χ².
  • Mutually exclusive categories: Each observation falls into exactly one category. No observation is counted in multiple categories under the standard setup.
  • Independent observations: Observations must arise independently. Do not apply the standard test to repeated measurements from the same participants as though they were independent observations.
  • Adequate expected frequencies: Expected counts should be large enough for the chi-square approximation to be reliable. See the guidance below on what "adequate" means in practice.
  • Externally specified null probabilities: The null probabilities must come from theory, prior data, or some other source independent of the sample being tested.

Expected Frequency Conditions

The accuracy of the chi-square approximation depends partly on expected counts, but common textbooks have stated this condition in different ways and with different strictness. A widely cited rule of thumb requires all expected frequencies to be at least 5. Some sources allow up to 20% of cells to fall below 5 provided none falls below 1. These are rules of thumb, not mathematical absolutes, and their practical importance depends on sample size, the number of categories, and the structure of the distribution.

The important nuance: expected-frequency conditions concern the expected counts, not the observed ones. A category may have a low observed count while its expected count is adequate, or vice versa. Checking only observed counts misapplies the assumption.

💡
What to Do When Expected Counts Are Small

If expected counts are too sparse for a reliable chi-square approximation, options include combining substantively defensible categories, using an exact multinomial test, applying Monte Carlo simulation methods, or collecting more data. Do not combine categories merely because their observed counts are low, and do not collapse them post-hoc to manufacture a significant result.

How to Interpret a Significant Result

A significant chi-square goodness-of-fit result means the observed distribution is inconsistent with the specified null distribution at the chosen significance level. It does not mean every category departs significantly from its expected proportion. The omnibus χ² statistic aggregates information from all categories, and one or two categories driving large contributions can push the total to significance even when other categories fit the null model well.

Which Categories Contributed Most?

After a significant omnibus result, examining the contribution table identifies which categories depart most from the null model. The contribution for category i is (Oᵢ − Eᵢ)² / Eᵢ — larger values signal greater departure relative to the expected count in that category.

For a directional picture, Pearson residuals add a sign:

Pearson Residual
rᵢ = (Oᵢ − Eᵢ) / √Eᵢ
rᵢ > 0 = observed exceeds expected rᵢ < 0 = observed falls below expected

For Example 1, the Pearson residuals are: A = (18 − 25)/√25 = −1.40; B = −0.60; C = +0.40; D = +1.60. Category D has the largest magnitude, showing the most departures from the equal-preference null model, though neither residual is extreme enough to suggest a single category alone is responsible for the overall pattern.

These residual magnitudes are not individually adjusted for the fact that you are examining multiple categories. If you want to conduct formal inferential comparisons for individual categories after a significant omnibus result, apply a suitable adjustment for multiple comparisons such as Bonferroni correction.

Effect Size for Goodness of Fit

Statistical significance tells you whether the data are inconsistent with the null model. It does not tell you how large the discrepancy is in practical terms. For a very large sample, even a trivially small distributional difference can produce a significant χ². Effect size addresses this by measuring magnitude independently of sample size.

Cohen's w is the standard effect size measure for chi-square goodness-of-fit tests. Using the sample relationship under the Pearson GOF framework:

Cohen's w (Effect Size)
w = √(χ² / n)

For Example 1: w = √(5.04 / 100) = √0.0504 ≈ 0.22. Cohen proposed conventional benchmarks of w = 0.10 (small), w = 0.30 (medium), and w = 0.50 (large), though what counts as meaningful depends heavily on the subject-matter context and should not be applied automatically. For more on effect size in hypothesis testing, the dedicated guide covers the broader picture.

Chi-Square Goodness of Fit vs Test of Independence

Both procedures use the same χ² formula with the same chi-square reference distribution, but they address completely different questions and apply to different data structures.

Feature Goodness of Fit Test of Independence
Number of variablesOne categorical variableTwo categorical variables
Research questionDoes the distribution match specified proportions?Are the two variables associated?
Data structureSingle list of category countsr × c contingency table
Expected count formulaEᵢ = n × pᵢ₀E = (row total × col total) / n
Degrees of freedomk − 1 (when probabilities fully specified)(r − 1)(c − 1)
Null probabilities fromTheory, prior data, external modelMarginal totals of the contingency table
Example questionAre four choices equally popular?Is product choice associated with age group?

Mixing the formulas produces wrong results. The contingency-table expected-count formula (row total × column total / n) is not appropriate for a one-variable goodness-of-fit test, and the df formula (r − 1)(c − 1) does not apply there either. If you are dealing with two categorical variables in a table, use the chi-square test of independence. The chi-square test examples page shows both procedures applied to concrete problems.

Goodness of Fit vs Homogeneity

Chi-square tests of homogeneity compare the distribution of a single categorical outcome across two or more separate populations. They use contingency-table structure (groups × categories) and compute expected counts from marginal totals, placing them closer in structure to the independence test than to goodness of fit. Computationally the formula is the same, but the research question differs: homogeneity asks whether the categorical distribution is the same across groups, while goodness of fit tests one group's distribution against a pre-specified null model.

Goodness of Fit vs Binomial Test

When there are exactly two categories, an exact binomial test is often an appropriate alternative and does not rely on the chi-square approximation. For small samples where expected counts are low, the exact binomial is preferable to the chi-square approximation for two-category data. With large samples, the two approaches converge.

What If Expected Counts Are Small?

Small expected frequencies reduce the accuracy of the chi-square approximation. Several responses are available, and the right choice depends on the specifics of the problem.

  • Combine categories: If two or more categories are substantively similar and combining them makes theoretical sense, doing so raises the combined expected count. The critical constraint is that category merging must be justified by the research question, not by the desire to clear an assumption threshold. Collapsing categories changes what is being tested.
  • Exact multinomial test: For small samples with multiple categories, exact multinomial procedures compute the probability of the observed table (and more extreme tables) under the null distribution without relying on the chi-square approximation.
  • Monte Carlo simulation: Randomly generates a large number of tables under H₀ and uses the proportion of simulated tables exceeding the observed χ² as the p-value. Many statistical packages offer this as an option.
  • Collect more data: Sometimes the most direct remedy is a larger sample, which raises expected counts without altering the research question.
🚫
Do Not Delete Categories Just Because They Fit Poorly

Observed discrepancies are exactly the evidence being examined. Removing a category because it has a large residual, or combining categories after seeing the results to reach significance, introduces data-dependent analysis choices that invalidate the test's stated error rates.

How to Report the Test

Thorough reporting gives readers everything they need to evaluate and reproduce the result. A complete report includes the observed counts, expected counts or proportions, the chi-square statistic, degrees of freedom, p-value, sample size, and an effect size where appropriate.

A general academic reporting structure:

📝 Reporting Template
  • A chi-square goodness-of-fit test indicated that the observed distribution [did / did not] differ significantly from the expected distribution, χ²(df, N = n) = value, p = value.
  • When reporting effect size: ... Cohen's w = value, indicating a [small / medium / large] effect.
  • Example (Significant): "A chi-square goodness-of-fit test indicated that the observed sales distribution differed significantly from the historical proportions, χ²(3, N = 200) = 8.06, p = .045, w = 0.20."
  • Example (Non-Significant): "A chi-square goodness-of-fit test indicated that the observed preference distribution did not differ significantly from equal proportions, χ²(3, N = 100) = 5.04, p = .169, w = 0.22."

Reporting only the p-value is insufficient. The pattern of observed versus expected counts, the size of the chi-square statistic, and an effect measure all contribute to a complete picture of the findings. For p-value interpretation guidance more broadly, the p-values guide covers the key concepts. For the mechanics of statistical decisions, see the decision rule guide.

Common Mistakes

Mistake 1 — Using percentages as observed counts
❌ Entering 25%, 35%, 40% as the observed values when n is unknown
✓ Convert percentages to actual counts using the known sample size: O = percentage × n
Mistake 2 — Deriving expected proportions from the observed data
❌ Setting expected proportions to match the sample proportions exactly, then running the test
✓ Expected proportions must come from an independently justified null model, not the same data being tested
Mistake 3 — Using the wrong degrees of freedom
❌ Using (r − 1)(c − 1), the contingency-table df, for a one-variable goodness-of-fit test
✓ Use df = k − 1 for a standard GOF test with k categories and fully specified null probabilities
Mistake 4 — Accepting the null hypothesis
❌ "p = 0.169, so the distribution fits perfectly" or "We proved the null hypothesis is true"
✓ "The data do not provide sufficient evidence to reject the null distribution at α = 0.05"
Mistake 5 — Treating the omnibus test as proof every category differs
❌ "χ²(3) = 8.06, p = .045, so each of the four categories is significantly different from expected"
✓ A significant omnibus result means the overall distribution departs from H₀. Individual category conclusions require residual analysis with appropriate multiplicity adjustment.
Mistake 6 — Checking expected-count conditions on observed counts
❌ Looking at observed frequencies to decide whether the chi-square approximation is reliable
✓ Expected-count conditions apply to Eᵢ = n × pᵢ₀, the counts predicted by the null model, not the observed counts
Mistake 7 — Confusing p-value interpretation
❌ "p = 0.045 means there is a 4.5% chance that H₀ is true" or "only a 4.5% chance the result happened by chance"
✓ The p-value is the probability, under H₀ and model assumptions, of obtaining a χ² statistic at least as large as observed. It is not the probability that H₀ is true.

Verify Your Calculation

Once you have computed the chi-square statistic and degrees of freedom by hand using the contribution table above, you can verify the result and obtain the exact p-value using the Chi-Square Calculator at Statistics Fundamentals. For finding critical values at a specific significance level, the critical value calculator supports the chi-square distribution with any degrees of freedom.

🔗
Quick Tools

Chi-square table for manual lookup: chi-square table. Not sure how to use a chi-square table? See the guide on how to read a chi-square table. For degrees of freedom background, the degrees of freedom guide covers the concept thoroughly.

Frequently Asked Questions

What is a chi-square goodness-of-fit test?
A chi-square goodness-of-fit test is a hypothesis test used to determine whether the observed frequencies across the categories of one categorical variable are consistent with a specified set of expected proportions. It uses the test statistic χ² = Σ[(O − E)² / E] and compares the result against a chi-square reference distribution to determine whether the observed discrepancies exceed what random sampling variation would typically produce.
When do you use a chi-square goodness-of-fit test?
Use the goodness-of-fit test when you have one categorical variable, observed frequency counts across its categories, and a specific null distribution specifying the expected probability for each category. Common applications include testing whether survey responses are equally distributed, whether a die is fair, whether genetic ratios match Mendelian predictions, or whether sales data matches a historical benchmark distribution.
What is the chi-square goodness-of-fit formula?
The formula is χ² = Σ[(Oᵢ − Eᵢ)² / Eᵢ], where Oᵢ is the observed frequency for category i, Eᵢ is the expected frequency, and the sum runs across all k categories. Each term measures how much that category's observed count departs from its expected count, scaled relative to the expected count. Summing these gives the total test statistic.
How do you calculate expected frequencies?
Expected frequency for category i is Eᵢ = n × pᵢ₀, where n is the total sample size and pᵢ₀ is the hypothesized probability for that category under the null hypothesis. The null probabilities must sum to 1, and consequently the expected counts will sum to n. Expected probabilities need not be equal — any justified null distribution works.
What are the degrees of freedom?
For a standard goodness-of-fit test with k categories and all null probabilities specified in advance from an external source, df = k − 1. The intuition is that once the total n is fixed, knowing k − 1 category counts determines the last one. If parameters of the null distribution were estimated from the same data, degrees of freedom are reduced accordingly: df = k − 1 − m, where m is the number of independently estimated parameters.
Do expected frequencies all have to be greater than 5?
The "expected count ≥ 5" guideline is a widely cited rule of thumb, not a mathematical absolute. It approximates the conditions under which the chi-square distribution provides an accurate reference for the test statistic. Some authorities allow up to 20% of expected counts to fall below 5, provided none falls below 1. The exact threshold depends on the number of categories, the sample size, and the distribution structure. When expected counts are too small, alternatives include combining categories, exact multinomial tests, or Monte Carlo methods.
Is the goodness-of-fit test one-tailed or two-tailed?
The test is effectively upper-tailed. Because χ² is built from squared differences, it is always non-negative. Larger values indicate poorer fit with the null distribution. The p-value is the probability of obtaining a chi-square statistic at least as large as observed, which is computed from the upper tail of the chi-square distribution. The term "two-tailed" does not apply in the conventional sense here.
What is the difference between goodness of fit and independence?
Goodness of fit tests one categorical variable against a pre-specified distribution. Independence tests whether two categorical variables are related, using a contingency table with r rows and c columns. Expected counts are computed differently: in GOF, Eᵢ = n × pᵢ₀ from the null model; in independence, expected cell counts are derived from marginal totals. Degrees of freedom also differ: k − 1 for GOF vs (r − 1)(c − 1) for independence. Using the wrong formula for either test produces incorrect results.
Can chi-square goodness of fit use unequal expected proportions?
Yes. The null distribution can assign any combination of probabilities to the categories as long as they sum to 1 and are specified from a justified external source. Unequal null proportions are common in practice: genetic ratios, historical benchmarks, regulatory standards, and theoretical model predictions all produce unequal expected proportions. The equal-proportions case (uniform distribution) is simply one special case of this general framework.
How do you report a chi-square goodness-of-fit test?
A complete report includes: observed counts per category, expected counts or null proportions, the chi-square statistic, degrees of freedom, p-value, total sample size, and an effect size (Cohen's w) where appropriate. Example: "A chi-square goodness-of-fit test indicated that the observed distribution differed significantly from the expected distribution, χ²(3, N = 200) = 8.06, p = .045, w = 0.20." Reporting only the p-value omits information needed to evaluate the practical significance of the result.
Pearson, K. (1900). "On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling." Philosophical Magazine, 50(302), 157–175. | Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Hillsdale, NJ: Erlbaum. | Agresti, A. (2013). Categorical Data Analysis (3rd ed.). Wiley.