What Is a Z-Test?
The shape of that general formula stays the same across all Z-tests. What changes is the estimate being tested (a sample mean, a sample proportion, a difference between two proportions) and the corresponding standard error formula. Understanding those differences is the heart of applying Z-tests correctly.
For a deeper look at the theory behind the one-sample mean case — assumptions, sampling distributions, and formal derivations — see the one-sample Z-test guide. This page concentrates on worked examples and calculation.
When Should You Use a Z-Test?
The right test depends on the parameter, what quantities are known, and the study design — not on a sample-size cutoff alone.
This is an oversimplification. For a one-sample mean test, the classical Z-test requires the population standard deviation σ to be known. If σ is estimated from the sample, a t-based method is generally appropriate for means — regardless of how large the sample is. The t-distribution approaches the normal as degrees of freedom increase, so the numerical difference shrinks with large samples, but the reasoning matters for understanding which framework you're in.
| Test | Use when… | Key condition |
|---|---|---|
| One-sample mean Z-test | Testing a population mean μ | Population SD σ is known |
| One-proportion Z-test | Testing a population proportion p | Normal approximation to binomial adequate |
| Two-proportion Z-test | Comparing two independent proportions | Independent groups; adequate expected counts |
| One-sample t-test (not Z) | Testing a mean when σ is unknown | Estimate σ from sample; use t-distribution |
When you're unsure which test fits your data, the statistical test selector walks through the decision systematically.
How to Solve a Z-Test Step by Step
Z-Test Workflow — 10 Steps
Z-Test Examples at a Glance
| Example | Test type | Alternative | Key formula |
|---|---|---|---|
| 1 — Product fill weight | One-sample mean Z | Two-sided (μ ≠ 500) | (x̄ − μ₀) / (σ/√n) |
| 2 — Processing time | One-sample mean Z | Upper-tailed (μ > 30) | (x̄ − μ₀) / (σ/√n) |
| 3 — Process improvement | One-sample mean Z | Lower-tailed (μ < 100) | (x̄ − μ₀) / (σ/√n) |
| 4 — Customer preference | One-proportion Z | Two-sided (p ≠ 0.60) | (p̂ − p₀) / √[p₀(1−p₀)/n] |
| 5 — Website conversion | Two-proportion Z | Two-sided (pA ≠ pB) | (p̂A − p̂B) / √[p̂pool(1−p̂pool)(1/n₁+1/n₂)] |
One-Sample Mean Z-Test Formula
The classical one-sample mean Z-test applies when testing H₀: μ = μ₀ and the population standard deviation σ is known. The test statistic is:
x̄ = sample mean
μ₀ = hypothesized population mean
σ = known population standard deviation
n = sample size
σ/√n = standard error
The denominator σ/√n — the standard error — scales the observed difference relative to how much variability you expect in sample means under repeated sampling. A positive z means the sample mean is above the null value; a negative z means it is below. The magnitude |z| tells you how many null-model standard errors separate the observed result from μ₀.
Example 1 — Two-Tailed One-Sample Mean Z-Test
Problem: A manufacturer states that the mean fill weight of a bottled product is 500 g. The known population standard deviation is σ = 12 g. A quality team randomly samples n = 64 bottles and records a mean fill weight of x̄ = 496.5 g. At α = 0.05, does the sample provide evidence that the mean fill weight differs from 500 g?
Parameter: The population mean fill weight μ (in grams).
Hypotheses:
H₀: μ = 500
H₁: μ ≠ 500 (two-tailed — testing for a difference in either direction)
Assumptions check: The sample is drawn randomly; observations are independent; σ = 12 is treated as known. The sampling distribution of x̄ is approximately normal (the population is continuous and n = 64 is reasonably large). A Z-test is appropriate here.
Significance level: α = 0.05. For a two-tailed test, critical values are approximately ±1.96.
Standard error:
SE = σ / √n = 12 / √64 = 12 / 8 = 1.5
Z statistic:
z = (x̄ − μ₀) / SE = (496.5 − 500) / 1.5 = −3.5 / 1.5 = −2.3333
P-value (two-tailed):
P(Z < −2.3333) ≈ 0.0098 (left tail).
Two-tailed p-value = 2 × 0.0098 = p ≈ 0.0196
Critical value check: |z| = 2.333 > 1.96. The statistic falls in the rejection region, consistent with the p-value decision.
Decision: p = 0.020 < α = 0.05 → Reject H₀.
Conclusion: At the 5% significance level, the sample provides evidence that the population mean fill weight differs from 500 g. The observed sample mean (496.5 g) is 3.5 g below the target.
✅ Reject H₀. p ≈ 0.020 < 0.05. Evidence that μ ≠ 500 g.
⚠️ What this does not prove: Statistical significance does not tell us whether a 3.5 g shortfall is practically important for consumers or the business. Effect size and context determine practical significance, which is a separate question from the p-value.
Example 2 — Upper-Tailed (Right-Tailed) Mean Z-Test
Problem: A delivery company's historical records give an average processing time of 30 minutes with a known process standard deviation of σ = 6 minutes. After a workflow change, a random sample of n = 100 orders yields a mean processing time of x̄ = 31.2 minutes. Has mean processing time increased? Test at α = 0.05.
Parameter: The population mean processing time μ (in minutes).
Hypotheses:
H₀: μ = 30
H₁: μ > 30 (upper-tailed — the question asks specifically whether time increased)
Assumptions check: Random sample; independent orders; σ = 6 is treated as known. With n = 100, the sampling distribution of x̄ is well-approximated by a normal. Z-test is appropriate.
Significance level: α = 0.05. For an upper one-tailed test, the critical value is z* = 1.645.
Standard error:
SE = 6 / √100 = 6 / 10 = 0.6
Z statistic:
z = (31.2 − 30) / 0.6 = 1.2 / 0.6 = 2.00
P-value (upper-tailed):
P(Z > 2.00) = 1 − Φ(2.00) ≈ 1 − 0.9772 = p ≈ 0.0228
Critical value check: z = 2.00 > 1.645 (upper critical value). Reject H₀.
Decision: p = 0.023 < α = 0.05 → Reject H₀.
Conclusion: At the 5% significance level, the data provide evidence that mean processing time has increased beyond 30 minutes. The estimated increase is 1.2 minutes.
✅ Reject H₀. p ≈ 0.023 < 0.05. Evidence that μ > 30 minutes.
⚠️ Note: The decision to test upper-tailed was made because the research question specifically asked whether time increased. Choosing one-tailed vs two-tailed after observing the data inflates the Type I error rate and is not valid practice. Learn more about one-tailed vs two-tailed tests.
Example 3 — Lower-Tailed (Left-Tailed) Mean Z-Test
Problem: An industrial process has a historical benchmark of μ₀ = 100 units per hour with a known standard deviation of σ = 15 units/hour. Engineers implement a new procedure and want to know whether mean output has dropped below the benchmark. A random sample of n = 36 shifts gives x̄ = 95 units/hour. Test at α = 0.05.
Parameter: The population mean hourly output μ.
Hypotheses:
H₀: μ = 100
H₁: μ < 100 (lower-tailed — question asks whether output dropped)
Assumptions check: Shifts sampled randomly; outputs approximately independent; σ = 15 is treated as known. Z-test is appropriate.
Significance level: α = 0.05. For a lower one-tailed test, the critical value is z* = −1.645. Reject H₀ if z < −1.645.
Standard error:
SE = 15 / √36 = 15 / 6 = 2.5
Z statistic:
z = (95 − 100) / 2.5 = −5 / 2.5 = −2.00
P-value (lower-tailed):
P(Z < −2.00) = Φ(−2.00) ≈ p ≈ 0.0228
Critical value check: z = −2.00 < −1.645. Falls in the rejection region.
Decision: p = 0.023 < α = 0.05 → Reject H₀.
Conclusion: At the 5% significance level, there is evidence that the new procedure is associated with a decrease in mean hourly output below 100 units. The observed mean is 5 units/hour lower than the benchmark.
✅ Reject H₀. p ≈ 0.023 < 0.05. Evidence that μ < 100 units/hour.
⚠️ What this does not prove: The test tells us the output decline is statistically detectable at α = 0.05. Whether a drop of 5 units/hour is operationally significant — and whether the new procedure caused it — are separate questions requiring effect-size analysis and knowledge of the study design.
Two-tailed α = 0.05: reject if |z| > 1.960 | Upper-tailed α = 0.05: reject if z > 1.645 | Lower-tailed α = 0.05: reject if z < −1.645. These values come from the standard normal (Z) table.
One-Proportion Z-Test Formula
When the parameter of interest is a population proportion p, and the hypothesis is H₀: p = p₀, the Z statistic uses the null proportion p₀ in the standard error — not the sample proportion p̂. This is a frequent mistake worth flagging explicitly.
p̂ = observed sample proportion
p₀ = hypothesized population proportion
n = sample size
The normal approximation to the binomial requires that np₀ and n(1 − p₀) are both sufficiently large. Many textbooks use thresholds of 5 or 10; the key point is to verify this condition before proceeding, not skip it.
Example 4 — One-Proportion Z-Test
Problem: A company claims that 60% of its customers prefer Product A over Product B. A market research team surveys n = 200 randomly selected customers; 134 say they prefer Product A. Does the data contradict the company's claim? Test at α = 0.05.
Parameter: The population proportion p of customers who prefer Product A.
Hypotheses:
H₀: p = 0.60
H₁: p ≠ 0.60 (two-tailed — testing whether the true proportion differs in either direction from 0.60)
Assumptions check:
– Random sample; observations are independent.
– Normal approximation: np₀ = 200 × 0.60 = 120 ≥ 10 ✓ and n(1 − p₀) = 200 × 0.40 = 80 ≥ 10 ✓.
Conditions are well satisfied.
Significance level: α = 0.05. Critical values: ±1.96.
Sample proportion:
p̂ = 134 / 200 = 0.67
Null standard error:
SE₀ = √[0.60 × 0.40 / 200] = √(0.24/200) = √0.0012 ≈ 0.03464
Z statistic:
z = (0.67 − 0.60) / 0.03464 = 0.07 / 0.03464 ≈ 2.0207
P-value (two-tailed):
P(Z > 2.0207) ≈ 0.0217 (upper tail).
Two-tailed p = 2 × 0.0217 ≈ 0.0433
Decision: p = 0.043 < α = 0.05 → Reject H₀.
Conclusion: At the 5% significance level, the data provide evidence that the population proportion who prefer Product A differs from 0.60. The sample estimate (p̂ = 0.67) is 7 percentage points above the claimed value.
✅ Reject H₀. p ≈ 0.043 < 0.05. Evidence that p ≠ 0.60.
⚠️ Critical note on the SE formula: The null standard error uses p₀ = 0.60, not the sample proportion p̂ = 0.67. Using p̂ in the denominator here would be incorrect for this hypothesis test. That substitution is appropriate in some confidence interval formulas, but not the hypothesis test. See the proportion hypothesis testing guide for details.
Two-Proportion Z-Test Formula
When comparing two independent proportions — H₀: p₁ = p₂ — the null hypothesis assumes a single common proportion. The test estimates that common value with a pooled proportion combining both samples:
p̂₁, p̂₂ = observed sample proportions
p̂pool = (x₁ + x₂)/(n₁ + n₂)
n₁, n₂ = sample sizes
Using pooled p̂ here is correct precisely because H₀ assumes both groups share the same true proportion. This pooled SE is specific to equality testing. If you are building a confidence interval for p₁ − p₂ instead, the appropriate SE formula uses p̂₁ and p̂₂ separately.
Example 5 — Two-Proportion Z-Test
Problem: A company runs an A/B test on two landing page versions. Version A receives 200 visitors and 80 convert. Version B receives 200 visitors and 60 convert. Does the conversion rate differ between the two versions? Test at α = 0.05.
Parameter: The population conversion proportions pA and pB.
Hypotheses:
H₀: pA = pB
H₁: pA ≠ pB (two-tailed)
Assumptions check:
– Visitors are independently assigned to each version (proper A/B testing protocol).
– Pooled p̂ = (80 + 60)/(200 + 200) = 140/400 = 0.35.
– Expected counts under H₀: n×p̂pool = 200×0.35 = 70 (successes) and 200×0.65 = 130 (failures) for each group — both well above 10. ✓
Significance level: α = 0.05. Critical values: ±1.96.
Sample proportions:
p̂A = 80/200 = 0.40
p̂B = 60/200 = 0.30
Pooled proportion and standard error:
p̂pool = (80 + 60) / (200 + 200) = 140 / 400 = 0.35
SE = √[0.35 × 0.65 × (1/200 + 1/200)]
= √[0.2275 × 0.01]
= √0.002275
≈ 0.047697
Z statistic:
z = (0.40 − 0.30) / 0.047697 = 0.10 / 0.047697 ≈ 2.097
P-value (two-tailed):
P(Z > 2.097) ≈ 0.0180 (upper tail).
Two-tailed p = 2 × 0.0180 ≈ 0.0360
Decision: p = 0.036 < α = 0.05 → Reject H₀.
Conclusion: At the 5% significance level, the data provide evidence of a difference in conversion proportions between Version A and Version B. The estimated difference is 0.40 − 0.30 = 0.10 (10 percentage points), with Version A converting at a higher rate.
✅ Reject H₀. p ≈ 0.036 < 0.05. Evidence that pA ≠ pB.
⚠️ Causation requires design, not just significance: A statistically significant difference does not by itself mean Version A caused higher conversions. Causal conclusions require proper randomization and experimental implementation. This applies to A/B testing generally.
Z-Test vs T-Test: Which One Applies?
Students often see a rule like "use Z when n ≥ 30." That shortcut gets the logic backwards. The right question is not "how large is my sample?" but "do I know σ?"
| Feature | Z-test for a mean | T-test for a mean |
|---|---|---|
| Population SD | Known (σ given) | Unknown (estimated as s from the sample) |
| Reference distribution | Standard normal | t-distribution with n−1 degrees of freedom |
| Critical values | Fixed at ±1.96 (two-tailed, α=.05) | Depend on degrees of freedom |
| n = 30 rule? | Insufficient basis on its own | Use t whenever σ is unknown, regardless of n |
| As n grows large | — | t-distribution approaches the normal; results converge numerically |
In most real-world applications, σ is unknown and estimated from the data. That makes a t-test the standard approach for means. The Z-test for means appears most often in educational examples — where σ is given as a known value — and in some process-control contexts where long-run σ has been established.
For a systematic approach to selecting between these and other tests, use the statistical test selector.
P-Value Method vs Critical-Value Method
Both methods draw on the same Z statistic and reach the same decision when applied consistently. Example 1 above illustrates both:
| Method | What you compare | Example 1 result | Decision |
|---|---|---|---|
| P-value method | p vs α | p = 0.020 vs α = 0.05 | 0.020 < 0.05 → Reject H₀ |
| Critical-value method | |z| vs z* | |−2.333| = 2.333 vs z* = 1.96 | 2.333 > 1.96 → Reject H₀ |
Both methods tell you to reject H₀ — they must, because they are mathematically equivalent for any standard normal test. Choose whichever your course or software reports. When using the Z-test calculator, both the p-value and the test statistic are reported so you can verify against either method.
How to Write a Z-Test Conclusion
- When you reject H₀: "At the α = [level] significance level, the data provide sufficient evidence to conclude that [population parameter statement in plain language]."
- When you fail to reject H₀: "At the α = [level] significance level, the data do not provide sufficient evidence to conclude that [H₁ in plain language]. This does not prove H₀ is true."
- Add practical context: "The estimated [mean/proportion/difference] is [value], which should also be evaluated for practical importance in context."
Never write "accept H₀" — a non-significant result means lack of sufficient evidence, not proof of the null. Never write "p = 0.02 means there is a 2% probability H₀ is true" — a p-value is not the probability the null hypothesis is correct. It is the probability of observing a result at least as extreme as yours under the null model's assumptions.
Common Z-Test Mistakes
| Mistake | Wrong approach | Correct approach |
|---|---|---|
| Wrong SE for means | Using σ alone in the denominator | Use SE = σ/√n (standard error, not SD) |
| Using s as if it were σ | Plugging sample SD s into a "Z-test" without noting the difference | σ unknown → use a t-test framework for the mean |
| n ≥ 30 rule | "n = 50 so I'll use Z" even when σ is unknown | Test choice depends on what is known, not just sample size |
| Wrong tail after seeing data | Switching to one-tailed because the observed direction looks significant | Set H₁ direction based on the research question before examining results |
| Wrong proportion SE | Using p̂ in the one-proportion null SE | Use p₀ under H₀: SE₀ = √[p₀(1−p₀)/n] |
| Forgetting pooled p̂ | Using separate p̂₁ and p̂₂ in the two-proportion equality test SE | Use pooled p̂ = (x₁+x₂)/(n₁+n₂) under H₀: p₁=p₂ |
| "Accept H₀" | p > α, so "H₀ is true" | "Fail to reject H₀" — the test is inconclusive, not confirmatory |
| p-value misinterpretation | "p = 0.04 means 4% chance H₀ is true" | p-value = probability of result this extreme if H₀ were true |
Practice Problems
Work through each problem fully before revealing the answer. Try to write your own hypotheses, calculate the SE, find z, obtain the p-value, and state a contextual conclusion.
A machine fills bags of flour with a mean of 1000 g and a known population SD of σ = 20 g. A sample of n = 100 bags yields x̄ = 997.5 g. At α = 0.05, does the mean fill weight differ from 1000 g?
Show Solution
H₀: μ = 1000 | H₁: μ ≠ 1000 (two-tailed)
α = 0.05. Critical values: ±1.96.
SE = 20/√100 = 20/10 = 2.0
z = (997.5 − 1000)/2.0 = −2.5/2.0 = −1.25
Two-tailed p-value: 2 × P(Z < −1.25) ≈ 2 × 0.1056 ≈ 0.2113
Decision: Fail to reject H₀. p = 0.211 > 0.05. |z| = 1.25 < 1.96.
Conclusion: At α = 0.05, the data do not provide sufficient evidence that the mean fill weight differs from 1000 g. This does not prove μ = 1000 — it means the 2.5 g shortfall is not statistically detectable at this significance level and sample size.
A school's standardized test historically averages 75 points (σ = 10, known). After a tutoring program, a sample of n = 49 students yields x̄ = 77.5 points. Did the program increase the mean score? Test at α = 0.01.
Show Solution
H₀: μ = 75 | H₁: μ > 75 (upper-tailed)
α = 0.01. Critical value: z* = 2.326.
SE = 10/√49 = 10/7 ≈ 1.4286
z = (77.5 − 75)/1.4286 = 2.5/1.4286 ≈ 1.750
Upper-tail p-value: P(Z > 1.750) ≈ 0.0401
Decision: Fail to reject H₀. p = 0.040 > α = 0.01. Also: z = 1.75 < 2.326.
Conclusion: At the 1% significance level, the data do not provide sufficient evidence that the tutoring program increased mean scores. Note: at α = 0.05 the result would be significant (p < 0.05). The choice of significance level matters and should be set before the analysis.
A website historically converts 25% of visitors. After a homepage redesign, a random sample of n = 400 visitors produces 112 conversions. Has the conversion rate changed? Test at α = 0.05.
Show Solution
H₀: p = 0.25 | H₁: p ≠ 0.25 (two-tailed)
Conditions: np₀ = 400 × 0.25 = 100 ≥ 10 ✓ n(1−p₀) = 300 ≥ 10 ✓
p̂ = 112/400 = 0.28
SE₀ = √[0.25 × 0.75 / 400] = √(0.1875/400) = √0.00046875 ≈ 0.02165
z = (0.28 − 0.25)/0.02165 = 0.03/0.02165 ≈ 1.386
Two-tailed p-value: 2 × P(Z > 1.386) ≈ 2 × 0.0828 ≈ 0.1655
Decision: Fail to reject H₀. p = 0.166 > 0.05.
Conclusion: At α = 0.05, the sample does not provide sufficient evidence that the conversion rate has changed from 25%. The observed rate (28%) could reasonably occur by chance when p = 0.25.
Email campaign A is sent to 300 recipients; 90 open it. Campaign B goes to 300 recipients; 72 open it. Is there a difference in open rates? Test at α = 0.05.
Show Solution
H₀: pA = pB | H₁: pA ≠ pB (two-tailed)
p̂A = 90/300 = 0.30 | p̂B = 72/300 = 0.24
p̂pool = (90 + 72)/(300 + 300) = 162/600 = 0.27
Conditions: 300 × 0.27 = 81 ≥ 10 ✓ 300 × 0.73 = 219 ≥ 10 ✓
SE = √[0.27 × 0.73 × (1/300 + 1/300)] = √[0.1971 × 0.006667] = √0.001314 ≈ 0.03625
z = (0.30 − 0.24)/0.03625 = 0.06/0.03625 ≈ 1.655
Two-tailed p-value: 2 × P(Z > 1.655) ≈ 2 × 0.0490 ≈ 0.0980
Decision: Fail to reject H₀. p = 0.098 > 0.05.
Conclusion: At α = 0.05, there is not sufficient evidence of a difference in open rates between the two campaigns. The observed 6 percentage point difference could plausibly arise by chance.
A researcher has n = 80 observations, a sample mean x̄, and a sample SD s. She does not know σ. Should she automatically use a one-sample Z-test because n > 30?
Show Solution
No. The sample size alone does not determine the correct test for a mean. The classical one-sample Z-test for a mean requires that the population SD σ is known. Here, only the sample SD s is available, so σ is unknown.
With σ unknown, a one-sample t-test is generally the appropriate method, using s in place of σ and the t-distribution with df = n − 1 = 79.
As the degrees of freedom increase, the t-distribution approaches the standard normal, so the numerical difference between Z and t becomes small with large n. But the conceptual framework — and the assumption set — differs. Using s in a Z formula without acknowledging this is imprecise at best and misleading at worst.
A test yields z = 1.45, p = 0.147, α = 0.05 (two-tailed). What is the correct decision and what does it mean?
Show Solution
Decision: Fail to reject H₀. p = 0.147 > α = 0.05.
This means: at the 5% significance level, the data do not provide sufficient evidence against H₀.
What it does not mean: It does not mean H₀ is true. A non-significant result may reflect a real but small effect that the study lacks the power to detect, high variability in the data, a small sample, or a genuinely absent effect. The result is inconclusive about the truth of H₀. For more on this, see the discussion of p-values and power.
Z-Test Calculator
Enter your values below to calculate the Z statistic and p-value for a one-sample mean Z-test. For proportion tests, use the dedicated Z-test calculator.
One-Sample Mean Z-Test Calculator
Frequently Asked Questions
What is the formula for a one-sample Z-test?
For testing H₀: μ = μ₀ with known population SD σ: z = (x̄ − μ₀) / (σ / √n). The denominator σ/√n is the standard error — not σ alone.
Can you use a Z-test when σ is unknown?
For a mean, no — not in the classical sense. If σ is unknown and estimated from the sample as s, a t-test is generally the appropriate method. The t-distribution accounts for the additional uncertainty of estimating σ. With very large samples the t and Z distributions are numerically close, but the conceptual difference remains.
What is the critical Z value at α = 0.05?
The answer depends on the direction of the test. For a two-tailed test at α = 0.05 the critical values are approximately ±1.96. For an upper-tailed test: +1.645. For a lower-tailed test: −1.645. The value 1.96 is not "the" Z critical value — it applies specifically to the two-sided case at α = 0.05. Use the critical value calculator for other α levels.
What does a negative Z statistic mean?
A negative z means the observed estimate (sample mean or sample proportion) is below the null value. It does not by itself indicate the result is significant. Whether to reject H₀ depends on the test direction and the p-value relative to α.
Is Z-test the same as Z-score?
They are related but distinct. A Z-score standardizes an individual data value relative to a distribution: (value − mean) / SD. A Z-test uses a Z statistic within a hypothesis-testing framework: (estimate − null value) / standard error. The denominators differ — SD vs SE — and the purpose is different. The Z-score examples page covers individual standardization specifically.
What is a one-proportion Z-test?
A one-proportion Z-test evaluates whether a population proportion p equals a specified null value p₀. The formula is z = (p̂ − p₀) / √[p₀(1−p₀)/n]. The key detail: the null proportion p₀ goes in the denominator, not the sample proportion p̂. See Example 4 above for a fully worked problem.
Why is the pooled proportion used in the two-proportion Z-test?
Under H₀: p₁ = p₂, both groups share the same true proportion. The pooled estimate p̂pool = (x₁ + x₂)/(n₁ + n₂) combines both samples to estimate that common value, which is then used in the standard error. If instead you were constructing a confidence interval for p₁ − p₂ (not testing equality), you would use separate p̂₁ and p̂₂ in the SE formula.
How do you interpret p = 0.03 in a Z-test?
A p-value of 0.03 means: if H₀ were true, the probability of observing a test statistic as extreme as yours (or more extreme) is 3%. At α = 0.05, this is below the threshold — reject H₀. It does not mean "there is a 3% chance H₀ is true." The p-value describes the data under the null model, not the probability of the null model itself. For more on this, see the p-values guide and p-value examples.
Key Takeaways
- The general Z-test structure is always: Z = (estimate − null value) / null-model standard error. What changes is the estimate and the SE formula.
- One-sample mean Z-test requires known σ. Use SE = σ/√n. Test choice depends on the parameter and what is known, not on sample size alone.
- Set the test direction (two-tailed, upper, lower) based on the research question — before examining the data.
- One-proportion Z-test uses p₀ in the null SE: √[p₀(1−p₀)/n]. Check the normal approximation conditions first.
- Two-proportion Z-test uses pooled p̂ in the SE under H₀. Pooling is correct for equality testing; separate proportions go in CI formulas.
- Fail to reject H₀ ≠ accept H₀. Statistical non-significance does not prove the null.
- Statistical significance ≠ practical importance. Always report the estimated effect alongside the p-value.
- A Z-test cannot establish causation. That depends on the study design and random assignment.
Sources and Further Reading
- NIST/SEMATECH e-Handbook of Statistical Methods — Section 7.3.2: One Sample Z-Test. itl.nist.gov
- OpenStax Statistics — Introductory Statistics (2013). Chapters 9–10 cover hypothesis testing for means and proportions with full examples.
- Moore, D.S., McCabe, G.P., & Craig, B.A. — Introduction to the Practice of Statistics, 9th ed. W. H. Freeman. Standard introductory reference for Z-test and t-test distinctions.
- Fisher, R.A. (1925) — Statistical Methods for Research Workers. Edinburgh: Oliver and Boyd. Foundational text for significance testing.