Interactive Sampling Distribution Simulator
Each repeated sample contributes exactly one sample mean to the histogram.
Advanced simulation settings
1. Population Distribution
The stable distribution from which every observation is drawn.
2. Current Random Sample
One random sample. Its observations vary from draw to draw.
3. Sampling Distribution of the Mean
Each bar summarizes sample means from repeated samples of the same size.
Theory
Simulation
What Are You Seeing?
In this simulator, every sample produces one sample mean. Repeating the process creates a distribution of those means. The center should settle near the population mean, while the spread is described by the standard error.
The three charts are intentionally connected. The first chart is the population. The second chart shows one actual random sample from that population. The third chart does not contain the original observations; it contains the means computed from many samples. This is the distinction that is easy to miss when a sampling distribution is shown as a finished bell curve without showing how it was produced.
Examples to Try
Use these presets as short experiments. Each one resets the simulation so the results are not mixed across different sampling conditions.
1. Normal population: compare spread
Start with a small sample, then increase n. The sampling distribution is normal in the independent normal model at both sizes, but the larger-sample distribution is narrower.
2. Right-skewed population: watch the shape change
With n = 2, the sample-mean distribution can still be visibly skewed. Increase n and compare it with the normal approximation.
3. Uniform population: start at n = 1
When n = 1, each sample mean is simply the one observation that was drawn. The sampling distribution therefore mirrors the population in theory.
4. Same population, four times the sample size
Compare n = 10 with n = 40. Because SE is proportional to 1/√n, multiplying n by four cuts the theoretical standard error in half.
What Is a Sampling Distribution?
A sampling distribution is the probability distribution of a statistic over repeated samples drawn under the same conditions. Here, the statistic is the sample mean, x̄. Imagine drawing a sample of size 10, calculating its mean, returning to the population, and repeating the same process thousands of times. The collection of those thousands of means forms the sampling distribution of x̄.
Sampling distributions matter because a statistic calculated from one sample is not fixed across all possible samples. Another random sample usually gives a different mean. The sampling distribution describes that sample-to-sample variation and provides the basis for standard errors, confidence intervals, and many hypothesis tests.
Population vs Sample vs Sampling Distribution
The distribution that generates individual observations. Its parameters include the population mean μ and population standard deviation σ.
The observations inside one random sample. A different sample usually has a different pattern and a different sample mean.
The distribution of a statistic across repeated samples. In this tool, every point represented in the histogram is one sample mean.
| Question | Population | One sample | Sampling distribution |
|---|---|---|---|
| What are the values? | Individual observations | Individual observations | Sample means |
| How many observations create one mean? | Not applicable | n observations | Each mean comes from n observations |
| What changes when you repeat sampling? | The theoretical population stays fixed | The observations and x̄ change | The histogram accumulates more sample means |
How the Simulator Works
- Choose a population. The tool uses a mathematically defined normal, exponential, uniform, or two-component normal mixture distribution.
- Choose the sample size, n. This determines how many independent observations are drawn for each sample.
- Draw a sample. The observations appear in the middle chart and their arithmetic mean is calculated.
- Add that mean to the histogram. One sample contributes one x̄, regardless of whether n is 2 or 100.
- Repeat the process. As the number of repeated samples grows, the simulated center and spread become easier to compare with their theoretical values.
Changing the population, any population parameter, or n resets the accumulated sample means. Mixing means generated under different sampling conditions would create a histogram that no longer represents one sampling distribution.
Central Limit Theorem
The Central Limit Theorem describes the sampling distribution of a properly standardized sum or mean, not the shape of the original data. In the classical independent, identically distributed setting with finite variance, the standardized sample mean approaches a normal distribution as n grows.
That does not mean a large sample makes the population normal, and it does not mean the observations within one sample become normal. A right-skewed population remains right-skewed. What becomes increasingly well approximated by a normal distribution, under the relevant conditions, is the distribution of sample means across repeated samples.
There is also no universal “n = 30” cutoff. A nearly symmetric population may need only a modest n for a useful normal approximation, while a strongly skewed or heavy-tailed population can require more. Use the right-skewed preset to see why a fixed threshold is a poor teaching rule.
Normal populations are a special case
If the population itself is normal and observations are independently sampled, the sample mean is normally distributed for every positive sample size. For non-normal populations with finite variance, the blue curve in the simulator is labeled as a CLT normal approximation, not as the exact sampling distribution.
How Sample Size Changes the Sampling Distribution
For independent observations from a population with standard deviation σ, the standard error of the sample mean is σ/√n. Increasing n therefore makes sample means less variable. The relationship follows a square root: doubling n does not halve the standard error.
| Change in sample size | SE multiplier | Meaning |
|---|---|---|
| n → 2n | 1/√2 ≈ 0.707 | About a 29.3% reduction in standard error |
| n → 4n | 1/2 | Standard error is cut in half |
| n → 9n | 1/3 | Standard error is one-third as large |
Standard Error of the Mean
Standard deviation and standard error answer different questions. The population standard deviation, σ, describes how individual observations vary around the population mean. The standard error of x̄ describes how sample means vary across repeated samples of the same size.
The simulator shows both a theoretical standard error and a simulated SD of sample means. With only a few repeated samples, the simulated value can bounce around. After many repetitions, it should move closer to σ/√n, subject to ordinary Monte Carlo variation.
Sampling Variability
Sampling variability is the natural difference that appears because random samples contain different observations. It is not automatically a data-quality problem or a mistake in calculation. Even if every sample is collected correctly, their means will differ. The sampling distribution quantifies how much variation to expect from that repeated-sampling process.
For a direct treatment of this idea, see sampling variability and the guide to the sampling distribution of the sample mean.
Central Limit Theorem vs Law of Large Numbers
These results are related but answer different questions. The Law of Large Numbers concerns what happens to a sample average as the sample itself grows: under appropriate conditions, the sample mean converges toward the population mean. The Central Limit Theorem concerns the shape and scale of the sampling distribution after suitable standardization as sample size grows.
If you increase the number of simulation repetitions while holding n fixed, you are not increasing the sample size. You are simply getting a clearer picture of the same sampling distribution. If you increase n, you change the sampling distribution itself by reducing its standard error.
Common Misconceptions
- “n = 30 guarantees normality.” No. The quality of the approximation depends on the population and assumptions.
- “The population becomes normal.” No. The original population shape does not change when n increases.
- “One large sample must look normal.” No. The CLT is about the distribution of a statistic across repeated samples.
- “1,000 simulations means n = 1,000.” No. Repetitions and sample size are separate quantities.
- “Standard error is the same as standard deviation.” No. SD describes individual observations; SE describes the variability of a statistic across samples.
Related Statistics Tools and Guides
Frequently Asked Questions
The histogram bars count sample means. Each repeated sample produces one x̄, and that one mean is placed into a histogram bin. The original observations are not placed directly into the sampling-distribution histogram.
For independent observations with a finite population mean, the expected value of the sample mean is the population mean: E(X̄) = μ. With many simulated samples, the average of the simulated sample means should therefore settle near μ.
With n = 1, the sample mean equals the single observation in the sample. In theory, the sampling distribution of x̄ then has the same distribution as the population, with the same mean and standard deviation.
No. Thirty is not a universal threshold. How quickly the sample-mean distribution becomes well approximated by a normal distribution depends on population shape, skewness, tail behavior, dependence, and whether the required moments exist.
A sample mean averages n observations. Under the independent identically distributed model, its variance is σ²/n, so its standard deviation is σ/√n. Larger samples therefore produce means that vary less from sample to sample.
No. For a normal population under the usual independent model, the distribution of x̄ is exactly normal. For the other finite-variance populations in this simulator, the blue curve is a CLT normal approximation whose quality generally improves as n increases.
Sample size n is the number of observations used to calculate each sample mean. Simulation repetitions are the number of separate samples you draw. Increasing repetitions makes the simulated histogram smoother; increasing n changes the theoretical standard error and the sampling distribution itself.