- Statistics is the science of collecting, organising, analysing, and interpreting numerical data to answer questions and make decisions.
- It has two main branches: descriptive statistics (summarising data you have) and inferential statistics (drawing conclusions about populations from samples).
- Every major field — medicine, business, sports, government, science — depends on statistics to make reliable decisions.
- The core concepts build on each other: start with data types, move to averages and spread, then probability, then hypothesis testing.
- You do not need calculus to learn introductory statistics. The ideas matter more than the formulas.
The Simple Definition of Statistics
A lot of people hear "statistics" and picture dusty textbooks full of formulas. The real subject is something more practical: it's a set of tools for turning raw numbers into answers.
"Statistics is the science of collecting, organising, analysing, and interpreting numerical data to answer questions and make decisions."
Statistics converts raw numbers into conclusions you can act on.
Whenever you want to know something about the world — whether a new drug works, which advertisement performs better, how reliable a manufacturing process is — you collect data, organise it, and use statistical methods to find the answer. Statistics tells you not just what the numbers say, but how confident you should be in what they say.
Statisticians define the field as a branch of mathematics concerned with collecting data, measuring variability within that data, and making inferences about populations from sample observations, with explicit quantification of uncertainty.
Jump to the two main branches, key concepts, or how to get started. The table of contents on the right covers every section.
The Two Main Branches of Statistics
Statistics splits into two distinct jobs. The first is to describe what you already see in data. The second is to use what you see in a sample to say something about a much larger group you couldn't measure directly.
To keep the ideas grounded, imagine a maths teacher with 30 students. Their end-of-term scores are the running example throughout this section.
Descriptive Statistics — Summarising What You Have
Descriptive statistics describe and summarise a dataset you already hold. The word "descriptive" is the clue: you're not guessing at anything beyond the numbers in front of you. You're computing the clearest possible picture of them.
Descriptive Statistics
Summarises data you already have. The goal is to describe it accurately and clearly — not to generalise beyond it.
Common tools: mean, median, mode, standard deviation, variance, histograms, bar charts, scatter plots.
Those three numbers (average, range, standard deviation) give the teacher a complete summary of how the class performed, without any guessing. That's descriptive statistics doing its job.
Inferential Statistics — Drawing Conclusions from Samples
Inferential statistics start where descriptive statistics end. Instead of describing the data you have, you use that data — drawn from a sample — to make claims about a larger population you couldn't fully measure.
Inferential Statistics
Uses a sample to draw conclusions about a population. Uncertainty is always part of the picture, which is why inferential statistics produces confidence levels alongside estimates.
Common tools: hypothesis tests, confidence intervals, p-values, regression, sampling distributions.
A poll of 1,000 voters used to forecast a national election is inferential statistics. So is a clinical trial of 500 patients used to approve a drug for millions of people. The sample is small; the population is large; the inference connects them.
Descriptive vs Inferential Statistics: At a Glance
| Feature | Descriptive Statistics | Inferential Statistics |
|---|---|---|
| Goal | Summarise and describe existing data | Draw conclusions about a population from a sample |
| Data scope | Works with the full dataset you have | Uses a sample to represent a larger group |
| Common techniques | Mean, median, standard deviation, charts | Hypothesis tests, confidence intervals, regression |
| Example question | What was the average score in this class? | Do students in this school score higher than the national average? |
| Certainty | Exact (describes what you measured) | Probabilistic (estimates with stated uncertainty) |
Why Statistics Matters — Real-World Examples
Statistics isn't a subject that lives only in textbooks. It's the infrastructure behind decisions in almost every field worth naming. The examples below aren't hypothetical — they're how things actually get decided.
Medicine and Healthcare
When Pfizer reported its COVID vaccine was 95% effective, that figure came from a randomised controlled trial analysed with statistical methods. Every drug approval, dosage recommendation, and screening guideline rests on the same kind of evidence. Statistics is the reason doctors can say "this treatment works" and mean something precise by it.
Business and Economics
When Netflix decides whether to renew a series, it uses statistical models built from viewing data across hundreds of millions of subscribers. When a retailer sets prices, adjusts inventory, or designs a promotional offer, the decisions flow from A/B tests and regression models. Statistics is the language business uses to stop guessing.
Sports
The Moneyball revolution in baseball — using statistics to find undervalued players — has spread to football's expected goals (xG), basketball's player efficiency ratings, and cycling's power-to-weight analysis. Modern sports teams run statistics departments as large as their coaching staffs.
Government and Policy
Census data determines how legislative seats are allocated and how public funds are distributed. Unemployment figures, inflation rates, and public health metrics — all products of statistical surveys — shape the laws that govern everyday life. Government without statistics is government flying blind.
Science and Research
Peer-reviewed science requires that findings pass statistical significance tests before publication. When a psychology study or a physics experiment reports a result, the p-value and confidence interval tell other scientists how much weight to give the finding. Without statistics, science has no standard for deciding what counts as a real discovery.
Everyday Life
A "70% chance of rain" is a probability statement produced by a statistical model. Your car insurance premium is set using actuarial statistics. A restaurant review rating is an average of a sample of customer opinions. You interact with statistical outputs dozens of times a day, whether or not you recognise them.
For a detailed look at how statistics powers business decisions, read how statistics powers A/B testing on the blog.
Key Statistical Concepts Every Beginner Should Know
The concepts below are the building blocks. Each one gets its own page on this site with worked examples, calculators, and practice questions. Here the goal is simpler: one clear idea and one real-world anchor per concept.
Population vs. Sample
A population is every member of the group you want to study. A sample is the subset you actually measure. You almost never have access to the full population, which is why inferential statistics exists. In our running example, the 30 students are a sample; all students in the district are the population.
Learn more: Population vs Sample →Mean, Median, and Mode
Three ways to find the centre of a dataset. The mean is the arithmetic average. The median is the middle value when data is sorted. The mode is the value that appears most often. When exam scores are skewed by a handful of very low marks, the median gives a better picture of the typical student than the mean does.
Learn more: Mean, Median and Mode →Standard Deviation and Variance
These measure how spread out a dataset is. A small standard deviation means scores cluster tightly around the mean; a large one means they're scattered widely. Two classes with the same average score of 74 can look very different: one might have everyone clustered between 68 and 80, while another ranges from 30 to 100.
Learn more: Standard Deviation →Probability
The language of uncertainty. Probability assigns a number between 0 and 1 to how likely an event is. Statistics relies on probability to describe how confident we are in our conclusions. A weather forecast of 70% rain is a probability. The p-value in a hypothesis test is a probability.
Learn more: Basic Probability →Distributions
A distribution describes how data values are spread across a range. The most famous is the normal distribution — the bell curve — where most values cluster in the middle and fewer appear at the extremes. Heights, exam scores, and measurement errors all tend to follow this pattern.
Learn more: Normal Distribution →Hypothesis Testing
A formal method for deciding whether an observed pattern in data is real or just the result of random chance. You start with a null hypothesis (the boring assumption: nothing interesting happened), collect data, and calculate how unlikely that data would be if the null hypothesis were true. If it's unlikely enough, you reject the null.
Learn more: Hypothesis Testing →Confidence Intervals
Rather than claiming an exact answer, a confidence interval gives a range that likely contains the true value. "The average score is 74, with a 95% confidence interval of 69 to 79" is more honest than claiming precision that the sample size doesn't support. Most published statistics come with a confidence interval attached.
Learn more: Confidence Intervals →Correlation vs. Causation
Two variables can move together — ice cream sales and drowning rates both rise in summer — without one causing the other. Statistics can measure how strongly variables are correlated. Establishing that one causes the other requires a more careful study design, usually a controlled experiment. This distinction matters every time you read a headline about a new study.
Learn more: Correlation vs Causation →Practice with real numbers
Our free mean calculator lets you enter any dataset and see the mean, median, and mode calculated step by step.
A Worked Example: The Class Exam Scores
The running example from earlier pulls everything together. This walkthrough shows how the same dataset gets different treatment depending on whether you need descriptive or inferential answers.
Setup: A class of 30 students takes a maths exam. The teacher wants to both summarise the class results and check whether this class performed differently from the school's historical average of 70.
Collect data: The 30 scores are recorded. This is the raw dataset.
Descriptive statistics: The teacher calculates a mean of 74, a median of 75, and a standard deviation of 12.3. A bar chart shows the distribution. These numbers describe this class — nothing more.
Inferential statistics: To test whether 74 is genuinely different from the school's historical average of 70, the teacher runs a one-sample t-test. The p-value comes back at 0.04 — there's only a 4% chance of seeing a sample mean this high if the true mean were 70.
Confidence interval: A 95% confidence interval for the true class mean runs from 69.4 to 78.6. The school average of 70 sits just inside the interval, which tells the teacher the difference, while statistically significant, is not dramatic.
The same 30 scores answered two completely different questions — one descriptive (what happened in this class?) and one inferential (is this class unusual compared to the school average?). This two-step pattern appears in nearly every real statistical analysis.
How Data Is Collected in Statistics
Every statistical analysis starts with data, and data doesn't appear by magic. The method used to collect it determines what questions can honestly be answered — and how far the conclusions can reach.
There are four main collection methods in practice:
Surveys and Questionnaires
Asking people directly. Useful for opinions, preferences, and self-reported behaviour. Customer satisfaction surveys, political polls, and census forms all belong here. The risk is response bias — people may not answer truthfully, or certain groups may be less likely to respond.
Controlled Experiments
Randomly assigning participants to conditions and measuring outcomes. Clinical drug trials are the gold standard. Because random assignment controls for other variables, experiments are the best tool for establishing causation. See randomised controlled trials for a full breakdown.
Observational Studies
Watching what happens without intervening. Tracking how many hours students study versus their grades, without controlling the hours, is observational. The data can reveal patterns, but confounding variables make causal claims tricky. See study design for how researchers handle this.
Existing Records and Secondary Data
Using data already collected for another purpose. Hospital records, transaction databases, and census archives all fall here. Secondary data is cheap and large, but the researcher didn't control how it was collected.
Statistics is only as reliable as the data it's built on. A biased sample produces biased conclusions, no matter how sophisticated the analysis that follows.
Statistics vs. Mathematics — What's the Difference?
Many beginners assume statistics is just another name for maths, or that struggling with algebra means struggling with statistics. Neither is quite right.
Maths deals with certainty. The answer to 2 + 2 is always 4; the proof of Pythagoras's theorem is either correct or it isn't. Statistics deals with uncertainty. The goal is to quantify how confident you can be in a conclusion drawn from imperfect, incomplete data — not to produce exact answers.
Statistics borrows extensively from mathematics. Probability theory, calculus (for continuous distributions), and linear algebra (for regression) all appear as you go deeper. But the concepts matter far more than the formulas at the introductory level. The ideas of variation, sampling error, and significance are more important to grasp than the mechanics of computing them by hand.
If you can calculate an average, you already understand the most fundamental idea in statistics. The rest builds from there.
Statistics and Probability — How They're Connected
Probability and statistics are deeply linked, but they run in opposite directions.
Probability starts with a known model and asks: what outcomes should we expect? Roll a fair die, and probability tells you the chance of any given number is 1/6. Statistics starts with observed outcomes and asks: what does this tell us about the underlying model? Roll a die 100 times, observe the results, and statistics tells you whether the die is likely fair.
Probability is the foundation inferential statistics is built on. When a hypothesis test produces a p-value of 0.03, it means: if the null hypothesis were true, there's a 3% probability of seeing data at least this extreme. That's a probability statement. Statistics and probability are two sides of the same coin: one predicts outcomes from models; the other infers models from outcomes.
A Brief History of Statistics
Statistics as a discipline grew out of two separate traditions. In the 17th and 18th centuries, governments began collecting systematic data about populations for tax and military purposes — this is where the word "statistics" (from the Latin statisticum collegium, meaning "council of state") originated. At the same time, mathematicians like Pascal and Fermat were developing probability theory to answer questions about games of chance.
Francis Galton and Karl Pearson formalised the ideas of correlation and regression in the late 19th century. Ronald Fisher, working in the early 20th century, developed much of the framework modern researchers still use: the analysis of variance, the p-value, and the randomised controlled trial. The digital age moved statistics from printed tables and hand calculation to the computational tools that now make it accessible to anyone with a laptop.
How to Get Started with Statistics
Statistics builds in a clear sequence. Skipping steps makes later topics harder than they need to be. Here's a practical path that works for students, career-changers, and curious readers alike.
Start with Types of Data and Descriptive Statistics
Learn what data types exist (categorical vs numerical, discrete vs continuous) before anything else. Then work through mean, median, mode, standard deviation, and the main chart types. These come up in every subsequent topic. Start at descriptive statistics.
Understand Probability
Probability is the bridge between descriptive and inferential statistics. Spend real time here before moving on. The ideas of independent events, conditional probability, and probability distributions will make hypothesis testing much clearer. Begin at basic probability.
Explore Distributions
The normal distribution is the most important shape in all of statistics — most inference relies on it directly or indirectly. The binomial and t-distributions follow naturally from there.
Learn Hypothesis Testing
This is where you move from describing data to drawing conclusions. Work through null and alternative hypotheses, then p-values, then specific tests for different situations.
Practice with Real Tools
Use our free calculators and visual tools as you learn. Seeing numbers change in real time builds intuition faster than reading alone. The bell curve generator, standard deviation visualiser, and p-value visualiser are good starting points.
Free statistics calculators
From mean and standard deviation to confidence intervals and hypothesis tests — all free, all with step-by-step outputs.
Frequently Asked Questions
Statistics is the mathematical discipline for drawing valid conclusions from data. Data science applies statistics alongside programming, machine learning, and domain knowledge to extract insights from large, often messy datasets. Statistics asks whether a finding is real; data science builds the pipelines to surface those findings at scale. If you're learning statistics, you're building the foundation that makes data science possible.
The concepts in statistics matter more than the formulas. If you understand what an average represents, you're already thinking statistically. The challenge for many beginners isn't the mathematics — it's building intuition for what uncertainty means and when a result is genuinely surprising. That intuition comes from working through real examples, which is exactly what each page on this site provides.
Weather forecasts, insurance premiums, medical test results, election polls, quality control on manufactured goods, A/B tests for website designs, and sports performance ratings all depend on statistics. The number attached to almost any real-world decision — a percentage, a rate, a probability — was produced by a statistical method of some kind.
Descriptive statistics and inferential statistics. Descriptive statistics summarise and describe data you already have — an average, a chart, a measure of spread. Inferential statistics use a sample to draw conclusions about a larger population, with explicit statements about how confident you are in those conclusions. The two work together: you describe the sample first, then use inference to generalise from it.
In research, statistics provides the tools to design studies, collect data systematically, analyse results, and decide whether findings are genuine or due to chance. Every published scientific study reports statistical results — usually a test statistic, a p-value, and a confidence interval — that allow other researchers to evaluate how much weight to give the conclusions. Without statistics, research would have no shared language for distinguishing signal from noise.
For introductory and applied statistics — the level that covers everything on this site — no. You need arithmetic, basic algebra, and the ability to interpret formulas rather than derive them. Calculus becomes relevant if you pursue mathematical statistics at the graduate level, where probability density functions and moment-generating functions require integration. For practical data analysis, tools like calculators and software handle the computation.
A parameter describes a characteristic of a whole population — the true average height of all adults in a country, for instance. A statistic describes a characteristic of a sample drawn from that population — the average height of 500 surveyed adults. Because populations are usually too large to measure fully, statistics estimate parameters. The distinction matters because statistics have sampling error; parameters, by definition, do not.
Excel covers most basic descriptive statistics and some inferential tests. R is the standard for academic statistics — free, powerful, and used across nearly every discipline. Python with libraries like NumPy, SciPy, and statsmodels is increasingly common, especially when statistics is part of a broader data science or machine learning pipeline. SPSS is widely used in social science research; SAS appears in pharmaceutical and large corporate settings. Most beginners start with Excel or simple online calculators, then move to R or Python as their needs grow.
Continue Learning
Every major concept introduced on this page has its own dedicated section. The links below cover the natural next steps, whether you want to go deeper on a specific method or explore how statistics applies in a particular field.