Who Invented Probability?
The short answer is that probability theory was built by several people across more than a century. Gerolamo Cardano, a 16th-century Italian physician and mathematician, wrote the earliest systematic study of dice probability around 1560. His work, Liber de Ludo Aleae (Book on Games of Chance), sat unpublished for over a century, and contained errors, so it had little direct influence on what came next.
The decisive moment came in the summer of 1654. Blaise Pascal, a French mathematician, received a set of gambling questions from the Chevalier de Méré. Pascal began corresponding with Pierre de Fermat, and their letters worked out how to divide a prize fairly when a game is interrupted. That exchange introduced mathematical expectation, the idea of working from future possibilities rather than past results. Most historians date the birth of probability theory to those letters.
Interactive History of Probability Timeline
Click any milestone to expand it. Use the era filters to narrow the view. The timeline extends past 1900, where most competing resources stop.
Why Did Probability Develop So Late?
People played dice games for at least three millennia before anyone wrote down a mathematical theory of chance. Knucklebones (astragali) and marked dice appear in ancient Egyptian, Greek, and Roman contexts. The gap puzzles historians. A few explanations have been proposed, though none is universally accepted.
Ian Hacking, in The Emergence of Probability, argued that ancient and medieval philosophy divided knowledge into certain truths and mere opinion. Probability, which quantifies the in-between, required a conceptual space that did not yet exist. Outcomes of dice were seen as controlled by fate, luck, or the will of the gods, which placed them outside the domain of mathematics.
There were also practical obstacles. Ancient dice were often made from knucklebones, which are not symmetric. Unequal faces produce unequal probabilities, and you cannot do classical probability theory cleanly if outcomes are not equally likely. Even many manufactured ancient dice were poorly shaped. The equal-outcomes assumption that Cardano used required reasonably fair dice.
Combinatorics, the branch of mathematics needed to count outcomes, was also underdeveloped before the Renaissance. Without convenient notation and the ability to enumerate cases systematically, solving even simple problems was difficult. The commercial revolution of the 14th and 15th centuries, which created demand for annuities, marine insurance, and loans, gave mathematicians both the problems and the patrons that probability theory needed.
Who Is the Father of Probability?
The question gets different answers depending on which criterion matters most. Here are the main candidates and the case for each.
| Candidate | Case | Limitation |
|---|---|---|
| Gerolamo Cardano (1501-1576) | First systematic treatment of dice probability; introduced the favorable-to-total ratio as the basic definition of chance | Work published a century after writing; contained errors; had little influence on his contemporaries |
| Blaise Pascal (1623-1662) and Pierre de Fermat (c. 1601 or 1607-1665) | Solved the problem of points; introduced mathematical expectation; their methods spread through Europe via Huygens and others | Built on problems already circulating; their work was a correspondence, not a published treatise |
| Christiaan Huygens (1629-1695) | Published De Ratiociniis in Ludo Aleae (1657), the first book on probability, making the theory widely available | Explicitly built on what Pascal and Fermat had done |
| Andrey Kolmogorov (1903-1987) | Gave probability its modern axiomatic foundation in 1933, putting it on the same rigorous basis as the rest of mathematics | Formalized what existed; did not originate the field |
Pascal and Fermat are most commonly credited as the founders of probability theory. Cardano is usually called its pioneer. Historian Oystein Ore, in Cardano: The Gambling Scholar (1953), argued that Cardano deserves the founding title. Both positions have scholarly support. The label "father of modern probability" is sometimes applied to Kolmogorov for his axioms.
The Problem That Started It All: The Problem of Points
In 1654, the Chevalier de Méré asked Pascal how to fairly divide the pot when a game is stopped early. This question, the problem of points, had been posed by Luca Pacioli in 1494 and left unanswered correctly for 160 years. Pascal's exchange with Fermat solved it and, in doing so, introduced the concept of mathematical expectation.
Pacioli proposed dividing the pot in proportion to rounds already won. If the score is 2-1, Player A gets 2/3 and Player B gets 1/3. But this ignores how close each player is to winning. Pascal and Fermat showed the split should depend on the rounds still needed.
Problem of Points: First to 3 Wins, Score 2-1
Player A needs 1 more win. Player B needs 2. Each round is fair (50-50). How should the pot be divided?
Find the maximum remaining rounds: If A needs 1 and B needs 2, the game ends within 2 more rounds (1 + 2 - 1 = 2).
List all equally likely outcomes (Fermat's method) for 2 rounds: AA, AB, BA, BB. Each has probability 1/4.
Determine the winner in each: AA (A wins, 1 win needed); AB (A wins after round 2); BA (A wins after round 1); BB (B wins, 2 wins needed). A wins in 3 out of 4 scenarios.
Calculate probabilities: P(A wins) = 3/4. P(B wins) = 1/4. The fair split of the pot is 75% to A, 25% to B.
Fair division: A receives 3/4 of the pot. B receives 1/4. This follows from each player's probability of winning from the current state.
Try the calculator below with your own values. The formula uses the binomial distribution: P(A wins) is the sum of C(a+b-1, k) times p^k times (1-p)^(a+b-1-k) for k from a to a+b-1.
Problem of Points Calculator
The problem of points is where expected value and the binomial distribution meet history. Every time you calculate what a gamble is worth, you are using the idea Pascal and Fermat worked out in 1654.
The Chevalier de Méré's Dice Problem
The same Chevalier who prompted Pascal also posed a puzzle that had cost him money at the tables. He believed two bets were equally favorable. Pascal showed him why one was and the other was not.
Two Dice Bets: Which Is Better?
Bet 1: At least one six in 4 rolls of one die. Bet 2: At least one double-six in 24 rolls of two dice.
Bet 1: P(no six in one roll) = 5/6. P(no six in 4 rolls) = (5/6)^4 = 625/1296 ≈ 0.4823. So P(at least one six) = 1 - 0.4823 = 0.5177.
Bet 2: P(no double-six in one roll) = 35/36. P(no double-six in 24 rolls) = (35/36)^24 ≈ 0.5086. So P(at least one double-six) = 1 - 0.5086 = 0.4914.
Bet 1 is favorable (probability 0.5177). Bet 2 is slightly unfavorable (probability 0.4914). De Méré's proportional reasoning was wrong. The correct method uses the complement rule.
de Méré Dice Simulator
Run many trials of both bets to watch the simulated proportions converge toward the theoretical values. This is a simulation, not a proof.
As trials increase, the simulated rates approach 51.77% and 49.14%. This illustrates the law of large numbers.
Galileo and the Three-Dice Puzzle
Around 1613-1623, a Florentine nobleman asked Galileo why a sum of 10 appears more often than 9 when three dice are thrown, even though both totals can be made from exactly six unordered combinations. Galileo's answer hinged on the difference between ordered and unordered outcomes.
For example, 1+2+6 = 9 appears in six ordered arrangements (1,2,6), (1,6,2), (2,1,6), (2,6,1), (6,1,2), (6,2,1). But 3+3+3 = 9 appears in only one. Adding up all the ordered ways to make each total gives 25 for sum 9 and 27 for sum 10, out of 6^3 = 216 equally likely outcomes. This is why 10 wins. The key insight, counting ordered outcomes, underlies all of modern counting methods and permutations and combinations.
Three Dice Sum Explorer
Click any bar to see the number of ordered ways to make that sum out of 216 total outcomes.
The History of Probability, Era by Era
Ancient Games of Chance
Astragali (knucklebones), polished flat stones, and marked dice appear in archaeological records from ancient Egypt, Mesopotamia, Greece, and Rome. They were used in games and in divination. No mathematical theory accompanied them. Outcomes were interpreted as divine messages or expressions of fate, not as quantities to be calculated.
Renaissance Beginnings
Luca Pacioli included the problem of points in his 1494 Summa de arithmetica but gave the wrong answer. Gerolamo Cardano's Liber de Ludo Aleae, written around 1560 but unpublished until 1663, defined probability as the ratio of favorable outcomes to total equally likely outcomes and worked through many dice problems. Niccolò Tartaglia and Galileo Galilei also addressed dice questions in this period. Galileo's three-dice analysis (c. 1613-1623) was the most careful counting argument before Pascal.
Pascal and Fermat (1654): The Founding Letters
Between July and October 1654, Blaise Pascal and Pierre de Fermat exchanged letters on the problem of points and de Méré's dice puzzles. Pascal's method used what we now call Pascal's triangle and a backward-induction argument. Fermat's method enumerated all possible future outcomes. Both arrived at the same answers. Their correspondence introduced mathematical expectation: the fair price of a gamble equals the weighted average of its possible outcomes.
Christiaan Huygens heard of the work and in 1657 published De Ratiociniis in Ludo Aleae, the first book on probability, making these ideas widely available. Huygens framed the subject around the concept of expected value, which he called the value of a chance (valor sortis).
Probability Meets Data
John Graunt's 1662 analysis of London's Bills of Mortality was the first systematic use of population data to draw statistical inferences. He noticed regularities: the ratio of male to female births, the age structure of deaths, and seasonal patterns in mortality. Edmond Halley produced the first actuarially usable life table in 1693, which allowed annuity prices to be set mathematically. These applications pulled probability out of the gaming table and into public affairs and commerce.
The Law of Large Numbers and the Bell Curve
Jacob Bernoulli's Ars Conjectandi, published posthumously in 1713, proved the first version of the law of large numbers: as the number of trials grows, the observed frequency of an event converges to its true probability. This gave the frequentist interpretation of probability its theoretical backing. Abraham de Moivre, in The Doctrine of Chances (1718, expanded 1733), discovered the normal approximation to the binomial distribution, which is the first appearance of the bell curve in probability theory.
Bayes and Inverse Probability
Thomas Bayes wrote an essay on inverse probability that was read to the Royal Society in 1763, two years after his death, by Richard Price. Bayes asked the reverse of the usual question: given observed data, what can you infer about an unknown probability? Pierre-Simon Laplace independently developed the same ideas more fully, publishing the most comprehensive treatment of probability of his era in Théorie analytique des probabilités (1812). Laplace also gave what is now called the classical definition: probability is the number of favorable cases divided by the total number of equally possible cases.
Rigor, Social Statistics, and Markov Chains
Siméon Denis Poisson derived the Poisson distribution in 1837, applicable to rare events. Adolphe Quetelet applied probability to social data, introducing the concept of the "average person" and showing that social phenomena followed statistical regularities. Pafnuty Chebyshev proved a general form of the law of large numbers using only moment conditions. His student Andrey Markov extended the theory of dependent random variables, introducing what are now called Markov chains: sequences where the next state depends only on the current state.
Kolmogorov's Axioms and the Modern Era
In 1933, Andrey Kolmogorov published Grundbegriffe der Wahrscheinlichkeitsrechnung (Foundations of the Theory of Probability), which placed probability on a rigorous mathematical footing using measure theory. His three axioms: probabilities are non-negative, the total probability of all outcomes is 1, and probabilities of mutually exclusive events add, are the basis of every modern probability textbook. In the 1940s, Stanislaw Ulam, John von Neumann, and Nicholas Metropolis developed Monte Carlo methods, using random sampling to solve problems that were analytically intractable. Probabilistic machine learning now permeates science, engineering, medicine, and technology.
Key Mathematicians in the History of Probability
| Mathematician | Life Dates | Nationality | Main Contribution |
|---|---|---|---|
| Luca Pacioli | 1447-1517 | Italian | Posed the problem of points in 1494; solution was incorrect |
| Gerolamo Cardano | 1501-1576 | Italian | Liber de Ludo Aleae (c. 1560, pub. 1663): first systematic dice probability; favorable/total ratio |
| Galileo Galilei | 1564-1642 | Italian | Three-dice puzzle (c. 1613-1623): demonstrated that ordered, not unordered, outcomes must be counted |
| Blaise Pascal | 1623-1662 | French | 1654 letters with Fermat: problem of points, expectation; Pascal's triangle applied to probability |
| Pierre de Fermat | c. 1601 or 1607-1665* | French | 1654 letters with Pascal: enumeration method for problem of points |
| Christiaan Huygens | 1629-1695 | Dutch | De Ratiociniis in Ludo Aleae (1657): first published probability text; systematic treatment of expectation |
| John Graunt | 1620-1674 | English | 1662 mortality table analysis: first demographic statistics from data |
| Jacob Bernoulli | 1655-1705 | Swiss | Ars Conjectandi (1713, posthumous): law of large numbers; binomial theorem applications |
| Abraham de Moivre | 1667-1754 | French (in England) | The Doctrine of Chances (1718, 1733): normal approximation to the binomial; first bell curve |
| Thomas Bayes | c. 1701-1761 | English | Essay on inverse probability (pub. 1763): foundation of Bayesian inference |
| Pierre-Simon Laplace | 1749-1827 | French | Théorie analytique des probabilités (1812): classical definition; central limit theorem ideas; applied probability |
| Siméon Denis Poisson | 1781-1840 | French | Poisson distribution (1837): model for rare events |
| Pafnuty Chebyshev | 1821-1894 | Russian | General law of large numbers; Chebyshev's inequality |
| Andrey Markov | 1856-1922 | Russian | Markov chains: dependent random processes with the memoryless property |
| Andrey Kolmogorov | 1903-1987 | Soviet | Grundbegriffe der Wahrscheinlichkeitsrechnung (1933): axiomatic foundations of modern probability |
* Sources disagree on Fermat's birth year, citing both 1601 and 1607. The MacTutor History of Mathematics at St Andrews notes 1607 as more likely but acknowledges the uncertainty.
How the History Connects to Modern Probability
Every topic in a first probability course has a historical origin. The table below maps the history onto the concepts you are likely studying, with links to the relevant pages on this site.
| Historical Idea | Modern Concept | Learn It Here |
|---|---|---|
| Cardano's favorable/total ratio | Classical probability | Basic Probability |
| Galileo's ordered outcome counting | Counting and permutations | Permutations & Combinations |
| Problem of points (Pascal & Fermat) | Expected value | Expected Value |
| Pascal's triangle applied to games | Binomial coefficients | Binomial Distribution |
| Bernoulli's law of large numbers | Convergence of frequency to probability | Law of Large Numbers |
| De Moivre's normal approximation | Normal distribution | Normal Distribution / Normal Approximation |
| Laplace's central limit ideas | Central limit theorem | Central Limit Theorem |
| Bayes and Price: inverse probability | Bayes' theorem, prior and posterior | Bayes' Theorem / Prior Probability |
| Frequentist vs Bayesian interpretation debate | Philosophical foundations | Bayesian vs Frequentist |
| De Méré's complement rule problem | Probability rules | Probability Rules |
| Pascal's triangle outcome tree | Probability trees | Probability Trees |
| Markov chains (Markov, 1900s) | Monte Carlo simulation | Markov Chain Monte Carlo |
Interpretations of Probability
One reason the history of probability is complicated is that the word "probability" was used to mean different things by different thinkers, and those tensions have never been fully resolved.
The classical interpretation (Laplace): probability is the number of favorable outcomes divided by the number of equally possible outcomes. It applies only when outcomes are symmetric and equally likely, which limits it.
The frequentist interpretation (Venn, von Mises): probability is the long-run relative frequency of an event in an indefinitely repeated series of trials. It connects probability to observation and avoids subjectivity, but it cannot be applied to one-off events.
The Bayesian or subjective interpretation (Ramsey, de Finetti): probability is a degree of belief that can be updated with evidence using Bayes' theorem. It applies to any uncertain proposition, including one-off events, but requires specifying a prior.
The axiomatic approach (Kolmogorov): probability is a mathematical function satisfying three axioms. It is agnostic about interpretation and compatible with all three views above. Modern mathematics uses Kolmogorov's framework.
Myths About the History of Probability
| Common Claim | What the Evidence Shows |
|---|---|
| Probability was invented to settle a single gambling dispute | De Méré's questions were the trigger, but Pascal and Fermat were solving a class of problems with a long history (Pacioli, Tartaglia, Cardano). The theory grew from multiple conversations and problems. |
| Pascal alone invented probability | Fermat's contribution was equal and, by some accounts, deeper. Cardano preceded both. Huygens spread the work. |
| Ancient Romans and Greeks understood probability mathematically | They used chance devices extensively but left no evidence of a mathematical theory. No ancient text treats probability as a calculated quantity. |
| Kolmogorov invented probability | Kolmogorov formalized probability rigorously in 1933. The subject had been developed by many mathematicians over three centuries before him. |
| Bayes' theorem was Bayes's alone | Laplace independently developed the same result and published the more general and more influential version. Price, who published Bayes's essay posthumously, also contributed substantially to the framing. |
History of Probability Quiz
Test Your Knowledge
Further Reading
The books below are arranged by difficulty. All are either primary sources or highly regarded scholarly works.
Popular and Accessible
Intermediate History
Scholarly and Philosophical
Primary Sources
Frequently Asked Questions
No single person invented probability. Gerolamo Cardano wrote the earliest systematic dice mathematics around 1560. Blaise Pascal and Pierre de Fermat founded the formal theory in their 1654 letters. Christiaan Huygens published the first book on probability in 1657. Andrey Kolmogorov gave probability its modern axiomatic foundation in 1933.
Pascal and Fermat are most commonly credited as the founders of probability theory because their 1654 work introduced expectation and spread through Europe. Cardano is usually called the pioneer, and historian Oystein Ore argued he deserves the founding credit. Kolmogorov is sometimes called the father of modern probability for his 1933 axioms.
Most historians date the beginning of formal probability theory to 1654, when Pascal and Fermat exchanged letters on the problem of points. Cardano's earlier work (c. 1560) is the most important precursor but had little influence at the time because it was not published until 1663.
The problem of points asks how to fairly divide a prize when a game of chance is stopped before a winner is decided. Luca Pacioli posed it in 1494 but solved it incorrectly. Pascal and Fermat solved it correctly in 1654 by calculating each player's probability of winning from the current state, which introduced mathematical expectation.
Christiaan Huygens's De Ratiociniis in Ludo Aleae (1657) was the first published book on probability. Cardano's Liber de Ludo Aleae was written earlier (c. 1560) but not published until 1663. Jacob Bernoulli's Ars Conjectandi (1713) was the first comprehensive probability text.
Scholars have proposed several reasons: outcomes were attributed to gods or fate rather than mathematics; early dice were irregular in shape; combinatorics was underdeveloped; and the philosophical divide between certain knowledge and mere opinion did not leave room for a theory of partial belief (Ian Hacking's argument). Commercial growth in insurance and annuities eventually created the demand for probability.
Bernoulli's Ars Conjectandi (1713, published after his death in 1705) proved the first rigorous version of the law of large numbers: as the number of trials increases, the observed frequency of an event converges to its true probability. He also contributed to binomial probability and discussed how probability could be applied to civil and moral life.
Andrey Kolmogorov published Grundbegriffe der Wahrscheinlichkeitsrechnung in 1933, which established probability on an axiomatic basis using measure theory. His three axioms (non-negativity, normalization to 1, and countable additivity) are the foundation of all modern probability mathematics. Every advanced probability text uses Kolmogorov's framework.
Ancient civilizations used dice and other chance devices extensively, but there is no evidence they developed a mathematical theory of probability. They understood that some outcomes were more common than others in a practical sense, but they attributed outcomes to fate or divine will rather than calculating probabilities.