Descriptive Statistics Distribution Shape Probability Distributions 30 min read Updated August 24, 2026
BY: Statistics Fundamentals Team
Reviewed By: Minsa A (Senior Statistics Editor)

Skewness and Kurtosis Explained: Formulas, Types, and Interpretation Guide

Two investment funds both average a 7% annual return with the same standard deviation. One consistently earns close to 7%. The other sometimes loses 40% and sometimes gains 60%. The mean and variance say they are identical — but they are not. Skewness and kurtosis are the two numbers that reveal the difference.

This guide covers definitions, symbols (γ₁, β₂), formulas, all three types of skewness and kurtosis, how to interpret values, acceptable ranges, APA reporting format, real-world case studies, and an interactive calculator at the bottom for your own data.

What You'll Learn
  • ✓ Symbols γ₁, g₁, β₂, g₂ — which to use and why
  • ✓ Fisher-Pearson formulas for skewness and excess kurtosis, explained
  • ✓ Types of skewness (positive, negative, zero) and kurtosis (leptokurtic, platykurtic, mesokurtic)
  • ✓ How to interpret values — acceptable ranges by field
  • ✓ How to report in APA format with standard errors (SES / SEK)
  • ✓ Alternative formulas: Galton, Pearson 2, and the Jarque-Bera test
  • ✓ Real-world case studies in finance, income, and exam scores
  • ✓ Interactive calculator — try it with your own dataset

The Shape Language of Data

The mean tells you where the center of your data sits. The standard deviation tells you how spread out the values are. These two numbers summarize most distributions well enough for everyday use — but they say nothing about shape.

Shape matters. A distribution can be symmetric or lopsided, can have thin tails or thick ones, can have most of its weight near the center or concentrated in the extremes. Two datasets with identical means and standard deviations can behave completely differently in practice. According to the NIST/SEMATECH Engineering Statistics Handbook, the third and fourth standardized moments — skewness and kurtosis — are the primary tools for detecting departures from normality and for understanding the shape of a distribution beyond its center and spread.

0
Skewness of a perfect normal distribution
3
Raw kurtosis of the normal distribution (0 in excess form)
4th
Statistical moment kurtosis is derived from
3rd
Statistical moment skewness is derived from

What is Skewness? (Direction of Data Asymmetry)

Definition — Third Standardized Moment
Skewness measures the asymmetry of a probability distribution around its mean. A positive value means the right tail is longer; a negative value means the left tail is longer; zero means the distribution is symmetric.
γ₁ = E[(X − μ)³] / σ³

Skewness Symbol (γ₁ / g₁)

Skewness uses two symbols depending on whether you are working with a population or a sample:

γ₁
Population skewness
Gamma-one (Greek)
g₁
Sample skewness
Fisher-Pearson estimate
Skew(X)
Functional notation
Used in software docs
b₁
Moment ratio
Alternative notation

In most academic writing you will see γ₁ for the population parameter and g₁ for the sample statistic. Excel's SKEW(), Python's scipy.stats.skew(), and SPSS all compute g₁ (the bias-corrected sample version) by default.

Skewness and Kurtosis Formulas

The formula for skewness expresses it as the third standardized moment. Take every value's deviation from the mean, cube it (which preserves the sign, unlike squaring), average those cubed deviations, and divide by the cubed standard deviation to make the result unitless. The cube operation is what gives skewness its directional sensitivity.

Sample Skewness Formula (Fisher–Pearson)
g₁ = [n / ((n−1)(n−2))] × Σ[(xᵢ − x̄) / s]³
Recommended for most sample data; corrects for small-sample bias
g₁ = sample skewness n = sample size xᵢ = each data point x̄ = sample mean s = sample standard deviation
Sample Excess Kurtosis Formula (Fisher's Definition)
Kurt = [(n(n+1)) / ((n−1)(n−2)(n−3))] × Σ[(xᵢ−x̄)/s]⁴ − [3(n−1)² / ((n−2)(n−3))]
Excess kurtosis = raw kurtosis − 3. Zero means same tail behavior as the normal distribution.
β₂ = raw kurtosis (normal = 3) β₂ − 3 = excess kurtosis (normal = 0) n = sample size s = sample standard deviation

Most statistical software — including SPSS and Excel's SKEW() function — uses the Fisher–Pearson formulation for skewness. The correction factor n / ((n−1)(n−2)) matters most for small datasets. For large samples it approaches 1 and the simpler population formula gives essentially the same result.

💡
Interpretation Rule of Thumb

A skewness value between −0.5 and +0.5 is generally considered approximately symmetric. Values between ±0.5 and ±1.0 indicate moderate skewness. Values beyond ±1.0 indicate substantial skewness that materially affects how you should interpret the mean. These thresholds are conventions, not laws — always look at a histogram alongside the number.

Types of Skewness and Kurtosis

Both skewness and kurtosis have three named types based on the direction or magnitude of their values relative to the normal distribution baseline.

Positive Skew (Right-Skewed Data)

Positive / Right Skew

Long tail extends right

Most data clusters on the left; a smaller number of very high values stretches the distribution rightward. The mean is pulled above the median.

Mean vs Median

Mean > Median > Mode

The extreme values in the right tail inflate the mean. The median is a more reliable summary of a typical value in this case.

Histogram showing a right-skewed (positively skewed) distribution with a long tail extending to the right. The mean is pulled right of the median, which sits right of the mode.

Income distribution in most countries shows clear positive skewness. The majority of households earn near the median — for the United States, the Census Bureau reports a consistent 20–25% gap between mean and median household income because a smaller number of very high earners pulls the mean upward. Other common examples include housing prices, wealth distribution, and waiting times for rare events.

📊 Real-World Example

U.S. Household Income (Positive Skew)

According to the U.S. Census Bureau's Current Population Survey, the 2023 mean household income was approximately $105,000 while the median was around $77,000 — a gap of nearly $28,000. This difference is a direct consequence of the right-skewed income distribution. The mean is being pulled upward by households earning several hundred thousand dollars or more per year. The median better represents what a randomly selected household actually earns.

Negative Skew (Left-Skewed Data)

Negative / Left Skew

Long tail extends left

Most data clusters on the right; a smaller number of very low values pulls the distribution leftward. The mean falls below the median.

Mean vs Median

Mean < Median < Mode

The extreme low values in the left tail drag the mean downward. The median resists this distortion.

Histogram showing a left-skewed (negatively skewed) distribution with a long tail extending to the left. Most data clusters near the high end, with the mean pulled below the median.

Negatively skewed distributions appear wherever there is a natural ceiling. Human lifespan in high-income countries is left-skewed: most people survive into their 70s or 80s, while a smaller number die young and pull the mean below the most common life expectancy. Easy exam scores — where most students score in the 80s and 90s — produce the same shape.

Zero Skew (Symmetry)

A skewness of zero describes a distribution that is perfectly symmetric around its mean. The normal distribution is the most familiar example: its bell curve is a mirror image on either side of the mean, and mean = median = mode all sit at the same point.

⚠️
Zero skewness ≠ normality

A dataset can have zero skewness while being far from normally distributed. A perfectly bimodal distribution (two equal humps on either side of the center) has zero skewness but is obviously not normal. Skewness tests symmetry; it does not test normality. Use the Shapiro–Wilk test or a Q-Q plot for normality assessment.

What is Kurtosis? (Tail Intensity of Data)

Definition — Fourth Standardized Moment
Kurtosis measures how heavy or light the tails of a distribution are relative to the normal distribution. It describes the probability of extreme values — how often the distribution produces outliers. A common misconception is that kurtosis measures the height of the peak; it does not. It measures the weight of the tails.
β₂ = E[(X − μ)⁴] / σ⁴

Kurtosis Symbol (β₂ / g₂)

Kurtosis also uses Greek-based symbols that differ between population and sample contexts:

β₂
Raw kurtosis
Beta-two (population)
β₂ − 3
Excess kurtosis
Also written γ₂
g₂
Sample excess kurtosis
Fisher-corrected estimate
Kurt(X)
Functional notation
Used in textbooks/software

The normal distribution has a raw kurtosis (β₂) of 3. Most software reports excess kurtosis (raw kurtosis minus 3), so the normal distribution's reference value is 0. Excel's KURT() function, Python's scipy.stats.kurtosis() by default, and SPSS all report excess kurtosis. When a research paper says "kurtosis = 0.5" they almost certainly mean excess kurtosis.

🔴
The Most Common Kurtosis Misconception

Kurtosis is widely described as measuring the "peakedness" of a distribution. This is incorrect. As statistician Peter Westfall demonstrated in a 2014 paper in The American Statistician, kurtosis is driven almost entirely by tail behavior, not by the shape of the peak. A distribution can be flat-topped and still have high kurtosis if its tails are fat enough.

Leptokurtic Distribution (Heavy Tails, Positive Kurtosis)

Leptokurtic — Excess Kurtosis > 0 — Positive Kurtosis

Heavier tails than normal

Extreme values occur more often than a normal distribution would predict. The distribution has more observations in its tails and near its center relative to the shoulders. In practice this means outliers are not rare events — they happen with meaningful regularity.

Daily stock market returns are the textbook example of leptokurtosis. Financial economists have known since Benoit Mandelbrot's 1963 research that stock return distributions have significantly heavier tails than the normal distribution. Crashes and rallies that a normal model would call "6-sigma events" appear in real markets every decade — the so-called "fat tail" problem.

Platykurtic Distribution (Light Tails, Negative Kurtosis)

Platykurtic — Excess Kurtosis < 0 — Negative Kurtosis

Lighter tails than normal

Extreme values are less common than the normal distribution predicts. More of the distribution's mass sits in the shoulders rather than the tails. Outcomes are more uniformly distributed — the data doesn't wander very far from the center as often as a bell curve would suggest.

The uniform distribution (where every outcome in a range is equally likely) is the clearest example of a platykurtic distribution, with excess kurtosis of −1.2. Rolling a single fair die produces a uniform distribution: each face appears with equal probability and the outcome never strays in any unusual way. Quality-controlled manufacturing processes often produce platykurtic output distributions.

Mesokurtic (Normal Baseline)

A mesokurtic distribution has excess kurtosis equal (or approximately equal) to zero — its tails behave like the normal distribution. The normal distribution itself is the reference case. IQ scores, heights, and many measurement errors approximate this shape on large samples.

Difference Between Skewness and Kurtosis

These two measures are frequently confused — partly because both describe departures from the normal distribution, and partly because the words "shape" and "distribution" appear in explanations of both. The distinction is exact and worth memorizing:

📐 The Data Shape Lens Model

Two Different Lenses on the Same Distribution

↔️

Skewness — The Direction Lens

Asks: "Which way does the data lean?" It measures asymmetry — whether the distribution is pulled toward higher values (right), lower values (left), or neither. It is derived from the third moment. A positive result means the right tail is longer; a negative result means the left tail is longer.

⚖️

Kurtosis — The Tail Lens

Asks: "How extreme are the extremes?" It measures tail weight — how often the distribution produces values far from the mean. It is derived from the fourth moment. A positive excess kurtosis means outliers are more common than in a normal distribution; a negative value means they are rarer.

These two properties are independent. Any combination is possible. The table below provides a detailed eight-dimension comparison:

Dimension Skewness (γ₁ / g₁) Excess Kurtosis (β₂−3 / g₂)
Core purposeMeasures asymmetry — left/right lean of the distributionMeasures tail heaviness — how common extreme values are
Statistical momentThird standardized momentFourth standardized moment
Normal distribution baseline0 (perfectly symmetric)0 in excess form; 3 in raw form
Positive values meanRight tail is longer (right-skewed); mean > medianHeavier-than-normal tails (leptokurtic); more outliers
Negative values meanLeft tail is longer (left-skewed); mean < medianLighter-than-normal tails (platykurtic); fewer outliers
Primary applicationChoosing between mean and median as center measureRisk modeling, outlier frequency, normality assumption checking
Practical implicationHigh skewness → report median, not meanHigh kurtosis → normal-based models underestimate tail risk
Real-world exampleIncome distribution (positive); exam scores on easy test (negative)Stock returns (positive leptokurtic); uniform die rolls (negative platykurtic)

Student's t-distribution with 5 degrees of freedom is perfectly symmetric (zero skewness) but substantially leptokurtic — this confirms that skewness and kurtosis measure genuinely different things.

How to Interpret Skewness and Kurtosis Values

Once you have computed skewness and kurtosis values, you need benchmark ranges to judge whether they indicate a problem. The thresholds below reflect conventions used across statistics, social science, finance, and data science. No universal cutoffs exist — always use judgment alongside the numbers and check a histogram.

Acceptable Range of Skewness and Kurtosis

Value Range Skewness Interpretation Recommended Action
|g₁| < 0.5 Approximately symmetric Mean is a reliable center measure; most parametric tests safe to proceed
0.5 ≤ |g₁| < 1.0 Moderate skewness Check histogram; mean may overstate or understate typical value; consider reporting median alongside
|g₁| ≥ 1.0 Substantial skewness Report median as primary center measure; consider log or square-root transformation before modeling
|g₁| ≥ 2.0 Severe skewness Distribution requires transformation or non-parametric methods; normality assumption likely violated
Excess Kurtosis Range Kurtosis Interpretation Recommended Action
|g₂| < 1.0 Approximately mesokurtic Tail behavior similar to normal; standard parametric methods appropriate
1.0 ≤ g₂ < 3.0 Moderate leptokurtosis Check for outliers; use robust standard errors in regression; consider t-distribution for modeling
g₂ ≥ 3.0 Strong leptokurtosis Normal-based risk models substantially underestimate tail probability; use fat-tail distributions
g₂ < −1.0 Platykurtic Tails lighter than normal; extreme values rarer than expected; uniform-like distribution
📋
Field-Specific Thresholds

Social science and psychology often use |skewness| < 2 and |excess kurtosis| < 7 as acceptable ranges for parametric tests (George & Mallery, 2010). Finance is more conservative — any excess kurtosis above 1 is typically flagged in risk modeling. Educational measurement uses ±2 for both. Check the conventions of your specific field.

How to Interpret Skewness and Kurtosis: Worked Examples

Worked Example 1 — Identifying Skewness from a Dataset

Seven annual bonuses paid to employees: $2,000 / $2,500 / $3,000 / $3,200 / $3,500 / $4,000 / $18,000

1

Calculate the mean: (2000 + 2500 + 3000 + 3200 + 3500 + 4000 + 18000) / 7 = 36,200 / 7 ≈ $5,171

2

Find the median: Sorted, the middle value (4th of 7) is $3,200

3

Compare: Mean ($5,171) > Median ($3,200). The mean is being pulled right by the $18,000 bonus.

4

Interpret: The dataset is positively (right) skewed. Using the range table above — skewness ≈ 1.7 falls in the "substantial" range. The median of $3,200 is the better central measure.

✓ Positive skewness (g₁ ≈ 1.7). Mean ($5,171) substantially overestimates typical earnings. Median ($3,200) is the more informative central measure here. Reporting the mean without noting the skewness would be misleading.

Worked Example 2 — Comparing Kurtosis Between Two Portfolios

Two investment portfolios report the same annualized mean return and standard deviation, but different distributions of monthly returns.

1

Portfolio A: Monthly returns are tightly clustered around the mean. No single month produced a return more than 2 standard deviations from average. Excess kurtosis = −0.8 (platykurtic).

2

Portfolio B: Most months produce modest returns close to the mean, but three months in the past five years produced losses greater than 4 standard deviations below the mean. Excess kurtosis = +8.2 (strongly leptokurtic).

3

Both portfolios share the same mean and standard deviation. Looking only at those two numbers, the portfolios appear equally risky.

4

Interpret kurtosis: Portfolio B (g₂ = 8.2) falls in the "strong leptokurtosis" range. Extreme losses are not black swan events — they are a predictable feature of its return distribution.

✓ Kurtosis reveals what standard deviation hides. A risk manager evaluating only mean and SD would treat these portfolios as equivalent. The kurtosis difference signals that Portfolio B requires different risk management — options-based hedging or more conservative position sizing.

How to Report Skewness and Kurtosis (APA Format)

Academic and clinical research routinely requires reporting skewness and kurtosis alongside descriptive statistics. APA style requires both the values and their standard errors — which are derived from the sample size.

Standard Error of Skewness and Kurtosis (SES / SEK)

APA Reporting Formulas

Standard Error of Skewness (SES):

SES = √(6/n)

Standard Error of Kurtosis (SEK):

SEK = √(24/n)

Normality z-score test (skewness):

z = skewness / SES  →  if |z| > 2, distribution is significantly skewed at α = .05

Example APA table format (n = 100):

VariableMSDSkewnessSEKurtosisSE
Exam score72.411.3−0.840.240.520.48
Response time (ms)423871.240.242.170.48

Note: SES = √(6/100) = 0.245; SEK = √(24/100) = 0.490. Values rounded to 2 decimal places per APA 7th edition.

In the body text of a results section you would write: "Descriptive statistics revealed that response time was positively skewed (skewness = 1.24, SE = 0.24), with the skewness z-score of 5.17 indicating significant departure from symmetry at p < .001. The distribution was leptokurtic (kurtosis = 2.17, SE = 0.48)."

Real-World Distribution Analysis

Case Study 1 — Salary Distribution

Case Study

Income Inequality: High Positive Skew, High Kurtosis

The distribution of individual annual incomes in most market economies shows two pronounced characteristics simultaneously: strong positive skewness (the right tail extends to multi-million-dollar incomes) and high excess kurtosis (extreme incomes, both very high and occasionally very low, are more common than a normal distribution predicts).

Median income is the standard measure of living standards precisely because positive skewness makes the mean a poor representative of the typical worker's experience. Meanwhile, the heavy upper tail — captured by kurtosis — is what drives Gini coefficient calculations and top-income-share statistics.

Sources: U.S. Census Bureau Current Population Survey (CPS); Piketty, T. & Saez, E. (2003). "Income Inequality in the United States, 1913–1998." Quarterly Journal of Economics, 118(1).

Case Study 2 — Stock Market Returns

Case Study

Financial Risk: Near-Zero Skew, Very High Kurtosis

Daily stock index returns are roughly symmetric in direction (skewness close to zero), but strongly leptokurtic — both very large gains and very large losses happen far more often than a normal distribution would predict. This is the "fat tails" property documented extensively in financial econometrics since the 1960s.

The practical implication is that Value-at-Risk (VaR) models built on normal distribution assumptions underestimate extreme losses. Research by Campbell, Lo, and MacKinlay documented that daily S&P 500 returns have an excess kurtosis of approximately 7–12, depending on the time period — meaning a normal model underestimates the probability of a single-day loss greater than 4 standard deviations by roughly 100×.

Sources: Fama, E.F. (1965). "The Behavior of Stock-Market Prices." Journal of Business, 38(1); Campbell, J.Y., Lo, A.W., & MacKinlay, A.C. (1997). The Econometrics of Financial Markets. Princeton University Press.

Case Study 3 — Exam Score Distribution

Case Study

Exam Scores: Varying Skew, Near-Normal Kurtosis

Exam score distributions change shape depending on test difficulty. On a well-calibrated exam, scores approximate a normal distribution (zero skewness, zero excess kurtosis). On a very easy exam, scores cluster near the top, producing negative skewness. On a very difficult exam, scores cluster near the bottom, producing positive skewness.

This matters for grading decisions. When a class's exam scores are negatively skewed, applying a strict normal curve to assign grades penalizes students unfairly — the distribution does not fit the assumption. Using the mean as the central benchmark fails in skewed score distributions.

Related reading: Lord, F.M. & Novick, M.R. (1968). Statistical Theories of Mental Test Scores. Addison-Wesley.

Skewness and Kurtosis Calculator

Enter your comma-separated data below. The calculator computes sample skewness (Fisher–Pearson), excess kurtosis, and their standard errors (SES / SEK) for any dataset with three or more values. Results also show the APA normality z-scores.

Skewness & Kurtosis Calculator

—
Sample Skewness (g₁)
—
Excess Kurtosis (g₂)
—
SES = √(6/n)
—
SEK = √(24/n)

How to Calculate Skewness and Kurtosis

Manual calculation using the definitions above is tedious for large datasets but straightforward for small ones. The steps below use a five-value dataset to show each moment calculation explicitly.

Hand Calculation — Dataset: 2, 4, 6, 8, 20

Calculate sample skewness and excess kurtosis for the values: 2, 4, 6, 8, 20

1

Mean (x̄): (2 + 4 + 6 + 8 + 20) / 5 = 40 / 5 = 8.0

2

Standard deviation (s): Deviations from mean: −6, −4, −2, 0, 12. Squared deviations: 36, 16, 4, 0, 144. Sum = 200. Sample variance = 200/(5−1) = 50. s = √50 ≈ 7.071

3

Standardized values [(xᵢ − x̄)/s]: −0.849, −0.566, −0.283, 0.000, 1.697

4

Cubed standardized values for skewness: −0.611, −0.181, −0.023, 0.000, 4.876. Sum = 4.061

5

Sample skewness: g₁ = [n/((n−1)(n−2))] × Σz³ = [5/(4×3)] × 4.061 = 0.4167 × 4.061 ≈ 1.69

✓ Skewness g₁ ≈ 1.69 — substantial positive skew, driven by the outlier value of 20. Excess kurtosis g₂ ≈ 2.41 (leptokurtic), indicating heavier tails than normal. SES = √(6/5) ≈ 1.10; skewness z-score = 1.69 / 1.10 ≈ 1.54 (not significant at n=5 — too small a sample to draw firm conclusions).

Alternative Skewness Formulas (Galton, Pearson 2)

The Fisher–Pearson formula is the most common, but two alternative measures are worth knowing for situations where the standard formula is sensitive to outliers:

Formula Definition When to Use
Fisher–Pearson (g₁)[n/((n−1)(n−2))] × Σ[(xᵢ−x̄)/s]³Default in most software; general use
Galton skewness (Bowley's)(Q3 + Q1 − 2×Q2) / (Q3 − Q1)Resistant to outliers; uses quartiles Q1, Q2 (median), Q3
Pearson 2 coefficientSk₂ = 3(mean − median) / sQuick approximation; useful when only M, Mdn, and SD are reported

Galton (Bowley's) skewness is particularly useful when your data has extreme outliers that dominate the standard formula. Since it is based on quartiles rather than all data points, a single extreme value cannot distort the result. The NIST Handbook covers this measure in depth for exploratory analysis of robust shape assessment.

Jarque-Bera Test: Combining Skewness and Kurtosis

The Jarque-Bera (JB) test is the most common formal statistical test that uses both skewness and kurtosis simultaneously to test whether a sample comes from a normal distribution. It is widely used in economics and finance as a diagnostic before model fitting.

Jarque-Bera Test Statistic
JB = (n/6) × (S² + K²/4)
Follows a chi-squared distribution with 2 degrees of freedom under the null hypothesis of normality
n = sample size S = sample skewness (g₁) K = excess kurtosis (g₂) JB > 5.99 → reject normality at α = .05

A large JB value (or small p-value against the chi-squared distribution with 2 df) means the data departs significantly from normality due to its combined skewness and/or kurtosis. If either skewness or kurtosis is large, the test rejects. In Python: scipy.stats.jarque_bera(data); in R: tseries::jarque.bera.test(data).

Skewness and Kurtosis in Machine Learning / EDA

In exploratory data analysis (EDA) and machine learning workflows, checking skewness and kurtosis is a standard pre-processing step. Here is the typical decision logic:

EDA Workflow — Shape Diagnostics Before Modeling

Checking skewness and kurtosis before fitting a model

1

Compute shape statistics: Calculate g₁ (skewness) and g₂ (excess kurtosis) for every numeric feature. Flag any with |g₁| > 1 or |g₂| > 2.

2

Apply transformation if needed: For right-skewed features (g₁ > 1): try log(x+1) or square root. For left-skewed features: try reflecting and then log. Box-Cox transformation automatically finds the optimal power parameter λ to minimize skewness.

3

Re-check after transformation: Recompute skewness and kurtosis. Most good transformations reduce |g₁| below 1 and bring |g₂| below 3.

4

Algorithm sensitivity: Linear models (OLS, logistic regression), LDA, and distance-based algorithms (k-NN, k-means) are sensitive to skewness and kurtosis. Tree-based methods (Random Forest, XGBoost) are generally robust to both — transformation is less critical for them.

✓ Skewness and kurtosis are part of the standard EDA checklist alongside missing values, correlations, and outlier detection. Libraries: df.skew() and df.kurtosis() in pandas give column-wise shape statistics in one call.

Skewness and Kurtosis in Software

# Python — using scipy.stats and pandas from scipy import stats import pandas as pd import numpy as np data = [2, 4, 6, 8, 20] n = len(data) skewness = stats.skew(data, bias=False) # Fisher-Pearson (unbiased) kurtosis = stats.kurtosis(data, bias=False) # excess kurtosis by default ses = np.sqrt(6 / n) # standard error of skewness sek = np.sqrt(24 / n) # standard error of kurtosis jb_stat, jb_p = stats.jarque_bera(data) # Jarque-Bera test print(f"Skewness: {skewness:.4f} (SE = {ses:.4f}, z = {skewness/ses:.2f})") print(f"Kurtosis: {kurtosis:.4f} (SE = {sek:.4f})") print(f"Jarque-Bera: {jb_stat:.4f}, p = {jb_p:.4f}") # For entire dataframe — pandas gives column-wise stats df = pd.DataFrame({'scores': data}) print(df.skew()) # all columns print(df.kurtosis()) # all columns (excess kurtosis)
# R — e1071 for corrected estimators library(e1071) library(tseries) # for jarque.bera.test data <- c(2, 4, 6, 8, 20) n <- length(data) skewness(data, type = 2) # type=2 = Fisher-Pearson kurtosis(data, type = 2) # excess kurtosis sqrt(6/n) # SES sqrt(24/n) # SEK tseries::jarque.bera.test(data) # Jarque-Bera test # SPSS: Analyze → Descriptive Statistics → Descriptives # Check "Kurtosis" and "Skewness" — SPSS reports both values AND their Std. Errors
=SKEW(A1:A5) // Sample skewness (Fisher-Pearson) =KURT(A1:A5) // Excess kurtosis (Fisher's correction applied) =SQRT(6/COUNT(A1:A5)) // SES = standard error of skewness =SQRT(24/COUNT(A1:A5)) // SEK = standard error of kurtosis

Kurtosis of Common Probability Distributions

The table below shows the skewness and excess kurtosis for 12 named distributions. This makes it easier to identify which distribution might model your data based on observed shape statistics.

Distribution Skewness Excess Kurtosis
Normal distribution00
Uniform distribution0−1.2 (platykurtic)
Exponential distribution2.0 (right-skewed)6.0 (leptokurtic)
Student's t (5 df)0 (symmetric)6.0 (leptokurtic)
Student's t (10 df)0 (symmetric)2.0 (leptokurtic)
Student's t (30 df)0 (symmetric)0.43 (slightly leptokurtic)
Log-normal (σ=1)6.18 (strongly right-skewed)110.9 (extreme leptokurtosis)
Beta(2,5)0.60 (right-skewed)−0.26 (slightly platykurtic)
Beta(5,2)−0.60 (left-skewed)−0.26 (slightly platykurtic)
Weibull (k=1.5)0.64 (right-skewed)0.07 (approximately mesokurtic)
Cauchy distributionUndefinedUndefined (tails too heavy)
Bernoulli (p=0.5)0−2.0 (platykurtic)

The Cauchy distribution is notable: its tails are so heavy that neither the mean, variance, skewness, nor kurtosis are defined in the traditional sense. The log-normal distribution with σ=1 shows extreme values for both statistics — which is why log-normal is used to model income and stock prices (which exhibit exactly these properties in real data).

Reading Histograms for Skewness and Kurtosis

Before computing any formula, a good histogram usually reveals the shape of your data. Here is what to look for:

Histogram Pattern Skewness Kurtosis
Symmetric, moderate peak≈ 0≈ 0 (excess)
Longer right tail, peak shifted left> 0 (positive)Varies
Longer left tail, peak shifted right< 0 (negative)Varies
Very tall sharp peak, long thin tails≈ 0> 0 (leptokurtic)
Flat, wide, no pronounced peak≈ 0< 0 (platykurtic)
Two humps, roughly equal≈ 0< 0 (platykurtic or bimodal)
✅
Always Use Numbers and Visuals Together

A Q-Q (quantile-quantile) plot complements skewness and kurtosis numbers: points bowing above the line indicate right skewness; S-shaped curves indicate kurtosis departures. The numbers tell you how much; the plot tells you where. For formal normality testing, use the Shapiro–Wilk test or normality tests overview.

From Beginner to Advanced: A Progressive Learning Path

Level 1 — Beginner: What Shape Means in Data

At the beginner level, the most important takeaway is that the mean and standard deviation do not fully describe a distribution. When someone tells you the average salary at a company is $90,000, you cannot judge whether that number is representative without knowing whether the distribution is skewed. If five executives each earn $1,000,000 and fifty employees each earn $50,000, the mean is pulled to $90,909 — but it describes no one's actual salary well.

Level 2 — Intermediate: Interpreting Shape in Datasets

At the intermediate level, the focus shifts to diagnosis. When you load a dataset and prepare to build a model or run a hypothesis test, checking skewness and kurtosis is part of the EDA workflow. Many statistical tests — including the one-sample t-test and ANOVA — assume that residuals are normally distributed. Substantial skewness or excess kurtosis is a signal that this assumption may need testing or that a transformation might be appropriate.

Level 3 — Advanced: Kurtosis and Risk Modeling

At the advanced level, kurtosis becomes a central concern in any model involving tail risk. Value-at-Risk (VaR) and Expected Shortfall (ES) — the two standard measures of market risk — are calculated from the tails of a return distribution. A model that assumes normality will assign incorrect probabilities to extreme outcomes. Modern approaches use distributions that explicitly accommodate non-zero kurtosis, including the Student's t-distribution and extreme value theory (EVT) models. (BIS: Minimum capital requirements for market risk)

Entity and Concept Glossary

Concept Formula / Symbol Interpretation Real-World Meaning Common Mistake
Skewness g₁ = Σ[(xᵢ−x̄)/s]³ × n/((n−1)(n−2)) Direction of data asymmetry Income inequality, exam score shape Confusing it with spread (standard deviation)
Kurtosis (excess) β₂ − 3; normal = 0 Tail heaviness relative to normal Financial crash probability, outlier frequency Thinking it measures peak height (it does not)
Positive skew g₁ > 0 Right tail longer; mean > median Wealth, housing prices, waiting times Using the mean as "typical" in this case
Negative skew g₁ < 0 Left tail longer; mean < median Easy exam scores, lifespan in rich countries Ignoring the tail when reporting results
Leptokurtic β₂ − 3 > 0 Heavier tails; more extreme values Stock returns, financial risk models Conflating with "tall peak"
Platykurtic β₂ − 3 < 0 Lighter tails; fewer extremes Uniform distributions, precision manufacturing Assuming low kurtosis means low variance
Mesokurtic β₂ − 3 ≈ 0 Tail weight similar to normal Height, IQ scores at population scale Assuming mesokurtic = normally distributed
SES √(6/n) Standard error of skewness Used in APA reporting and normality z-scores Omitting from APA results tables
SEK √(24/n) Standard error of kurtosis Used in APA reporting and normality z-scores Omitting from APA results tables
Jarque-Bera test JB = (n/6)(S² + K²/4) ~ χ²(2) Tests normality using both skewness and kurtosis Pre-model diagnostic in econometrics and finance Using it on very small samples (n < 30)
Galton skewness (Q3 + Q1 − 2Q2) / (Q3 − Q1) Quartile-based skewness resistant to outliers Exploratory analysis when outliers are suspected Using it when tails (not outliers) are the concern
Fat tails High positive excess kurtosis Extreme values more probable than normal Market crashes, insurance losses Underestimating tail probability using normal models

Frequently Asked Questions

What is skewness in simple terms?
Skewness tells you which direction a dataset leans. Imagine a seesaw: a symmetric distribution sits balanced in the middle. A positively skewed distribution has more weight on the left (most values are low) but the right side extends further due to a small number of very high values. A negatively skewed distribution is the mirror image. The number itself — calculated from cubed deviations — captures both the direction and the degree of that lopsidedness.
What does kurtosis tell you about data?
Kurtosis tells you how often your data produces extreme values relative to what a normal distribution would predict. High kurtosis (positive excess kurtosis, leptokurtic) means outliers are more common than expected — the tails of the distribution are heavier than normal. Low kurtosis (platykurtic, negative excess kurtosis) means outliers are less common than expected. It is fundamentally a measure of tail behavior, not of the peak's sharpness.
What is the acceptable range of skewness and kurtosis?
For most statistical tests: skewness between −0.5 and +0.5 is considered approximately symmetric. Values between ±0.5 and ±1.0 indicate moderate skewness. Values beyond ±1.0 indicate substantial skewness. For excess kurtosis: values between −1 and +1 are close to normal; values beyond ±2 warrant attention. These thresholds vary by field — social science often accepts |skewness| < 2 and |excess kurtosis| < 7 for parametric tests, while finance is more conservative.
Can kurtosis be negative?
Yes. Negative excess kurtosis (platykurtic) means the distribution has lighter tails than the normal distribution — extreme values are less common than expected. The uniform distribution has excess kurtosis of approximately −1.2. Rolling a fair die produces a platykurtic result: each outcome is equally likely and "extreme" values are no more common than any other. Negative kurtosis indicates a flatter, more evenly spread distribution.
What is the difference between excess kurtosis and kurtosis?
Raw kurtosis (β₂) is the fourth standardized moment; the normal distribution has raw kurtosis of 3. Excess kurtosis subtracts 3 from raw kurtosis (β₂ − 3), making the normal distribution's reference value 0. Most software — including Excel's KURT(), Python's scipy.stats.kurtosis(), and SPSS — reports excess kurtosis. When researchers write "kurtosis = 0.5" in a paper, they almost certainly mean excess kurtosis. The word "excess" may be dropped in context.
What is the difference between skewness and kurtosis?
Skewness and kurtosis are independent properties. Skewness measures asymmetry — whether data leans left or right — and is derived from the third statistical moment (symbol: γ₁ / g₁). Kurtosis measures tail heaviness — how likely extreme values are — and is derived from the fourth moment (symbol: β₂ / g₂). A distribution can be perfectly symmetric (zero skewness) while having very heavy tails (high kurtosis), as with the symmetric Student's t-distribution. They answer different questions and should be interpreted separately.
What is the Jarque-Bera test?
The Jarque-Bera test is a goodness-of-fit test that uses both skewness and kurtosis to test whether sample data comes from a normal distribution. The test statistic is JB = (n/6) × (S² + K²/4), where S is sample skewness and K is excess kurtosis. Under the null hypothesis of normality, JB follows a chi-squared distribution with 2 degrees of freedom. A p-value below 0.05 indicates the data departs significantly from normality. It is widely used in economics and finance as a pre-modeling diagnostic. Note: it has low power for small samples (n < 30).
How do you report skewness and kurtosis in APA format?
In APA format, report the mean (M), standard deviation (SD), skewness with its standard error SE (SES = √(6/n)), and kurtosis with its standard error (SEK = √(24/n)). Example for n = 100: "M = 52.3, SD = 8.4, skewness = 1.24 (SE = 0.24), kurtosis = 2.17 (SE = 0.49)." To test whether skewness is significantly different from zero, divide skewness by SES: if |z| > 1.96, it is significant at α = .05. Some journals also require a Shapiro–Wilk test result alongside these values.
What is positive kurtosis?
Positive kurtosis (positive excess kurtosis) means the distribution is leptokurtic — it has heavier tails than the normal distribution. Values more than 3 standard deviations from the mean occur more frequently than they would in a normal bell curve. The distribution may appear sharper-peaked, but the defining characteristic is tail heaviness, not peak height. Stock returns, insurance claims, and income distributions often exhibit positive kurtosis.
What does high kurtosis mean?
High kurtosis (typically taken as excess kurtosis above 1 in data analysis, or above 3 in finance) means outliers and extreme values are more common than a normal model would predict. The tails of the distribution are "fat." In practice: in regression, high kurtosis in residuals may indicate influential outliers. In finance, it means normal-distribution-based risk models underestimate the probability of large losses. In quality control, it signals that defects or process deviations occur more frequently than expected.
How does skewness affect mean and median?
In a positively skewed distribution, the mean is pulled upward by extreme high values, so mean > median > mode. The few very high values in the right tail inflate the mean. In a negatively skewed distribution, extreme low values drag the mean down, so mean < median < mode. The median is resistant to skewness because it is the middle ranked value and cannot be pulled by a few extremes. This is why median income is always reported alongside (or instead of) mean income for skewed distributions.
What is leptokurtic distribution?
A leptokurtic distribution has positive excess kurtosis (g₂ > 0), meaning its tails are heavier than the normal distribution — extreme values occur more often than a bell curve predicts. Stock market daily returns are the classic example: large gains or losses that a normal model would assign a probability of once in several thousand years actually occur every few years. Other leptokurtic real-world examples include insurance claim sizes, earthquake magnitudes, and internet traffic volumes.
📚 Academic Sources

Key References Used in This Guide

NIST/SEMATECH e-Handbook of Statistical Methods: Section on Skewness and Kurtosis — the primary technical reference for both measures.

Westfall, P.H. (2014): "Kurtosis as Peakedness, 1905–2014. R.I.P." The American Statistician, 68(3), 191–195. Documents the peer-reviewed evidence that kurtosis is a tail measure, not a peak measure.

DeCarlo, L.T. (1997): "On the Meaning and Use of Kurtosis." Psychological Methods, 2(3), 292–307. (APA PsycNet)

Groeneveld, R.A. & Meeden, G. (1984): "Measuring Skewness and Kurtosis." The Statistician, 33(4), 391–399. Covers alternative definitions and their properties.