What Is Heteroscedasticity?
The name comes from the Greek roots hetero (different) and skedasis (dispersion). When dispersion is the same across all observations, the condition is called homoscedasticity. When it differs, it is heteroscedasticity. Both spellings — heteroscedasticity and heteroskedasticity — refer to exactly the same concept; different textbooks and disciplines simply favour one over the other.
The clearest mental image: plot the regression residuals on the vertical axis and fitted values on the horizontal axis. Under constant variance, residuals form a band of roughly equal width from left to right. Under heteroscedasticity, that band fans out, narrows, or shifts in some systematic way. The pattern is a diagnostic clue, not an automatic verdict.
Heteroscedasticity occurs in a regression model when the variance of the error term is not constant across observations. Instead of a uniform band, residuals fan outward or narrow as fitted values change. It does not automatically bias OLS coefficient estimates but can make conventional standard errors, p-values, and confidence intervals unreliable.
The Constant-Variance Assumption
In a standard simple linear regression model:
Yᵢ = response variable for observation i
β₀ = intercept
β₁ = slope
Xᵢ = predictor variable
εᵢ = error term (unobserved)
The classical ordinary least squares framework assumes, among other things, that the error variance is the same for every observation:
Heteroscedasticity replaces that fixed σ² with a value that can differ across observations:
The subscript on σᵢ² signals that the variance is now observation-specific. It may grow with X, shrink with X, follow a nonlinear variance relationship, or change in some other structured way. Residuals — calculated as eᵢ = yᵢ − ŷᵢ — are the observable proxies for these unobserved errors, so analysts examine residuals to diagnose the problem. They are not identical to the true errors, but they carry the signal needed for practical diagnostics. For a full treatment of what regression residuals are and how to interpret them, see the regression residuals guide.
Homoscedasticity vs Heteroscedasticity
The two conditions sit at opposite ends of the variance-assumption spectrum. Both are properties of the conditional error distribution — how spread out errors are given a particular predictor value — not simply how spread out the raw Y values are.
| Feature | Homoscedasticity | Heteroscedasticity |
|---|---|---|
| Conditional error variance | Constant across observations (σ²) | Changes across observations (σᵢ²) |
| Residual-vs-fitted plot | Roughly equal vertical spread throughout | Fan, funnel, bow-tie, or shifting spread |
| OLS coefficients | Unbiased under standard conditions | Not necessarily biased by this alone |
| Conventional OLS standard errors | Consistent under classical assumptions | May be inconsistent (too large or too small) |
| Gauss-Markov efficiency | OLS has minimum variance in linear unbiased class | OLS may lose that efficiency property |
| Typical cause | Well-specified model with stable variance | Scale effects, omitted variables, wrong functional form |
Heteroscedasticity does not automatically make OLS coefficient estimates biased. Under an otherwise correctly specified linear conditional mean with appropriate exogeneity, the OLS slope and intercept can remain unbiased even when variance is not constant. The primary inferential damage falls on the standard errors, not the point estimates themselves.
What Does Heteroscedasticity Look Like?
The residual-versus-fitted-values plot is the primary visual diagnostic. The question to ask is simple: does the vertical spread of residuals remain roughly stable from left to right, or does it change in some systematic way?
Residuals scatter with roughly equal width across all fitted values. The band stays parallel to zero.
Residuals fan outward as fitted values increase. The variance is clearly larger on the right side of the plot.
Four Common Variance Patterns
Heteroscedasticity is not only an increasing fan. The defining feature is a systematic change in spread, regardless of the direction. Here are the patterns analysts most often encounter.
Constant Width
Variance is stable. Homoscedasticity is plausible.
Widening Fan
Variance grows with fitted values. The classic heteroscedastic shape.
Narrowing Funnel
Variance decreases with fitted values. Also heteroscedastic — just the reverse direction.
Bow-Tie
Variance is large at extremes and small in the middle — a more complex pattern.
One extreme residual does not prove heteroscedasticity. It may represent an outlier, a data error, or an influential observation with high leverage. Heteroscedasticity refers to a systematic change in variance across observations, not merely the presence of one large residual. Use the influential points guide to separate outlier effects from variance patterns.
Heteroscedasticity vs Nonlinearity
A curved band of residuals — where the center of the residual cloud bows above or below zero — points toward a misspecified conditional mean, not changing variance. Changing residual width is the variance signal. Changing residual center is the mean-specification signal. Both can appear simultaneously, which is why careful visual inspection of the full residual plot matters before any remedy is chosen.
Why Does Heteroscedasticity Occur?
There is no single cause. Understanding what produced the changing variance in a particular dataset is the first step toward choosing a sensible response. Common origins include the following.
| Cause | Description | Example setting |
|---|---|---|
| Scale effects | Absolute variability naturally grows with the level of the outcome | Household spending: high-income households vary by thousands more than low-income ones |
| Omitted variables | Unmodeled groups have different error dispersion; their absence leaks into the residuals | Wages model missing occupation subgroups with very different pay dispersion |
| Wrong functional form | A linear mean function poorly approximates a nonlinear relationship, creating structured residual patterns | Fitting a line to an exponential growth curve |
| Skewed positive outcomes | Outcomes bounded at zero with right-skewed distributions often have variance that grows with their mean | Insurance claims, house prices, production volumes |
| Measurement process | Measurement precision differs across the range of X | Self-reported survey data, laboratory instruments with range-dependent accuracy |
| Mixed populations | Different subgroups with different variability are pooled without group indicators | Pooling small businesses and large corporations in a single revenue regression |
| Structural changes | Variance shifts at a threshold or over time | Financial returns before and after a market shock |
Identifying the likely cause matters because it shapes the remedy. A scale effect suggests a response transformation. Omitted variables suggest model respecification. Unknown variance structure suggests robust standard errors. Do not reach for a fix before understanding the source.
Why Is Heteroscedasticity a Problem?
Effect on OLS Coefficients
This is the most widely misunderstood consequence. Under a correctly specified linear conditional mean model with the appropriate error-predictor independence conditions, heteroscedasticity alone does not make OLS coefficient estimates biased. The OLS regression line can still estimate the conditional mean appropriately. The slope β̂₁ and intercept β̂₀ are determined by the first-order conditions that minimize squared residuals, and those conditions do not require constant variance to produce an unbiased result.
Effect on Standard Errors
This is where heteroscedasticity creates real inferential damage. The conventional OLS standard-error formula was derived under the assumption of constant variance. When error variance changes across observations, that formula uses the wrong variance information, and the resulting standard errors can be inconsistent — meaning they do not converge to the right value even in large samples.
Effect on P-Values and Confidence Intervals
Because t-statistics divide coefficient estimates by their standard errors, distorted standard errors flow through to distorted t-statistics, distorted p-values, and confidence intervals with incorrect nominal coverage. A slope that appears statistically significant at the 5% level under conventional standard errors may not be significant under a correct variance estimate — or vice versa. This is why the inferential chain matters:
OLS coefficient estimate → standard error → t-statistic → p-value / confidence interval. Heteroscedasticity disrupts the standard error step. The coefficient may be unchanged while every subsequent inferential quantity shifts.
Effect on OLS Efficiency
Under the Gauss-Markov conditions — which include the constant-variance assumption — OLS has the familiar property of being the best (minimum variance) linear unbiased estimator. When heteroscedasticity is present, OLS no longer necessarily holds that efficiency property. Alternative estimators such as weighted least squares, properly specified, can be more efficient. Whether this matters practically depends on how severe the variance imbalance is and what precision is required for the analysis.
Effect on Prediction Uncertainty
When conditional error variance changes across X, prediction intervals of constant width are inadequate. A single RMSE figure summarises average prediction-error magnitude across all observations, but it cannot capture that predictions at some fitted values are inherently more uncertain than predictions at others. Understanding the variance structure allows better-calibrated prediction intervals. For more on how prediction accuracy is measured overall, see the RMSE guide, bearing in mind that a single RMSE does not diagnose whether variance is constant.
How to Detect Heteroscedasticity
Step-by-Step: Residual-vs-Fitted Plot
Fit the regression model
Run the OLS regression and obtain the fitted (predicted) values ŷᵢ for each observation.
Calculate residuals
Compute eᵢ = yᵢ − ŷᵢ for every observation. These are the observable estimates of the unobserved true errors.
Plot residuals against fitted values
Place ŷᵢ on the horizontal axis and eᵢ on the vertical axis. Add a horizontal reference line at zero.
Inspect the spread
Ask: does the vertical band of residuals remain roughly equal in width from left to right, or does it widen, narrow, or shift? Systematic changes in spread are the visual signal of heteroscedasticity.
Look for outliers and subgroups
A few isolated extreme points may be outliers or influential observations rather than evidence of changing variance. Also look for distinct clusters or bands that might signal omitted group structure.
Consider formal tests if useful
Visual inspection can reveal shape and direction; formal tests provide a p-value under stated assumptions. Use both — and the Residual Plot Generator to examine your own data interactively.
The interactive Residual Plot Generator on StatisticsFundamentals.com lets you paste your own actual and fitted values, then examine the residual pattern directly. Visual diagnosis of changing spread is most informative when done on real data rather than textbook illustrations.
Breusch-Pagan Test
The Breusch-Pagan test is a formal procedure for evaluating whether residual variance is systematically related to the explanatory variables. Rather than relying on visual judgment alone, it produces a test statistic and p-value under clearly stated assumptions.
Null hypothesis: Var(εᵢ | Xᵢ) = σ² — error variance is constant.
Alternative: Error variance depends systematically on the specified explanatory variables.
The core logic: if the squared residuals (which proxy for error variance at each observation) are predictable from the explanatory variables, the null of constant variance is undermined. At a high level, the test auxiliary regression relates squared residuals to the predictors and evaluates how much explained variation exists. The resulting LM test statistic follows an approximate chi-squared distribution under the null, with degrees of freedom corresponding to the number of predictors in the auxiliary regression.
Insufficient evidence against constant variance under the test specification. The constant-variance assumption is not contradicted — but this does not prove homoscedasticity definitively.
Evidence that residual variance is systematically related to the specified predictors. Investigate the residual plot and consider whether inference should use robust standard errors or another remedy.
A non-significant result does not prove homoscedasticity — the test may lack power, especially in small samples. A significant result does not mean the regression is unusable — it means conventional standard errors warrant scrutiny. Residual plots and formal tests work best together.
White Test
The White test takes a more general approach. Rather than regressing squared residuals only on the original predictors, it also includes squares and cross-products of those predictors, allowing it to detect forms of heteroscedasticity that do not follow a simple linear pattern with the original X variables.
This generality is both an advantage and a limitation. The White test can detect more complex variance relationships, but in models with several predictors the auxiliary regression accumulates many terms, which consumes degrees of freedom and can reduce statistical power in smaller samples. The chi-squared test statistic follows the same distributional reference as Breusch-Pagan under the null, but with a larger degrees-of-freedom parameter.
| Feature | Breusch-Pagan Test | White Test |
|---|---|---|
| Main purpose | Detect variance linearly related to specified predictors | Detect more general heteroscedasticity patterns |
| Auxiliary regression terms | Original predictors | Predictors, their squares, and cross-products |
| Degrees of freedom | Lower (number of predictors) | Higher (grows quickly with model size) |
| Power in small samples | Generally better with simple variance patterns | May be reduced due to many terms |
| Interpretability | Points toward which predictors drive variance | Less directional — flags general heteroscedasticity |
| Best used when | You have a plausible theory about which variable drives variance | You want a general-purpose check with no specific variance hypothesis |
In practice, plotting residuals first gives direction and shape; then a formal test adds inferential grounding. The choice between Breusch-Pagan and White depends on the model size and whether there is prior theory about the variance structure.
How to Fix Heteroscedasticity
No single remedy applies universally. The right choice depends on what caused the changing variance, what the goal of the analysis is, and how well the variance structure can be modelled. Avoid choosing a remedy solely because it makes a diagnostic-test p-value non-significant.
Robust Standard Errors
Adjust the covariance matrix used for inference without changing OLS coefficients. Best when the conditional mean model is correct but variance is non-constant.
Transform the Response
A log or other transformation may stabilise variance when variability grows proportionally with the outcome. Changes interpretation — must be justified.
Weighted Least Squares
Assign lower weights to higher-variance observations. Requires a defensible model of the variance structure — wrong weights can worsen things.
Re-specify the Model
Add omitted predictors, use a more appropriate functional form, or include interaction terms. Variance problems sometimes reflect mean misspecification.
Alternative Statistical Model
For counts, proportions, or positive skewed outcomes, a generalized linear model may be more natural than forcing homoscedastic Gaussian OLS.
Heteroscedasticity-Robust Standard Errors
Robust standard errors — sometimes called heteroscedasticity-consistent (HC) standard errors or sandwich standard errors — recompute the covariance matrix of the OLS estimator using a formula that does not assume constant variance. Crucially, they do not change the OLS coefficient estimates. The same β̂₀ and β̂₁ remain; only the estimated covariance matrix, and therefore the standard errors, t-statistics, and confidence intervals, changes.
OLS coefficients: β̂ (unchanged — same OLS point estimates) Covariance matrix: V̂ (changed — uses HC-corrected formula) Standard errors: SE(β̂) = sqrt(diag(V̂)) (changed) t-statistics: β̂ / SE(β̂) (changed) p-values: from corrected t (changed) Confidence intervals: β̂ ± t* · SE(β̂) (changed)
HC0, HC1, HC2, HC3 — Which Variant?
Several versions of the heteroscedasticity-consistent covariance estimator exist, each applying different small-sample corrections.
| Estimator | Adjustment | When often preferred |
|---|---|---|
| HC0 | Original sandwich form; no leverage correction | Large samples where leverage is balanced |
| HC1 | Scales HC0 by n/(n−k) for degrees of freedom | Slight improvement over HC0 in moderate samples |
| HC2 | Adjusts each squared residual by (1 − hᵢᵢ), where hᵢᵢ is the leverage | Better finite-sample properties when leverage varies |
| HC3 | Stronger leverage adjustment; more conservative | Small samples or when high-leverage points are present; often the default recommendation |
HC3 is often the default recommendation in applied econometrics for small to moderate samples because its stronger leverage correction tends to produce better-calibrated inference. No single variant is universally optimal, and software packages use different defaults — always check which is being applied.
Response Transformation
A logarithmic transformation of the response variable may stabilise variance when variability is roughly proportional to the mean — a pattern common in positive, right-skewed outcomes. After taking logs, the model becomes multiplicative on the original scale, which changes both the variance structure and the interpretation of every coefficient.
Weighted Least Squares
Weighted least squares assigns each observation a weight inversely related to its estimated error variance. Observations with larger error variance receive lower weight, reducing their influence on the estimated coefficients. When the variance structure is correctly modelled, WLS can be more efficient than OLS under heteroscedasticity.
The challenge is that the variance structure Var(εᵢ | Xᵢ) is rarely known precisely. Incorrect weights introduce their own problems. Feasible generalised least squares (FGLS) estimates the variance parameters first, then uses those estimates as weights. If the variance model is misspecified, the resulting estimator may be worse than OLS with robust standard errors. WLS requires substantive justification — do not weight observations simply to make residual plots look flat.
| Remedy | Changes β̂? | Changes SE? | Requires variance model? | Best suited for |
|---|---|---|---|---|
| Robust standard errors (HC) | No | Yes | No | Correct inference when mean model is right |
| Response transformation | Yes | Yes | No exact model needed | Scale effects; positive, skewed outcomes |
| Weighted least squares | Yes | Yes | Yes — defensible weights needed | Known or well-estimated variance structure |
| Model respecification | Yes | Yes | No | Omitted variables or wrong functional form |
| Alternative GLM | Yes | Yes | Distributional model needed | Counts, binary, or other non-Gaussian outcomes |
Heteroscedasticity vs Other Assumption Violations
Heteroscedasticity vs Non-Normality
They are separate assumptions. Heteroscedasticity concerns whether the conditional variance is constant across observations. Non-normality concerns whether the conditional error distribution follows a normal distribution — relevant for exact finite-sample inference in classical regression. A model can have constant variance with non-normal errors, or changing variance with approximately normal conditional errors, or both, or neither. A normality test (such as a Shapiro-Wilk test on residuals) does not diagnose heteroscedasticity, and a heteroscedasticity test does not diagnose non-normality.
Heteroscedasticity vs Autocorrelation
Autocorrelation concerns dependence among errors across observations — a pattern common in time series or spatial data where nearby observations share information. Heteroscedasticity concerns changing variance magnitude, not error dependence. The two can coexist. Standard heteroscedasticity-robust standard errors do not correct for serial correlation; cluster-robust or HAC (heteroscedasticity and autocorrelation consistent) standard errors are needed when both are present.
Heteroscedasticity vs Outliers
A single extreme residual can locally distort the apparent spread in a residual plot, mimicking a heteroscedastic pattern. Before concluding that variance is systematically non-constant, identify whether any specific observations have unusually high leverage or influence on the regression. The influential points guide covers leverage, Cook's distance, and how to assess whether individual observations are driving apparent diagnostic patterns.
A Numeric Illustration
The table below uses two small hypothetical datasets to show the contrast between constant and increasing variance. The regression relationship (a slope of approximately 2) is similar in both; only the residual spread differs.
Dataset A — Roughly Constant Variance
| Obs (i) | X | Y (actual) | Ŷ (fitted) | Residual eᵢ | eᵢ² |
|---|---|---|---|---|---|
| 1 | 1 | 3 | 2.8 | 0.2 | 0.04 |
| 2 | 2 | 5 | 4.8 | 0.2 | 0.04 |
| 3 | 3 | 7 | 6.8 | 0.2 | 0.04 |
| 4 | 4 | 9 | 8.8 | 0.2 | 0.04 |
| 5 | 5 | 11 | 10.8 | 0.2 | 0.04 |
| Residual spread stays near ±0.2 throughout — constant variance plausible | 0.20 | ||||
Dataset B — Increasing Variance
| Obs (i) | X | Y (actual) | Ŷ (fitted) | Residual eᵢ | eᵢ² |
|---|---|---|---|---|---|
| 1 | 1 | 3.2 | 2.8 | 0.4 | 0.16 |
| 2 | 2 | 3.8 | 4.8 | −1.0 | 1.00 |
| 3 | 3 | 9.5 | 6.8 | 2.7 | 7.29 |
| 4 | 4 | 5.5 | 8.8 | −3.3 | 10.89 |
| 5 | 5 | 16.0 | 10.8 | 5.2 | 27.04 |
| Residual spread grows from ±0.4 at X=1 to ±5 at X=5 — clear fan pattern | 46.38 | ||||
In both datasets the fitted slope is approximately 2. The OLS coefficient is not obviously wrong in Dataset B. What is wrong is that if you apply the constant-variance standard-error formula to Dataset B, you get a single variance estimate that blends the small errors at X=1 with the large errors at X=5, producing standard errors that do not accurately reflect precision at any particular X value.
- Definition: Var(εᵢ | Xᵢ) changes across observations instead of remaining at a fixed σ².
- Spellings: Heteroscedasticity and heteroskedasticity refer to the same concept — both are correct.
- OLS coefficients: Not necessarily biased by heteroscedasticity alone under a correct mean model.
- Standard errors: Conventional homoscedastic SEs can be too large or too small — bias direction is not fixed.
- Primary visual tool: Residual-vs-fitted-values plot — inspect the width of the residual band.
- Formal tests: Breusch-Pagan for variance linearly related to predictors; White for more general patterns.
- Robust SEs: Correct inference without changing OLS coefficients — a common first-response remedy.
- WLS: Can improve efficiency but requires a justified variance model — not automatic.
- Log transforms: May stabilise variance for positive outcomes with proportional variability — not universal.
- Remedy choice: Driven by the cause, not by which option produces a non-significant diagnostic test.
Frequently Asked Questions
In a regression model, the errors attached to each prediction are not all equally spread out. For some observations, the model might be very close to the truth; for others, it might swing widely. When that unequal spread follows a pattern — usually growing or shrinking as the predictor changes — the model is said to have heteroscedasticity. The word is a technical way of saying "the prediction errors are not uniformly dispersed."
No. Both spellings describe the same statistical condition. Econometrics textbooks, particularly those with a British or European tradition, more often use heteroskedasticity (from the Greek skedasis). Statistics and data-science literature often uses heteroscedasticity (from the Latin root). If you see both in the same field, they still refer to identical assumptions about the regression error variance.
Not by itself. Under a correctly specified linear conditional mean with the standard error-predictor independence assumption, OLS coefficient estimates remain unbiased even when error variance is non-constant. The damage from heteroscedasticity falls primarily on the estimated standard errors, which then distort t-statistics, p-values, and confidence intervals. The regression line itself can still estimate the mean relationship accurately.
A funnel or fan shape in a residual-vs-fitted-values plot suggests that residual spread is not constant — it either widens or narrows as fitted values increase. This is the classic visual signature of heteroscedasticity. That said, one unusual cluster or outlier can create a superficially funnel-like appearance; confirm that the pattern is systematic rather than driven by a handful of observations before concluding that variance is truly non-constant.
No. R² measures the proportion of variance in the outcome explained by the predictors; it says nothing about whether error variance is constant across observations. A model can explain 95% of the variation in Y and still have a pronounced fan shape in its residuals. For more on what R² does and does not tell you, see the R-squared guide. Residual plots are the correct tool for assessing the variance assumption, not R².
No — robust standard errors do not remove or correct the heteroscedasticity itself. They adjust the way standard errors are estimated so that inference is more reliable in the presence of non-constant variance. The underlying data still has heteroscedastic errors; the robust approach simply avoids relying on the false assumption of constant variance when computing confidence intervals and hypothesis tests.
The Breusch-Pagan test is a formal statistical procedure that evaluates whether the squared residuals from a regression are systematically predicted by the explanatory variables. Its null hypothesis is constant error variance (homoscedasticity). A small p-value provides evidence that variance changes with the predictors. The test has a chi-squared null distribution with degrees of freedom equal to the number of predictors in the auxiliary regression. It works best when the variance pattern is expected to be linearly related to the original predictors.
Yes. Heteroscedasticity is not limited to simple linear regression with one predictor. In multiple linear regression, error variance can change with any combination of predictors, fitted values, or omitted variables. The same visual and formal diagnostic tools apply: plot residuals against fitted values, inspect the spread pattern, and use formal tests if needed. The remedies — robust standard errors, transformation, WLS, or respecification — are also the same in principle.
It depends on the analysis goal. If the objective is prediction and RMSE averaged across the sample is adequate, heteroscedasticity may have limited practical impact. If the goal is inference — testing whether a slope is statistically significant, constructing confidence intervals, or making policy recommendations — then ignoring non-constant variance means relying on standard errors that may be wrong in either direction. In most inferential applications it is worth at minimum using heteroscedasticity-robust standard errors when evidence of non-constant variance exists.
Quick Reference
| Item | Detail |
|---|---|
| Full name | Heteroscedasticity (also: heteroskedasticity — same concept) |
| Formal condition | Var(εᵢ | Xᵢ) = σᵢ² varies across observations |
| Opposite condition | Homoscedasticity: Var(εᵢ | Xᵢ) = σ² constant |
| Primary visual tool | Residual-vs-fitted-values plot — look for changing band width |
| Classic plot pattern | Fan shape (widening or narrowing residual spread) |
| Effect on OLS β̂ | Not necessarily biased under correct mean model conditions |
| Effect on conventional SEs | Can be too large or too small — direction depends on variance structure |
| Effect on p-values/CIs | Can be distorted — inference may be unreliable |
| Effect on efficiency | OLS may lose Gauss-Markov minimum-variance property |
| Formal test 1 | Breusch-Pagan — variance linearly related to predictors |
| Formal test 2 | White — more general variance patterns, uses squares/cross-products |
| Remedy 1 | Robust standard errors (HC0–HC3) — corrects inference, not coefficients |
| Remedy 2 | Response transformation (e.g., log) — for scale-driven variance growth |
| Remedy 3 | Weighted least squares — requires justified variance structure |
| Remedy 4 | Model respecification — for omitted variables or wrong functional form |
| Different from | Non-normality, autocorrelation, and outliers (related diagnostics, distinct concepts) |
| Relevant regression assumptions | See the full statistical assumptions overview |