What Are The Assumptions Of Linear Regression

7 min read

Linear regression stands as one of the most fundamental and widely used statistical techniques in data science, economics, and social sciences. Its popularity stems from its interpretability and the relative simplicity of its mathematical foundation. That said, the validity of the inferences drawn from a linear regression model—specifically the reliability of coefficient estimates, confidence intervals, and p-values—rests entirely on a set of core assumptions. Violating these assumptions can lead to biased estimates, inefficient predictions, and misleading significance tests, rendering the model useless for decision-making. Understanding these conditions is not merely an academic exercise; it is a practical necessity for anyone building predictive models or testing hypotheses.

The Gauss-Markov Assumptions: The Theoretical Bedrock

The classical linear regression model (CLRM) relies on the Gauss-Markov theorem, which states that under specific conditions, the Ordinary Least Squares (OLS) estimator is the Best Linear Unbiased Estimator (BLUE). "Best" here refers to having the minimum variance among all linear unbiased estimators. These conditions form the primary assumptions that must be satisfied for OLS to possess these desirable properties.

1. Linearity in Parameters

The first and most fundamental assumption is that the relationship between the dependent variable ($Y$) and the independent variables ($X_1, X_2, ..., X_k$) is linear in the parameters (coefficients $\beta$). This does not strictly require the variables themselves to have a straight-line relationship. You can model curves by including polynomial terms (e.g., $X^2$), interaction terms ($X_1 \times X_2$), or logarithmic transformations ($\ln(X)$), provided the equation remains linear in the betas.

Model Form: $Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_k X_k + \epsilon$

If the true relationship is inherently non-linear in parameters (e.g., $Y = \beta_0 + X^{\beta_1} + \epsilon$), standard OLS cannot estimate the parameters correctly without transformation. Diagnostic Tip: Plot residuals versus fitted values or residuals versus each predictor. A systematic pattern (like a U-shape) suggests a violation of linearity.

2. Random Sampling (Independence of Observations)

The data must represent a random sample drawn from the population. This implies that the observations are independent and identically distributed (i.i.d.). In practical terms, the value of the error term for one observation should not provide any information about the error term for another observation.

This assumption is frequently violated in time series data (autocorrelation), where today's error correlates with yesterday's error, or in panel/clustered data, where observations within the same group (e.Violation leads to underestimated standard errors, inflating t-statistics and producing false positives. g., students in the same school) share unobserved characteristics. Diagnostic Tip: Use the Durbin-Watson test for time series or examine the intraclass correlation coefficient (ICC) for clustered data That alone is useful..

3. Zero Conditional Mean (Exogeneity)

This is arguably the most critical assumption for causal inference. It states that the expected value of the error term ($\epsilon$), given any value of the independent variables ($X$), is zero: $E(\epsilon | X) = 0$.

In plain English, the unobserved factors captured in the error term must be uncorrelated with the included regressors. If an omitted variable correlates with both $Y$ and an included $X$, or if there is simultaneous causality (reverse causality), or measurement error in $X$, this assumption fails. The result is omitted variable bias (endogeneity), causing the OLS estimators to be biased and inconsistent—they will not converge to the true population parameters even with infinite data Worth keeping that in mind..

4. No Perfect Multicollinearity

The independent variables cannot have an exact linear relationship among themselves. Perfect multicollinearity occurs when one predictor is a perfect linear combination of others (e.g., including both "Height in inches" and "Height in centimeters," or a dummy variable trap where all categories are included without dropping a reference category) That's the part that actually makes a difference. That alone is useful..

While perfect multicollinearity makes estimation mathematically impossible (the $X'X$ matrix cannot be inverted), imperfect (high) multicollinearity is a more common practical nuisance. On the flip side, it inflates the variance of coefficient estimates, making them unstable and statistically insignificant (large standard errors), even if the predictors are jointly significant. Now, Diagnostic Tip: Calculate Variance Inflation Factors (VIF). A VIF above 5 or 10 typically signals problematic multicollinearity But it adds up..

5. Homoscedasticity (Constant Variance)

The variance of the error term must be constant across all levels of the independent variables: $Var(\epsilon | X) = \sigma^2$.

When this assumption is violated, we have heteroscedasticity. Still, this often happens in cross-sectional data where the spread of residuals increases with the magnitude of $X$ (e. Think about it: g. , predicting food expenditure based on income; high-income households have highly variable spending, low-income households have very consistent, low spending) Small thing, real impact..

While heteroscedasticity does not bias the coefficient estimates, it makes the standard OLS standard errors incorrect. Because of this, hypothesis tests (t-tests, F-tests) and confidence intervals become unreliable. Diagnostic Tip: Plot residuals vs. fitted values. Day to day, a "funnel" or "cone" shape indicates heteroscedasticity. On the flip side, formal tests include the Breusch-Pagan test and White test. Remedy: Use Heteroscedasticity-Consistent Standard Errors (Huber-White or "reliable" standard errors) Worth keeping that in mind..

The Normality Assumption: Inference and Small Samples

6. Normality of Error Terms

The Gauss-Markov assumptions (1-5) are sufficient for OLS to be BLUE. Even so, to perform exact hypothesis testing (t-tests, F-tests) and construct exact confidence intervals in small samples, we require the error terms to be normally distributed: $\epsilon \sim N(0, \sigma^2)$.

Thanks to the Central Limit Theorem (CLT), in large samples, the sampling distribution of the OLS estimators approximates normality regardless of the error distribution. So, normality is technically an assumption for finite-sample inference, not for the consistency or unbiasedness of the estimates themselves Simple as that..

Diagnostic Tip: Use a Q-Q plot (Quantile-Quantile plot) of the residuals. If points deviate significantly from the 45-degree line, normality is suspect. The Jarque-Bera test provides a formal statistical check. If normality fails in small samples, consider bootstrapping standard errors or transforming the dependent variable It's one of those things that adds up. Turns out it matters..

Additional Practical Considerations

Beyond the strict mathematical requirements, several "practical assumptions" dictate whether a regression model is useful and trustworthy in the real world.

Correct Model Specification

This encompasses linearity but goes further. It assumes you have included all relevant variables and excluded all irrelevant ones.

  • Omitted Variable Bias: Leaving out a relevant variable correlated with included $X
Dropping Now

Just Went Live

Explore More

Good Reads Nearby

Thank you for reading about What Are The Assumptions Of Linear Regression. We hope the information has been useful. Feel free to contact us if you have any questions. See you next time — don't forget to bookmark!
⌂ Back to Home