Errors and residuals

In

sample mean). The distinction is most important in regression analysis, where the concepts are sometimes called the regression errors and regression residuals and where they lead to the concept of studentized residuals

. In econometrics, "errors" are also called disturbances.^[1]^[2]^[3]

Introduction

Suppose there is a series of observations from a

location model

). In this case, the errors are the deviations of the observations from the population mean, while the residuals are the deviations of the observations from the sample mean.

A statistical error (or disturbance) is the amount by which an observation differs from its expected value, the latter being based on the whole population from which the statistical unit was chosen randomly. For example, if the mean height in a population of 21-year-old men is 1.75 meters, and one randomly chosen man is 1.80 meters tall, then the "error" is 0.05 meters; if the randomly chosen man is 1.70 meters tall, then the "error" is −0.05 meters. The expected value, being the mean of the entire population, is typically unobservable, and hence the statistical error cannot be observed either.

A residual (or fitting deviation), on the other hand, is an observable estimate of the unobservable statistical error. Consider the previous example with men's heights and suppose we have a random sample of n people. The

sample mean

could serve as a good estimator of the population mean. Then we have:

The difference between the height of each man in the sample and the unobservable population mean is a statistical error, whereas
The difference between the height of each man in the sample and the observable sample mean is a residual.

Note that, because of the definition of the sample mean, the sum of the residuals within a random sample is necessarily zero, and thus the residuals are necessarily not

independent. The statistical errors, on the other hand, are independent, and their sum within the random sample is almost surely

not zero.

One can standardize statistical errors (especially of a

studentized residuals

.

In univariate distributions

If we assume a normally distributed population with mean μ and standard deviation σ, and choose individuals independently, then we have

X_{1},\dots ,X_{n}\sim N\left(\mu ,\sigma ^{2}\right)\,

and the sample mean

{\overline {X}}={X_{1}+\cdots +X_{n} \over n}

is a random variable distributed such that:

{\overline {X}}\sim N\left(\mu ,{\frac {\sigma ^{2}}{n}}\right).

The statistical errors are then

e_{i}=X_{i}-\mu ,\,

with

expected values of zero,^[4]

whereas the residuals are

r_{i}=X_{i}-{\overline {X}}.

The sum of squares of the statistical errors, divided by σ², has a chi-squared distribution with n degrees of freedom:

{\frac {1}{\sigma ^{2}}}\sum _{i=1}^{n}e_{i}^{2}\sim \chi _{n}^{2}.

However, this quantity is not observable as the population mean is unknown. The sum of squares of the residuals, on the other hand, is observable. The quotient of that sum by σ² has a chi-squared distribution with only n − 1 degrees of freedom:

{\frac {1}{\sigma ^{2}}}\sum _{i=1}^{n}r_{i}^{2}\sim \chi _{n-1}^{2}.

This difference between n and n − 1 degrees of freedom results in

sample variance

of a population with unknown mean and unknown variance. No correction is necessary if the population mean is known.

Remark

It is remarkable that the

sum of squares of the residuals and the sample mean can be shown to be independent of each other, using, e.g. Basu's theorem

. That fact, and the normal and chi-squared distributions given above form the basis of calculations involving the t-statistic:

T={\frac {{\overline {X}}_{n}-\mu _{0}}{S_{n}/{\sqrt {n}}}},

where ${\overline {X}}_{n}-\mu _{0}$ represents the errors, $S_{n}$ represents the sample standard deviation for a sample of size n, and unknown σ, and the denominator term $S_{n}/{\sqrt {n}}$ accounts for the standard deviation of the errors according to:[5]

\operatorname {Var} \left({\overline {X}}_{n}\right)={\frac {\sigma ^{2}}{n}}

The probability distributions of the numerator and the denominator separately depend on the value of the unobservable population standard deviation σ, but σ appears in both the numerator and the denominator and cancels. That is fortunate because it means that even though we do not know σ, we know the probability distribution of this quotient: it has a Student's t-distribution with n − 1 degrees of freedom. We can therefore use this quotient to find a confidence interval for μ. This t-statistic can be interpreted as "the number of standard errors away from the regression line."^[6]

Regressions

In

homoscedasticity

.

However, a terminological difference arises in the expression mean squared error (MSE). The mean squared error of a regression is a number computed from the sum of squares of the computed residuals, and not of the unobservable errors. If that sum of squares is divided by n, the number of observations, the result is the mean of the squared residuals. Since this is a biased estimate of the variance of the unobserved errors, the bias is removed by dividing the sum of the squared residuals by df = n − p − 1, instead of n, where df is the number of degrees of freedom (n minus the number of parameters (excluding the intercept) p being estimated - 1). This forms an unbiased estimate of the variance of the unobserved errors, and is called the mean squared error.^[7]

Another method to calculate the mean square of error when analyzing the variance of linear regression using a technique like that used in

ANOVA (they are the same because ANOVA is a type of regression), the sum of squares of the residuals (aka sum of squares of the error) is divided by the degrees of freedom (where the degrees of freedom equal n − p − 1, where p is the number of parameters estimated in the model (one for each variable in the regression equation, not including the intercept)). One can then also calculate the mean square of the model by dividing the sum of squares of the model minus the degrees of freedom, which is just the number of parameters. Then the F value can be calculated by dividing the mean square of the model by the mean square of the error, and we can then determine significance (which is why you want the mean squares to begin with.).^[8]

However, because of the behavior of the process of regression, the distributions of residuals at different data points (of the input variable) may vary even if the errors themselves are identically distributed. Concretely, in a

regression coefficients

: endpoints have more influence.

Thus to compare residuals at different inputs, one needs to adjust the residuals by the expected variability of residuals, which is called

outliers

, where the case in question is somehow different from the others in a dataset. For example, a large residual may be expected in the middle of the domain, but considered an outlier at the end of the domain.

Other uses of the word "error" in statistics

The use of the term "error" as discussed in the sections above is in the sense of a deviation of a value from a hypothetical unobserved value. At least two other uses also occur in statistics, both referring to observable

prediction errors

:

The

mean squared error

(MSE) refers to the amount by which the values predicted by an estimator differ from the quantities being estimated (typically outside the sample from which the model was estimated). The root mean square error
(RMSE) is the square-root of MSE. The sum of squares of errors (SSE) is the MSE multiplied by the sample size.

Sum of squares of residuals (SSR) is the sum of the squares of the deviations of the actual values from the predicted values, within the sample used for estimation. This is the basis for the least squares

estimate, where the regression coefficients are chosen such that the SSR is minimal (i.e. its derivative is zero).

Likewise, the

sum of absolute errors (SAE) is the sum of the absolute values of the residuals, which is minimized in the least absolute deviations

approach to regression.

The mean error (ME) is the bias. The mean residual (MR) is always zero for least-squares estimators.

References

ISBN 978-1-4051-8257-7
. Retrieved 2022-05-13.

ISBN 978-1-337-67133-0
. Retrieved 2022-05-13.

ISBN 978-981-329-019-8
. Retrieved 2022-05-13.

OCLC 7779780
.

^
OCLC 262680588
.

OCLC 987251007
.

^ Steel, Robert G. D.; Torrie, James H. (1960). Principles and Procedures of Statistics, with Special Reference to Biological Sciences. McGraw-Hill. p. 288.

ISBN 9780521761598
.

^ "7.3: Types of Outliers in Linear Regression". Statistics LibreTexts. 2013-11-21. Retrieved 2019-11-22.

Further reading

Cook, R. Dennis; Weisberg, Sanford (1982). Residuals and Influence in Regression (Repr. ed.). New York:
ISBN 041224280X
. Retrieved 23 February 2013.

JSTOR 2984505
.

Weisberg, Sanford (1985). Applied Linear Regression (2nd ed.). New York: Wiley.
ISBN 9780471879572
. Retrieved 23 February 2013.

"Errors, theory of", Encyclopedia of Mathematics, EMS Press, 2001 [1994]

External links

Media related to Errors and residuals at Wikimedia Commons

v
t
e
Least squares and regression analysis
Computational statistics

Least squares

Linear least squares

Non-linear least squares

Iteratively reweighted least squares

Correlation and dependence

Pearson product-moment correlation

Rank correlation (Spearman's rho

Kendall's tau
)

Partial correlation

Confounding variable

Regression analysis

Ordinary least squares

Partial least squares

Total least squares

Ridge regression

Regression as a
statistical model
Linear regression

Simple linear regression

Ordinary least squares

Generalized least squares

Weighted least squares

General linear model

Predictor structure

Polynomial regression

Growth curve (statistics)

Segmented regression

Local regression

Non-standard

Nonlinear regression

Nonparametric

Semiparametric

Robust

Quantile

Isotonic

Non-normal errors

Generalized linear model

Binomial

Poisson

Logistic

Decomposition of variance

Analysis of variance

Analysis of covariance

Multivariate AOV

Model exploration

Stepwise regression

Model selection
Mallows's C_p

AIC

BIC

Model specification

Regression validation

Background

Mean and predicted response

Gauss–Markov theorem

Errors and residuals

Goodness of fit

Studentized residual

Minimum mean-square error

Frisch–Waugh–Lovell theorem

Design of experiments

Response surface methodology

Optimal design

Bayesian design

Numerical approximation

Numerical analysis

Approximation theory

Numerical integration

Gaussian quadrature

Orthogonal polynomials

Chebyshev polynomials

Chebyshev nodes

Applications

Curve fitting

Calibration curve

Numerical smoothing and differentiation

System identification

Moving least squares

Regression analysis category

Statistics category

Mathematics portal

Statistics outline

Statistics topics

Retrieved from "https://en.wikipedia.org/w/index.php?title=Errors_and_residuals&oldid=1186819298"

[Kennedy_2008_p._576-1] ISBN 978-1-4051-8257-7
. Retrieved 2022-05-13.

[Wooldridge_2019_p._57-2] ISBN 978-1-337-67133-0
. Retrieved 2022-05-13.

[Das_2019_p._7-3] ISBN 978-981-329-019-8
. Retrieved 2022-05-13.

[4] OCLC 7779780
.

[modernintro-5] 
OCLC 262680588
.

[6] OCLC 987251007
.

[7] Steel, Robert G. D.; Torrie, James H. (1960). Principles and Procedures of Statistics, with Special Reference to Biological Sciences. McGraw-Hill. p. 288.

[8] ISBN 9780521761598
.

[9] "7.3: Types of Outliers in Linear Regression". Statistics LibreTexts. 2013-11-21. Retrieved 2019-11-22.

[1]

[2]

[3]

[4]

[6]

[7]

[8]

Introduction

In univariate distributions

Remark

Regressions

Other uses of the word "error" in statistics

See also

References

Further reading

External links