Skip to main content
CalcMax

Sum of Squares Calculator

Result

32.0000

Sum of squared deviations

Sum of squared values (Σx²)
232.0000
Sum
40.0000
Mean
5.0000
Count
8

The sum of squares on this page is the sum of squared deviations from the mean: take every value, subtract the mean, square the result, and add those squares up. It is the number the variance and the standard deviation are built from — the variance is this quantity divided by the degrees of freedom — and it is the answer to a question those two pages never stop to ask, which is why the differences are squared at all. Plain deviations add up to exactly zero, every time, so something has to be done to them before they can be added. This sum of squares calculator prints the deviation sum of squares together with the three quantities of the identity that gives the same number another way, so both routes can be checked against each other.

Formula

SS = Σ(x − x̄)² = Σx² − (Σx)² / n

x
Each value in the list, one at a time. Every value contributes its own squared gap, so no single value can be dropped without changing the total
x̄
The mean of the values. It is the point the gaps are measured from, and it is chosen precisely because the plain deviations around it cancel to zero
SS
The sum of squared deviations — the main result. It is zero exactly when every value is identical, and it grows with both the size of the deviations and how many of them there are
Σx²
The sum of the squared values, printed as raw sum of squares too. It is the first term on the right of the identity and the easiest of the two to confuse with the answer, since both are called a sum of squares
Σx
The plain sum of the values, before squaring. Squaring it and dividing by n gives the correction term that turns Σx² into the deviation sum of squares
n
How many values went in, up to 200 here. It appears twice — as the divisor in the correction term and as the count printed for checking

Reach for it when you want to see the step rather than the finished measure: when checking a variance or standard deviation by hand, when explaining why those two quantities square their deviations, or when comparing how much of a total's variation a model accounts for. It is also the natural way to express spread when you are working with sums rather than averages, as in an analysis of variance, where the whole method is bookkeeping on sums of squares. In a regression with an intercept, this same quantity is what is called the total sum of squares, and the model's fit is measured by how much of it the fitted line removes. The identity on the right is worth knowing for a second reason: it is how the number was computed before machines, and it is still how many formulas are written, even though computing it that way loses precision when the values are large and their spread is small. That is why the calculator above uses the definitional route and prints the terms of the identity only so you can check the two agree.

Worked examples

  1. Eight values: 32 by both routes

    1. The mean is 40 / 8 = 5
    2. The deviations are −3, −1, −1, −1, 0, 0, 2 and 4, and they add to 0 — as they always do
    3. Squaring them gives 9, 1, 1, 1, 0, 0, 4 and 16, which add to 32
    4. The other route: Σx² = 232 and (Σx)² / n = 1600 / 8 = 200, so 232 − 200 = 32

    Both routes land on 32, and seeing that is most of the point of the page. The first three deviations are negative and the last two positive, which is exactly why the squares are needed — the raw deviations sum to zero and could not tell one data set from another.

  2. Two values with a fractional mean

    1. The mean is 3 / 2 = 1.5, so the deviations are −0.5 and +0.5
    2. Squaring gives 0.25 and 0.25, so the sum of squares is 0.5 — not 1
    3. The other route: 5 − 3² / 2 = 5 − 4.5 = 0.5, where the fractional part appears in the subtraction

    Two values can never produce a sum of squares larger than half their squared gap, and this is the smallest case: each of the two deviations carries half the distance. A common slip is to square the gap itself, which would give 1 here instead of 0.5 — the gap is already the difference of two deviations, so it gets counted twice.

  3. A single value: 0 against 1764

    1. The mean of a single value is that value, so 42 is the mean
    2. The deviation is 42 − 42 = 0, and 0² is 0, so the sum of squares is 0
    3. The raw sum of squares is the other quantity entirely: 42² = 1764

    Two things called a sum of squares, three orders of magnitude apart, in the same panel. This is the case that keeps them apart: with a longer list the two numbers are usually the same order of magnitude and it is easy to reach for the wrong one. The deviation sum of squares is 0 because a single value cannot deviate from itself — that is a real answer, not a failure, and it is why this page does not require two values the way the variance does.

  4. Four identical values

    1. The mean is 12 / 4 = 3, so every deviation is 0
    2. The sum of squares is 0
    3. The raw sum of squares is 3² + 3² + 3² + 3² = 36, which is not 0 and never will be

    Zero here means the data has no spread at all, and it is a different situation from the single-value case above even though both print 0: here there are four observations and they agree exactly, there just one observation exists. The raw sum of squares moves in the opposite direction, growing with every value added regardless of spread.

Limitations

The name is the main hazard. Two different quantities are both called a sum of squares and both appear on this page: the deviation sum of squares Σ(x − x̄)², which is the answer, and the raw sum of squares Σx², which is only a term in the identity. On a single value they differ by three orders of magnitude. Read the label on the row before quoting a number. The second limitation is what this page does not do: a regression splits the total sum of squares into a part explained by the fitted line and a residual part, and that split is not available here. It needs a second variable and a line fitted through the data, and no page in this section fits one — the covariance page stops at the correlation coefficient. What is true, and is worth knowing, is that the total sum of squares in a regression with an intercept is this same quantity applied to the outcome variable; it is only the split that is missing. There is also a practical reason the calculator uses the definitional route rather than the identity: subtracting two large, nearly equal numbers loses precision, so on a list of large values with a small spread the right-hand route can give an answer that is visibly wrong. The identity is printed for checking, not for speed. The list is capped at 200 values, and a token such as 1,500 is refused rather than guessed at, because a comma between digits is a decimal point in some countries and a thousands separator in others.

Frequently asked questions

How do you calculate the sum of squares?
Find the mean, subtract it from each value, square every result and add them up. For 2, 4, 4, 4, 5, 5, 7 and 9 the mean is 5, the deviations are −3, −1, −1, −1, 0, 0, 2 and 4, and their squares add to 9 + 1 + 1 + 1 + 0 + 0 + 4 + 16 = 32. Every value takes part, including the ones that happen to sit on the mean, which contribute 0.
Why square the deviations instead of just adding them?
Because the plain deviations always add up to exactly zero. The values below the mean cancel the values above it, for any list whatsoever, so the sum of the deviations carries no information about spread at all. Squaring does two things at once: it makes every term positive so nothing cancels, and it gives values far from the mean proportionally more weight — a value two units out contributes 4, one twenty units out contributes 400. That second effect is deliberate. It is also why the sum of squares, and the variance built from it, react strongly to a single extreme value.
Why are there two different numbers called a sum of squares on this page?
Because two quantities in statistics share the name, and this page prints both. The sum of squared deviations, Σ(x − x̄)², is the answer here and the numerator of the variance. The raw sum of squares, Σx², is the sum of the values each squared before any mean is subtracted. They are related by the identity Σ(x − x̄)² = Σx² − (Σx)² / n, which is why the raw version is on screen, but they are not interchangeable. With the single value 42 the deviation sum of squares is 0 while the raw one is 1764.
Does the sample or population choice change the answer?
No, and that is why there is no selector on this page. The sample and population formulas differ only in their divisor — n − 1 against n — and the sum of squares does not divide by anything. It is one number for a given list. The choice appears on the variance and standard deviation pages because the division happens there, and it is exactly why the same data set can produce two variances but only one sum of squares.
What is the connection to the variance and standard deviation?
The variance is this number divided by the degrees of freedom: n − 1 for a sample or n for a population. The standard deviation is the square root of the variance. So the sum of squares is the raw material, and the division is what makes it comparable between data sets of different sizes — a sum over eight values is naturally larger than a sum over four, even when the spread is the same. Dividing by n − 1 rather than n also corrects a slight underestimate that comes from measuring deviations around the sample mean instead of the true one.
Can the sum of squares be negative?
Never. Every term is a square, so every term is zero or positive and the total is too. It is zero exactly when every value equals the mean, which means every value equals every other value — either because the list is a single value, or because all the values are identical. A total of zero is a real result meaning the data has no spread, not a sign that something went wrong. If you see a negative number, it is the raw sum of squares Σx² you are looking at with a sign error, or a difference of squares computed by hand.

References

Related calculators