T-Test Calculator
Result
T statistic
- P-value
- 34.6594%
- Degrees of freedom
- 8
- Standard error
- 1.0000
A t test calculator runs the most common test of a mean: it turns a sample mean, a standard deviation and a sample size into a t statistic, a p-value and the degrees of freedom the other two are read at. One sample mode compares a mean against a fixed value — a target, a specification, a historical average. Two sample mode compares two groups against each other, and the version implemented here is the pooled two sample t test, which assumes the two groups share a variance: the two standard deviations are averaged with weights of n − 1, and that single pooled variance drives the standard error. The alternative, Welch's test, drops the equal-variance assumption and is the safer default when the groups differ in size and spread, but its degrees of freedom come out fractional and it is a different computation in a different place in the code; this page does the equal-variance test and says so rather than quietly approximating it. The p-value is the tail area of the t distribution beyond the statistic, and it is printed as a percentage.
Formula
t = (x̄ − μ₀) / (s / √n) t = (x̄₁ − x̄₂) / √(s_p²(1/n₁ + 1/n₂)) s_p² = ((n₁−1)s₁² + (n₂−1)s₂²) / (n₁ + n₂ − 2)
- t
- The test statistic — how many standard errors the observed mean sits from the value being tested. Its sign carries the direction, which is why a one-tailed test in the wrong direction returns a large p-value rather than an error: a t of −1 against a greater-than alternative is not a small effect, it is evidence pointing the other way, and the tail area says so
- x̄, μ₀
- The sample mean and the value it is being tested against, in the one-sample mode. The numerator is their difference, so the statistic measures a distance in the units of the data — the same difference counts for much less if the spread is wide or the sample small, which is what the denominator supplies
- s
- The sample standard deviation, estimated from the data rather than known in advance. That estimation is the entire reason the t distribution is used instead of the normal: the extra uncertainty it introduces makes the tails heavier, and the heavier tails are what the t critical value accounts for. It must be greater than zero — a sample with no variation has no denominator to divide by
- n
- The sample size, giving n − 1 degrees of freedom in the one-sample mode. For the two-sample test the two sizes add up to one more than either alone: n₁ + n₂ − 2, because the pooled variance spends one degree of freedom on each group's mean before measuring spread about it
- s_p²
- The pooled variance — the weighted average of the two groups' variances, with each weighted by its own degrees of freedom. This is where the equal-variance assumption enters: the two groups are treated as two samples from a single population spread, so a group with more observations, or one whose own variance is smaller, contributes proportionally to a single shared estimate. If the two variances differ substantially the pooling is what goes wrong, and the Welch test is the repair
- p-value
- The probability of a statistic at least this far from zero, in the direction the test asks about, if the two means were really the same. Two tails doubles the area beyond |t|; the greater and less alternatives take one tail each. It is reported as a percentage here, so 5% is the familiar 0.05, and it is a statement about the data given the hypothesis rather than about the hypothesis given the data
Use it when you have a mean and a standard deviation from a sample and want to know whether the difference from a target, or from another group, is larger than sampling noise would explain. The one-sample mode handles the case where the comparison value is a fixed number — a machine set to fill 500 grams, a specification of 5 mm, last year's average — and the two-sample mode handles the case where both quantities were measured and neither is privileged. Two things are worth deciding before reading the answer. The first is one tail or two: a two-tailed test asks whether the groups differ in either direction and is the right default, while a one-tailed test asks a directional question and should be chosen because the direction was specified in advance, not because the answer would be more favourable. The second is the equal-variance assumption, which the two-sample mode makes and cannot check for you. The tell is on the panel: if the two standard deviations you typed are close, pooling them costs little, and if they differ by a factor of two or more the Welch test is the one you want. This page computes the pooled test, which is what most textbooks and most statisticians mean by "the t test" in the equal-variance case, and reports the two standard deviations you supplied so the assumption can be judged rather than assumed.
Worked examples
The default: a sample mean of 6 against a target of 5
- Standard error: s / √n = 3 / 3 = 1
- t statistic: (6 − 5) / 1 = 1 — the sample mean sits one standard error above the target
- Degrees of freedom: n − 1 = 8
- Two-tailed p-value: the area beyond ±1 on a t distribution with 8 degrees of freedom, which is 34.66%
A difference of one whole unit sounds like a lot until the standard error is put next to it: one standard error is well within the range that sampling noise produces, and the p-value says so. The comparison with the normal curve is instructive — the same t of 1 would give 31.7% if the distribution were normal, and the extra three points are the price of having estimated the standard deviation from nine observations rather than knowing it.
A t statistic of 2.306 gives exactly 5%
- Standard error: 3 / √9 = 1, so the t statistic is the mean itself
- t statistic: (2.306 − 0) / 1 = 2.306
- The 95% two-tailed t boundary at 8 degrees of freedom is 2.306
- Sitting exactly on the boundary means the two-tailed p-value is exactly 5%
This is the pair the critical value page produces from the other direction, and putting the two together is the clearest demonstration that they are one fact read from two ends. Ask that page for a 95% two-tailed t boundary with 8 degrees of freedom and it returns 2.306; feed 2.306 into this one and the p-value comes back at exactly 5%. Nothing in between is rounded — the number was chosen because the standard deviation of 3 and the sample size of 9 make the standard error exactly 1, so the statistic and the boundary are the same figure. A statistic just below it would land above 5%, and one just above it below 5%.
Two groups of ten with the same spread
- Pooled variance: ((9 × 4) + (9 × 4)) / 18 = 72 / 18 = 4, so the pooled standard deviation is 2
- Standard error: √(4 × (1/10 + 1/10)) = √0.8 = 0.8944
- t statistic: (5 − 4) / 0.8944 = 1.118
- Degrees of freedom: 10 + 10 − 2 = 18, giving a two-tailed p-value of 27.83%
When the two variances are equal the pooling is a no-op — the pooled variance comes out at 4, exactly the value both groups supplied — and the standard error is smaller than either group's own standard deviation divided by its own √n. Comparing two means of ten is less precise than estimating a single mean of ten — a standard error of 0.8944 against 0.6325 — because two uncertain quantities are being subtracted rather than one being measured, and it is only worth that cost when the comparison answers a question a single sample cannot. Note that the population mean field is on screen and unused in this mode; both quantities are measured, so neither is the fixed target.
Unequal spreads and unequal sizes, where the pooling has to be watched
- Pooled variance: ((9 × 4.41) + (11 × 3.24)) / 20 = (39.69 + 35.64) / 20 = 3.7665
- Standard error: √(3.7665 × (1/10 + 1/12)) = √0.6905 = 0.831
- t statistic: (5.2 − 4.1) / 0.831 = 1.3237
- Degrees of freedom: 10 + 12 − 2 = 20
The pooled variance is 3.7665, which is not the midpoint of 4.41 and 3.24 — the group of twelve is weighted more heavily because it contributes eleven degrees of freedom to the group of ten's nine. That weighting is the whole content of the word pooled, and it is why the answer cannot be reproduced by averaging the two standard deviations. The two spreads here are close enough that the equal-variance assumption is comfortable; a pair like 2.1 and 0.6 with these group sizes would not be, and Welch's test would give a visibly different p-value.
A t of 1.96 from 900 observations, where t and z have converged
- Standard error: 3 / √900 = 3 / 30 = 0.1
- t statistic: (5.196 − 5) / 0.1 = 1.96
- Degrees of freedom: 899
- Two-tailed p-value: 5.0304% — just over 5, not exactly at it
The famous 1.96 gives 5% on the normal curve and 5.0304% here, and the gap is the point of the example. With 899 degrees of freedom the t distribution is nearly indistinguishable from the normal, but not exactly — it is still slightly heavier in the tails, so the same statistic sits slightly further out. This is also why a t of 1.96 is not automatically significant at 5% in a large sample, and why software that reports 0.0499 and software that reports 0.0501 can disagree about the same study at the third decimal place.
Limitations
The two-sample mode assumes the two groups share a variance, and that assumption is doing real work rather than tidying up notation. When the group sizes are equal the pooled test is fairly robust even if the variances differ, but when the sizes are unequal and the variances differ too, the test can be seriously wrong in either direction — this is the Behrens–Fisher problem, and Welch's test exists to solve it. The two standard deviations are on the panel for that reason: compare them before trusting the p-value, and if one is more than about twice the other while the group sizes differ, the equal-variance test is not the one you want. The one-sample mode has no such issue but assumes the observations are independent, which fails for repeated measurements on one subject and for matched or clustered data, where a paired or mixed model is needed. Both modes assume the underlying data is roughly normal, and both become less sensitive to that as the sample grows — the central limit theorem is what makes a t test on 900 skewed observations reasonable and a t test on 9 skewed observations questionable. Finally, the page reports a statistic and a p-value and no conclusion. It has no significance level to compare against, and the number 0.05 is a convention rather than a result, so the decision is left where it belongs.
Frequently asked questions
- Does this page use the pooled or Welch's two-sample t test?
- The pooled one — the equal-variance test, with n₁ + n₂ − 2 degrees of freedom. Welch's test is the better choice when the two groups differ in spread and in size, and the two standard deviation fields are on the panel so that you can check: if one is more than roughly twice the other and the sample sizes differ, Welch is what you want and this page will give a p-value that is off, sometimes substantially. The equal-variance version is implemented here because it is the textbook test, it is what a reference manual calls the two-sample t test for equal means, and its whole-number degrees of freedom keep the arithmetic on this page exact rather than approximated.
- Should I use a one-tailed or a two-tailed test?
- Two-tailed unless the direction was specified before the data was seen. A two-tailed test asks whether the groups differ at all and splits the significance level between the two ends; a one-tailed test asks whether one group is larger and puts the whole level in that tail, which makes it easier to reach significance — exactly twice as easy at the same level, since 1.96 becomes 1.645. That is a legitimate choice when the question really is directional, and it is a form of result-chasing when the direction was chosen after seeing which way the difference went.
- What is the difference between the t statistic and the p-value?
- The t statistic measures the difference in standard errors and carries the sample size inside it, while the p-value converts that figure into a probability using the t distribution at the matching degrees of freedom. So t answers "how many standard errors apart are these?" and the p-value answers "how often would sampling alone produce a gap this large?" Both are reported because the first is comparable across studies with the same design and the second is what a significance level is compared against. Neither is the size of the effect — a large sample can make a trivial difference produce a large t and a tiny p-value.
- Why does a t statistic of 1.96 not give exactly 5%?
- Because 1.96 is the normal curve's 5% boundary, and the t distribution is slightly heavier in the tails at any finite degrees of freedom. With 899 degrees of freedom a t of 1.96 gives 5.0304%, just over the line; at 8 degrees of freedom the same statistic gives 8.57%. The t distribution converges on the normal as the degrees of freedom grow, and it is the same curve throughout — only the weight of the tails changes. This is why software can report a p-value of 0.0499 where a hand calculation against 1.96 says significant, and neither is wrong.
- What does a negative t statistic mean?
- That the sample mean came out below the value being tested against, or below the second group's mean in the two-sample mode. The sign is direction information and nothing more — its size relative to the distribution is what the p-value is computed from. In a two-tailed test the sign is discarded because either extreme counts. In a one-tailed test it matters and can be informative: a t of −2.5 against a greater-than alternative gives a p-value near 99%, which is not a null result but evidence pointing the other way, and reporting it as "not significant" would throw that information away.
- How many observations do I need for a t test?
- The page requires at least two per group, because the standard deviation of a single observation is undefined. That is a mathematical floor rather than a practical recommendation — a t test on two observations per group has one degree of freedom and almost no power to detect anything but an enormous difference, and the p-value will be large unless the two groups barely overlap. The useful question is not the minimum the formula allows but the sample size needed to detect the difference you care about, which depends on the spread of the data and the effect size you would want to notice; the sample size page is where that calculation lives, and it works backwards from a margin of error rather than forwards from a dataset.
- Where is the table of t critical values?
- On the critical value page, and only there in this set — this page does not print one because it would have to be indexed by degrees of freedom as well as by significance level, which is a page of the book per significance level rather than a table. That is also why the answer is computed here rather than looked up: a table can only list the degrees of freedom somebody chose to print, and the degrees of freedom a two-sample test produces, n₁ + n₂ − 2, is almost never one of them. The one boundary worth carrying in your head is that the t values converge on 1.96 as the sample grows, so the table you are looking for is the normal one with a penalty attached that shrinks as the degrees of freedom rise.
References
- 1.3.5.3. Two-Sample t-Test for Equal Means — e-Handbook of Statistical Methods (the pooled two-sample test this page implements, including the pooled standard deviation and its n₁ + n₂ − 2 degrees of freedom) — National Institute of Standards and Technology (NIST)
- 1.3.5.1. Measures of Location — e-Handbook of Statistical Methods (the sample mean and the standard error that the one-sample statistic is built from) — National Institute of Standards and Technology (NIST)
- 1.3.6.7.1. Cumulative Distribution Function of the Standard Normal Distribution — e-Handbook of Statistical Methods (the curve the t distribution approaches as the degrees of freedom grow, which is why 1.96 and 5% are only approximately each other's partner) — National Institute of Standards and Technology (NIST)