Skip to the calculator
SwiftSums
Math & statistics

P-Value Calculator. From any test statistic.

Turn a test statistic into a p-value for a z, t, chi-square, or F test.

Live calculation

What's the p-value?

Which test?
Tail
p-value 0.0500

Your figures are kept on this device only.

How it works

A p-value is the probability of seeing a test statistic at least this extreme, if there were truly no effect. It's computed from the area under the relevant distribution's curve beyond your statistic — the normal curve for a z-test, or the shape appropriate to t, chi-square, and F tests, each of which accounts for sample size via degrees of freedom.

z: p = 2 × (1 − Φ(|z|))  — two-tailed, Φ = standard normal CDF
t: p = Iν/(ν+t²)(ν/2, 1/2)  — two-tailed, ν = degrees of freedom
χ²: p = Q(df/2, χ²/2)  — upper tail, Q = regularized upper incomplete gamma
F: p = 1 − Ix(df₁/2, df₂/2)  — upper tail, x = df₁F/(df₁F+df₂)

Chi-square and F tests are conventionally reported as a single upper-tail p-value — there's no separate "two-tailed" version, since both distributions only take non-negative values and a large statistic is what signals evidence against the null either way.

Worked example

A two-tailed z-test with statistic z = 1.96.

Step by step
  1. Find the area to the right of 1.96 under the standard normal curve: 1 − Φ(1.96) ≈ 0.0250.
  2. Double it for a two-tailed test: 2 × 0.0250 = 0.0500.

p ≈ 0.0500 — the classic z = 1.96 ↔ p = 0.05 landmark.

Common questions

How small should a p-value be to matter?

There's no universal cutoff, but 0.05 is the most common convention for "statistically significant" in many fields — treat it as a widely used threshold, not a law of nature, and check what your field or study protocol actually requires.

When do I use one-tailed vs. two-tailed?

Two-tailed tests whether the result differs from the null in either direction; one-tailed tests only one specific direction, which you must decide before seeing the data. One-tailed tests should be used sparingly — deciding the direction after seeing results inflates your false positive rate.

Which test should I pick?

z-tests suit large samples with known population variance; t-tests suit smaller samples or unknown variance; chi-square tests categorical/count data and goodness-of-fit; F-tests compare variances (e.g. in ANOVA). If you're unsure which your study design calls for, that choice happens before this calculator — it just converts the statistic you already have.