How it works
A p-value is the probability of seeing a test statistic at least this extreme, if there were truly no effect. It's computed from the area under the relevant distribution's curve beyond your statistic — the normal curve for a z-test, or the shape appropriate to t, chi-square, and F tests, each of which accounts for sample size via degrees of freedom.
z: p = 2 × (1 − Φ(|z|)) — two-tailed, Φ = standard normal CDF
t: p = Iν/(ν+t²)(ν/2, 1/2) — two-tailed, ν = degrees of freedom
χ²: p = Q(df/2, χ²/2) — upper tail, Q = regularized upper incomplete gamma
F: p = 1 − Ix(df₁/2, df₂/2) — upper tail, x = df₁F/(df₁F+df₂)
Chi-square and F tests are conventionally reported as a single upper-tail p-value — there's no separate "two-tailed" version, since both distributions only take non-negative values and a large statistic is what signals evidence against the null either way.
Worked example
A two-tailed z-test with statistic z = 1.96.
- Find the area to the right of 1.96 under the standard normal curve: 1 − Φ(1.96) ≈ 0.0250.
- Double it for a two-tailed test: 2 × 0.0250 = 0.0500.
p ≈ 0.0500 — the classic z = 1.96 ↔ p = 0.05 landmark.
Common questions
How small should a p-value be to matter?
There's no universal cutoff, but 0.05 is the most common convention for "statistically significant" in many fields — treat it as a widely used threshold, not a law of nature, and check what your field or study protocol actually requires.
When do I use one-tailed vs. two-tailed?
Two-tailed tests whether the result differs from the null in either direction; one-tailed tests only one specific direction, which you must decide before seeing the data. One-tailed tests should be used sparingly — deciding the direction after seeing results inflates your false positive rate.
Which test should I pick?
z-tests suit large samples with known population variance; t-tests suit smaller samples or unknown variance; chi-square tests categorical/count data and goodness-of-fit; F-tests compare variances (e.g. in ANOVA). If you're unsure which your study design calls for, that choice happens before this calculator — it just converts the statistic you already have.