Statistical Significance Calculator
Two-proportion z-test for A/B experiments. Enter your numbers, get a verdict.
What is statistical significance?
When you run an A/B test, you see a difference in conversion rates between control and variant. Statistical significance tells you whether that difference is real or just noise. Even identical pages will show some variation in conversion rates due to random visitor behavior. A significance test quantifies the probability that the observed difference happened by chance.
How this calculator works
This tool runs a two-proportion z-test. It takes the conversion rate from each group, pools them to estimate variance, then computes a z-score — how many standard deviations apart the two rates are. That z-score maps to a p-value via the normal distribution.
The bell curves above show the sampling distribution of each group's conversion rate. When the curves overlap heavily, the difference could easily be noise. When they separate, you have evidence of a real effect.
Interpreting results
At 95% confidence (the standard), a p-value below 0.05 means the result is significant — the observed difference is unlikely to be random. A p-value above 0.05 doesn't mean the pages are identical; it means you can't tell yet. You may need more data.
The confidence interval shows the range where the true difference between variants likely falls. If the interval doesn't cross zero, the result is significant at your chosen level.
Common mistakes
Peeking at results before you've collected enough data inflates false positives — you'll see significance that isn't real. Decide your sample size upfront and wait. Stopping at the first sign of significance is the same problem; early p-values fluctuate and often reverse. And if you don't have enough traffic to begin with, consider whether an A/B test is the right tool. Use a sample size calculator to plan ahead.
FAQ
What p-value threshold should I use?
0.05 (95% confidence) is the standard. Use 0.01 for high-stakes decisions like pricing. Use 0.10 for exploratory tests where speed matters more than certainty.
How many visitors do I need?
Depends on your baseline rate and the effect you want to detect. Rough guide: detecting a 10% relative lift on a 5% conversion rate at 95% confidence needs ~30,000 visitors per variant.
One-tailed vs two-tailed?
Two-tailed tests whether the variant differs in either direction. One-tailed only tests if it's better. Default to two-tailed unless you're certain the variant can't hurt (rare in practice).
Can I peek at results early?
Looking is fine, acting on incomplete data is not. Repeated checking inflates your false positive rate. If you need continuous monitoring, use sequential testing methods that adjust for multiple looks.
What does “not significant” mean?
Not enough evidence to conclude the versions differ. It doesn't mean they're the same — the true difference might just be too small to detect with your sample size.
SplitMonk runs significance tests automatically across all your experiments — no manual number entry needed. Try it free