You probably just saw p=0.04 in your dashboard
If you're reading this, you're staring at an A/B test result and want to know if you can ship. Short answer: if p < 0.05 and you ran the test properly, probably yes. But "properly" is doing a lot of work in that sentence.
Let me explain what these numbers actually mean, for people who build products and don't have time for a statistics textbook.
The coin flip analogy
Imagine I hand you a coin and claim it's rigged. You flip it 10 times and get 7 heads. Is the coin rigged?
Probably not. A fair coin gives 7+ heads about 17% of the time. That's not rare enough to be suspicious.
Now you flip it 100 times and get 70 heads. Same ratio (70%), but now the probability of this happening with a fair coin is 0.00004%. That coin is rigged.
That's what statistical significance measures: how unlikely your results would be if there were actually no difference between your variants. The more data you have, the more confident you can be that a difference is real and not just luck.
The p-value is that probability. p=0.04 means "if the variants were identical, there's a 4% chance we'd see a difference this large just by random chance." Since 4% is below our threshold (usually 5%), we call it significant.
What p-values don't tell you
Here's where most people go wrong. A p-value of 0.04 does NOT mean:
- "There's a 96% chance the variant is better" (this is a Bayesian statement; p-values don't work this way)
- "The variant improves conversion by 4%" (p-value says nothing about the size of the effect)
- "If we run this test again, we'll get the same result 96% of the time" (also wrong)
What it does mean: "If the truth is that A and B are exactly the same, there's a 4% chance we'd observe data this extreme or more extreme." That's it. It's a statement about the data, not about the truth.
I know this feels like a distinction without a difference. But it matters. A p-value of 0.04 for a 0.1% conversion lift means the difference is real but tiny. A p-value of 0.04 for a 15% lift means the difference is real and worth shipping. The p-value is the same in both cases.
Confidence intervals: what you actually want
If p-values answer "is there a difference?", confidence intervals answer "how big is the difference?" This is the number you should be making decisions on.
A 95% confidence interval of [+2.1%, +8.7%] means:
- The effect is definitely positive (the entire interval is above zero)
- The true lift is probably somewhere between 2.1% and 8.7%
- The result is statistically significant (zero is not in the interval)
A 95% confidence interval of [-0.8%, +5.2%] means:
- The result is NOT significant (zero is in the interval)
- But the effect is probably positive (most of the interval is above zero)
- You need more data to be sure
Think of the confidence interval as a "plausible range" for the true effect. The narrower it is, the more precise your estimate. It narrows as you collect more data.
The "I have low traffic" problem
This is the question I get most often. "We only get 2,000 visitors/month. Can we even run A/B tests?"
It depends on your baseline conversion rate and the minimum effect you're trying to detect. Here are some real numbers:
Baseline rate: 3%
Minimum detectable effect: 20% relative (3.0% → 3.6%)
Required sample: ~8,200 per variant → 16,400 total
At 2,000/month: ~8 months. Not practical.
Baseline rate: 3%
Minimum detectable effect: 50% relative (3.0% → 4.5%)
Required sample: ~1,400 per variant → 2,800 total
At 2,000/month: ~6 weeks. Doable.
Baseline rate: 10%
Minimum detectable effect: 20% relative (10% → 12%)
Required sample: ~3,600 per variant → 7,200 total
At 2,000/month: ~4 months. Borderline.
The rule of thumb: if you can't reach the required sample size within 4 weeks, either test bolder changes (bigger MDE) or focus on pages with more traffic.
In SplitMonk, we show you the estimated time to significance before you start a test. If the estimate is longer than 6 weeks, we flag it and suggest testing a larger change. Most people waste months running underpowered tests when they could've gotten a clear answer in 2 weeks by testing a bigger swing.
The four things that control your test
Every A/B test has four dials. Turn one, and the others adjust:
-
Significance level (alpha): Your tolerance for false positives. Industry standard is 5%. Lower = need more data but fewer false alarms.
-
Power: Your ability to detect real effects. Standard is 80%. This means 20% of the time, you'll miss a real winner and call it inconclusive.
-
Minimum detectable effect (MDE): The smallest improvement you care about. Smaller MDE = need way more data. A 5% relative MDE needs roughly 4x the data of a 10% MDE.
-
Sample size: Falls out of the other three. It's not a number you pick -- it's a number you calculate.
// This is roughly how SplitMonk calculates required sample size
function requiredSampleSize(
baselineRate: number,
mde: number, // relative, e.g. 0.10 for 10%
alpha: number = 0.05,
power: number = 0.80
): number {
const p1 = baselineRate;
const p2 = baselineRate * (1 + mde);
const zAlpha = 1.96; // two-sided, alpha=0.05
const zBeta = 0.84; // power=0.80
const n = Math.pow(
zAlpha * Math.sqrt(2 * p1 * (1 - p1)) +
zBeta * Math.sqrt(p1 * (1-p1) + p2 * (1-p2)),
2
) / Math.pow(p2 - p1, 2);
return Math.ceil(n); // per variant
}
// Example: 3% baseline, want to detect 15% relative lift
requiredSampleSize(0.03, 0.15);
// → ~24,100 per variant → ~48,200 total
The cheat sheet I use
I keep these rules taped to my monitor. Literally.
-
Before starting: Calculate required sample size. If it's more than 4 weeks of traffic, test a bigger change.
-
While running: Don't peek. Or if you must peek (you will), use a tool with Bayesian stats or proper sequential testing that handles continuous monitoring. SplitMonk uses Bayesian by default for exactly this reason.
-
At conclusion:
- p < 0.05 (or >95% Bayesian probability) AND the confidence interval doesn't include effects you'd consider trivial? Ship it.
- p < 0.05 but the effect is tiny (like +0.3% conversion)? Probably not worth the maintenance cost. Skip it.
- p > 0.05 but the confidence interval is mostly positive? Extend the test if you can. You probably need more data.
- p > 0.05 and the confidence interval straddles zero evenly? No effect. Move on to a different hypothesis.
-
Always: Check SRM (sample ratio mismatch) before trusting any result. Check mobile vs. desktop segments. Check guardrail metrics.
-
Never: Stop a test early because the numbers "look good enough." Never run more than 4 variants without correcting for multiple comparisons. Never extrapolate results from one page to another.
Statistical significance isn't complicated. It's just a measure of "can I trust this result?" Combined with confidence intervals that tell you "how big is the effect?", you have everything you need to make good shipping decisions. The math handles the uncertainty. Your job is to ask good questions and be patient enough to let the data answer them.

Michał Pogoda-Rosikoń
Founder
Founder of SplitMonk and bards.ai. Data scientist from Wroclaw University of Technology, specializing in NLP and machine learning. Building AI-powered tools that optimize conversions on autopilot.



