I've watched hundreds of A/B tests fail
Over the past year building SplitMonk, I've had a front-row seat to how teams run experiments. I've looked at the raw data from hundreds of tests across our platform. The pattern is depressingly consistent: about 1 in 8 produces a real, actionable winner. The rest are noise.
Most people blame low traffic. "We just don't have enough visitors." That's sometimes true, but it's rarely the real problem. The real problem is methodology. Here are the five patterns I keep seeing.
1. You're stopping tests too early
This is the one that kills me. You launch a test on Monday. By Wednesday you've got 800 visitors per variant and the dashboard shows +15% with 92% confidence. You screenshot it, share it in Slack, and ship the winner.
Two weeks later, if you'd kept it running, that +15% would be sitting at -2%.
I've seen this exact scenario play out dozens of times. Early results are wildly unreliable. Here's why: with small samples, random fluctuations look like real effects. Your 800-visitor "winner" is basically a coin flip that happened to land on heads a few extra times.
The math is brutal. If you peek at results daily over a 14-day test and stop whenever you see p < 0.05, your actual false positive rate isn't 5% -- it's closer to 30%. You're six times more likely to be wrong than you think.
-- This is roughly how we check for premature stopping in SplitMonk.
-- We flag any experiment that was stopped before reaching
-- its pre-calculated required sample size.
SELECT
e.id,
e.name,
e.required_sample_size,
SUM(v.visitors) AS actual_visitors,
ROUND(SUM(v.visitors) * 100.0 / e.required_sample_size, 1) AS pct_of_required
FROM experiments e
JOIN variants v ON v.experiment_id = e.id
WHERE e.status = 'completed'
GROUP BY e.id
HAVING pct_of_required < 80;
Fix: Calculate your required sample size before you start. Then don't touch the results until you hit it. We built auto-stopping into SplitMonk specifically because humans can't resist peeking.
2. You're testing things that are too small to measure
Changing a button from blue to slightly-different-blue. Swapping "Submit" for "Send." Moving a CTA 20 pixels to the left.
If a change only improves conversion by 0.1%, you need roughly 3.8 million visitors per variant to detect it at 80% power. Most sites get that in... never.
Fix: Test bold changes. Rewrite the entire headline. Change the value proposition. Go from "Start your free trial" to "See your results in 2 minutes." The bigger the swing, the faster you'll know.
3. The multiple testing problem is eating your results alive
You test 8 headline variants at a 5% significance level. The probability that at least one looks significant by pure chance?
1 - (0.95)^8 = 33.7%
One in three times, you'll find a "winner" that's pure noise. I've seen teams run 10+ variants and celebrate the top performer without any correction. It's basically a random number generator with extra steps.
Fix: Limit yourself to 2-3 variants. If you must test more, apply a Bonferroni correction (divide your alpha by the number of variants) or use a method like Holm-Bonferroni. In SplitMonk, we handle this automatically in our significance engine.
4. Sample ratio mismatch is silently breaking your tests
Before you trust any result, check that traffic actually split correctly. We run a chi-squared test on every experiment:
function checkSRM(visitors: number[], expectedRatios: number[]): boolean {
const total = visitors.reduce((a, b) => a + b, 0);
const expected = expectedRatios.map(r => r * total);
const chiSquared = visitors.reduce(
(sum, obs, i) => sum + Math.pow(obs - expected[i], 2) / expected[i],
0
);
const pValue = 1 - chi2cdf(chiSquared, visitors.length - 1);
return pValue > 0.001; // true = ratio looks fine
}
A 50/50 test that ends up 53/47 has a bug. Common causes: bot traffic hitting one variant harder, caching layers serving stale assignments, or redirect-based tests losing users mid-flight.
Fix: Check SRM before you look at conversion data. If the split is off, your results are garbage regardless of what the p-value says.
5. Novelty effects are fooling you
You redesign the pricing page and see a 20% lift in week one. By week three it's down to 3%. That's novelty -- returning visitors engaging with something new because it's new, not because it's better.
Fix: Run tests for at least two full weeks. Segment by new vs. returning visitors. If the effect lives entirely in returning visitors and fades over time, you're measuring curiosity, not improvement.
The checklist I use before calling any test
- Did we calculate the required sample size before starting?
- Did we actually reach that sample size?
- Does the SRM check pass?
- Did we run for at least one full business cycle (7 days minimum)?
- Did we apply correction for multiple variants?
- Is the effect consistent across segments (new vs. returning, mobile vs. desktop)?
- Is the lift practically meaningful, not just statistically significant?
If any answer is "no," I don't trust the result. Neither should you.

Michał Pogoda-Rosikoń
Founder
Founder of SplitMonk and bards.ai. Data scientist from Wroclaw University of Technology, specializing in NLP and machine learning. Building AI-powered tools that optimize conversions on autopilot.



