Back to blog
A/B Testing10 min read

How to A/B Test: The Complete Guide (With Real Examples)

M

Michał Pogoda-Rosikoń

Founder · 2026-04-12

How to A/B Test: The Complete Guide (With Real Examples)

I'll explain A/B testing the way I wish someone had explained it to me

Last year I changed one headline on a SaaS landing page. The original said "Build better products with AI." The new one said "Ship 10x faster -- your competitors already are." Signups went up 34%. One line of copy, five extra words, a third more revenue.

That wasn't a lucky guess. It was an A/B test. And before I ran it, I would have bet money the original was better. My gut was wrong. The data wasn't.

Google runs over 10,000 experiments per year on Search alone. Booking.com runs 1,000+ simultaneous tests at any given time. They don't do this because it's trendy -- they do it because guessing costs money.

Here's exactly how to run your first A/B test, step by step.

What is A/B testing?

A/B testing (also called split testing) is a controlled experiment where you compare two or more versions of something to see which performs better. You show version A to one group and version B to another, measure a specific outcome -- clicks, signups, purchases -- and let the data decide.

Think of it like a blind taste test. You don't ask which soda label looks nicer -- you hand people two unmarked cups and see which they drink more of. A/B testing does the same thing for your website, emails, or ads.

The key word is controlled. You change exactly one thing. Everything else stays identical. That way, any difference in results can only be caused by the change you made.

Concrete example: your SaaS landing page headline reads "The all-in-one platform for modern teams." You suspect "Stop drowning in 6 different tools" might convert better. You split traffic 50/50, run both simultaneously, and compare signup rates. That's it.

Step 1: Pick ONE thing to test

Start with whatever your visitors see first and has the biggest conversion impact. For most sites, the priority is:

  1. Headline -- can swing conversion by 20-50%. This is where the biggest wins live.
  2. Call to action -- "Start free trial" vs "See it in action" vs "Get started in 2 minutes."
  3. Hero section layout -- everything above the fold.
  4. Pricing presentation -- monthly vs annual, anchoring, visibility.
  5. Social proof -- testimonials, logos, case studies.

Don't test footer font sizes or subtle button color changes. If a change only improves conversion by 0.1%, you'd need millions of visitors to detect it. Start bold.

Good first tests:

  • "All-in-one project management" vs "Stop losing tasks across 6 different apps"
  • "Sign up free" vs "See your dashboard in 60 seconds"
  • Long-form landing page vs short hero + single CTA
  • Pricing visible on landing page vs gated behind "View plans"

Step 2: Write your variants

You can test more than two options (A/B/n testing). More variants need more traffic, but testing 3-4 headlines at once is more efficient than running them sequentially.

Say your current headline is:

A (Control): "The easiest way to manage your team"

Your variants:

B (Pain-point): "Stop wasting 5 hours a week on status meetings"

C (Provocative): "Your team is drowning in Slack messages. Fix it."

D (Aspirational): "Team management that doesn't feel like management"

Notice each takes a fundamentally different angle. Testing tiny wording tweaks ("easy" vs "simple") rarely moves the needle. Different value propositions do.

Write down your hypothesis for each. You won't always be right -- 60% of A/B tests deliver under 20% lift, and Kohavi et al. found only about 12-15% of experiments produce a statistically significant winner. But hypotheses keep your program focused and help you learn from every test.

Step 3: Calculate your sample size

This is the step most people skip -- and it causes the most wasted time.

Before launching, answer: how many visitors do I need to trust the results? It depends on three inputs:

  1. Baseline conversion rate. Average ecommerce: ~2.5-3%. SaaS landing pages: ~3-5%. If you don't know yours, measure it for two weeks first.

  2. Minimum detectable effect (MDE). A 20% relative improvement (e.g., 3% to 3.6%) is a common threshold.

  3. Statistical power. Industry standard is 80% -- accepting a 20% chance of missing a real winner.

Rough benchmark: detecting a 20% relative lift on a 5% baseline requires about 5,000 visitors per variation. For a 10% lift, you need 20,000+.

Use this calculator for your exact number:

Interactive tool

Sample Size Calculator

3%
20%
Visitors per variant
10,948
Total visitors needed
21,896
Est. duration (500 visitors/day)
22 days

Based on a two-tailed z-test with 80% statistical power. Adjust the sliders to match your site.

Evan Miller's sample size calculator is the gold standard for double-checking. The formula uses power analysis for two-proportion z-tests -- in plain English: "How many coin flips to tell a fair coin from a slightly biased one?"

Write down your required sample size. Commit to it. This is your finish line.

Step 4: Set up and launch

Traffic split: 50/50. Some guides suggest 90/10 to "limit risk," but unequal splits dramatically increase time to significance. Even splits are almost always better.

Sticky randomization. Each visitor gets randomly assigned to a variant and sees it every return visit -- via cookie or user ID hash. With SplitMonk, this happens at the edge before your page loads, so there's zero flicker.

Full business cycles. Start on Monday. Run for full weeks. B2B traffic on Tuesday looks nothing like Saturday. At least 2 full weeks captures these cycles.

QA before launch. 52% of businesses don't QA experiments before going live. Check every variant on desktop and mobile. Verify tracking fires. Five minutes of QA saves two weeks of garbage data.

Step 5: Wait (seriously, just wait)

Here's what happens after launch -- I've seen it hundreds of times:

  • Day 1-2: One variant "crushes it" at +40%. This means nothing.
  • Day 3-5: Results flip. Also means nothing.
  • Day 7: Stabilizing, but still below your sample size.
  • Day 10-14: Confidence interval tightening. Getting close.
  • Day 14-28: Now you can look.

The temptation to peek is overwhelming. Resist it. Peeking inflates your false positive rate from 5% to 20-30%. Every time you check and think "it's significant, I'll stop," you're running multiple tests and cherry-picking the best-looking moment.

Kohavi, Tang, and Xu documented this extensively in "Trustworthy Online Controlled Experiments." Peeking is the single most common reason A/B tests produce false winners.

Minimum runtime: 2 weeks. Ideal: 4 weeks. Even if you hit your sample size in 5 days, keep running to capture weekly seasonality.

Step 6: Read the results

P-value in plain English: "If both variants were truly identical, what's the probability I'd see results this extreme?" A p-value of 0.03 means a 3% chance the difference is noise. Convention: p < 0.05 is statistically significant.

Confidence intervals matter more. A "significant" result with a confidence interval of [+0.5%, +8%] is actionable. One with [-2%, +15%] isn't -- the true effect could go either way.

Declare a winner when all four are true:

  1. You reached your pre-calculated sample size.
  2. The test ran for at least 2 full weeks.
  3. P-value is below 0.05.
  4. The confidence interval doesn't cross zero.

If you don't hit significance, that's still a result. The difference is too small to matter. Pick your preferred version and move to a bolder test.

Common mistakes that ruin A/B tests

1. Testing too many things at once. Changing headline, CTA, and hero image simultaneously means you don't know which caused the result. That's not an A/B test -- it's a coin flip between two different pages. Multivariate testing exists but requires 10x more traffic.

2. Stopping at the first "significant" result. Pre-commit to a sample size. Don't touch the test until you reach it. If you can't resist peeking, use a tool with auto-stopping.

3. Ignoring segments. No overall winner? Variant B might be crushing it on mobile while losing on desktop. Always check device type, traffic source, and new vs returning visitors.

4. No hypothesis. "Let's test 10 random headlines" isn't a program -- it's a lottery. Read support tickets. Watch session recordings. Then form a hypothesis about why a change should work. Wrong hypotheses still teach you something.

5. Ignoring external factors. Product launches, press mentions, holidays -- all distort results. Always run control and variant simultaneously (never sequentially), and run for full business cycles.

What to test next

Here's the priority framework I use:

  1. Impact x Confidence. Score each idea by potential impact (funnel position, visitor volume) and your confidence it'll move the needle (based on research). Test highest scores first.
  2. Move down the funnel. Headline, then CTA, then below-fold content, then pricing, then checkout.
  3. Build on winners. Pain-point headline won? Test different pain points next. Double down on what works.
  4. Re-test periodically. Audiences change. What won 6 months ago might not win today.

The companies that win at A/B testing run continuously, learn from every experiment, and build compounding knowledge about their audience. Start with one bold test. Calculate your sample size. Launch, wait, read honestly. That's how to A/B test. Everything else is refinement.


Sources: Kohavi, Tang & Xu, "Trustworthy Online Controlled Experiments" | Evan Miller's Sample Size Calculator | VWO A/B Testing Statistics | Convert A/B Testing Stats

Michał Pogoda-Rosikoń

Michał Pogoda-Rosikoń

Founder

Founder of SplitMonk and bards.ai. Data scientist from Wroclaw University of Technology, specializing in NLP and machine learning. Building AI-powered tools that optimize conversions on autopilot.