📄

How to Calculate A/B Test Sample Size

Sample size is decided before a test runs, not after you look at the results and decide they "feel" done. Get it wrong and you either run underpowered tests that miss real effects or waste months waiting for a sample you never needed.

How Many Visitors Does an A/B Test Need?

Enough to detect your target minimum effect at your desired confidence level given your current baseline conversion rate — there is no single universal number, and any test claiming "1,000 visitors is enough" is ignoring the underlying math.

The relationship is non-linear: halving the effect size you want to detect roughly quadruples the required sample. This is the single most common reason conversion tests fail to reach statistical significance — teams underestimate how much traffic a modest, realistic lift actually requires to detect reliably.

What Inputs Drive a Sample Size Calculation?

Four inputs determine sample size: your current baseline conversion rate, the minimum effect size worth detecting, your desired statistical significance level, and your desired statistical power (typically 80%).

Power is the probability of detecting a real effect if one exists — the "false negative" counterpart to significance's "false positive" control. Most calculators default to 80% power and 95% significance; raising either number increases the required sample, sometimes substantially.

Why Does Your Baseline Conversion Rate Matter So Much?

Lower baseline conversion rates require larger sample sizes to detect the same relative lift, because the absolute number of conversion events — not just visitors — is what drives statistical power.

A page converting at a low single-digit rate needs many more total sessions to accumulate enough conversion events than a page converting at a high rate, even testing the identical relative improvement. This is a big part of why e-commerce conversion benchmarks vary so much across categories — a low-baseline niche store will need proportionally longer tests than a high-converting subscription checkout to reach the same statistical confidence.

What Is Minimum Detectable Effect and Why Does It Matter?

Minimum detectable effect (MDE) is the smallest lift you've decided is worth detecting — setting it too small inflates your required sample size far beyond what's practical; setting it too large risks missing real, valuable improvements.

A useful discipline is to set MDE based on what lift would actually change a business decision. If a 2% lift wouldn't justify the engineering cost of shipping a change, there's no point sizing a test to detect it — size for the smallest effect that would actually be worth acting on.

What Should Low-Traffic Sites Do Instead?

Test changes with larger expected effects, run fewer simultaneous tests so traffic isn't split further, and accept longer test durations — or shift to qualitative signals like session recordings and heatmaps for lower-traffic pages that will never reach a statistically valid sample.

A formal split test is the wrong tool for a page that gets a few hundred visits a month; you will wait years to reach significance for anything but a dramatic effect. In that situation, session recordings and direct user feedback tell you more per hour of effort than an underpowered A/B test ever will.

Should You Plan by Duration or by Sample Size?

Calculate the required sample size first, then divide by your expected daily traffic to estimate duration — and commit to running at least one full business cycle (typically one to two weeks) regardless of how fast you hit the number, to avoid day-of-week bias.

Stopping purely because you hit a sample size on a Tuesday afternoon ignores that weekday and weekend traffic often convert at meaningfully different rates. Combine the sample-size target with a minimum-duration rule, covered in more detail in when to stop an A/B test, to avoid both underpowered and biased results.

How Do You Actually Run the Calculation?

Use any standard two-proportion sample size calculator: enter your current baseline conversion rate, your chosen minimum detectable effect, and your significance and power targets, and it will output the required visitors per variant.

Do this before launching every test, not just the first one — your baseline rate shifts over time (seasonally, after redesigns, after traffic-mix changes), so a sample size calculated a year ago is not automatically valid for a test you're launching today.

Summary

Sample size is driven by your baseline conversion rate and the smallest effect worth detecting — calculate it before the test starts, plan duration around a full business cycle, and use qualitative methods on pages that will never reach a valid sample.

Ready to Add Social Proof to Your Website?

Get started free and increase conversions in minutes.

Get Started Free

Ready to Increase Your Conversions?

Start using NotiProof free today and turn visitors into customers with social proof. No credit card required.

Free forever plan · No credit card required