"When can I stop this test?" is the question that causes more bad decisions than any other part of A/B testing. The honest answer is decided before the test starts, not while staring at a dashboard hoping the line crosses a threshold. Get the stopping rule right and the rest of the analysis takes care of itself.
How Long Should You Run an A/B Test?
Run it until it reaches the sample size calculated in advance for your baseline conversion rate and minimum detectable effect, and for at least one to two full weekly cycles — commonly two to four weeks for most sites, longer for low-traffic pages.
Duration and sample size are two separate constraints and you need both satisfied. A high-traffic page might hit its required sample in three days, but stopping there risks catching only weekday behavior. Calculate the number using sample-size calculation first, then let the calendar set the floor.
Fixed-Horizon vs Sequential Testing: What's the Difference?
Fixed-horizon testing commits to a sample size and duration up front and checks the result exactly once at the end; sequential testing methods are statistically designed to let you check results as they accumulate without inflating the false-positive rate.
Most standard significance tests (including simple two-proportion z-tests) are fixed-horizon by design — checking them daily and stopping early breaks the guarantee. If you want the ability to stop early legitimately, you need a testing platform or method that explicitly supports sequential analysis, not a standard test checked impatiently.
What Minimum Sample and Duration Do You Need?
The minimum sample is whatever your power calculation returns for your baseline rate, desired minimum detectable effect, and significance/power thresholds — there is no universal number, and claims of "a few hundred conversions" are usually too small for realistic effect sizes.
Smaller expected effects require dramatically larger samples, which is why most tests of subtle design changes need far more traffic than teams initially budget for. Confirm the required significance threshold matters with statistical significance before committing to a test plan.
What Signs Mean You Should Keep Running the Test?
Keep running if you haven't reached the pre-calculated sample size, if you haven't completed at least one full weekly cycle, or if traffic mix or promotional activity during the test window was unusual.
A launch coinciding with a spike from a paid campaign or a seasonal sale skews the visitor mix in ways that won't repeat under normal traffic. Extend the test, or restart it, rather than reporting a result drawn from an atypical week.
What Signs Mean You Can Stop?
Stop once you've reached the pre-committed sample size and duration, whether or not the result is significant — the stopping condition is about reaching the plan, not about the result looking good.
This is the part teams find hardest to accept: a test that hits its planned end date with a flat result is a completed test with a null finding, not an unfinished one. Ending on schedule protects the validity of every other test you run afterward.
What Do You Do If a Test Stays Inconclusive?
Accept the null result, check whether the test was powered for a realistic effect size, and decide whether the change is worth re-testing with a larger sample or simply isn't going to move the metric enough to matter.
An inconclusive test is still informative — it tells you the effect, if real, is smaller than your minimum detectable effect. Feed that back into funnel analysis to decide whether a different stage of the funnel deserves the next test instead.
How Should You Document the Result?
Record the hypothesis, the pre-registered sample size and duration, the primary metric, the final result, and whether it was significant — before any post-hoc segment analysis — so future readers can tell a planned result from an exploratory one.
This record is what keeps a testing program credible over time. Feed validated wins into your conversion dashboard as confirmed changes, and keep exploratory findings clearly labeled as hypotheses for the next test.
Summary
Decide your sample size and duration before launch, run through at least one full weekly cycle, and stop on schedule regardless of what the interim numbers show. The discipline of the stopping rule is what makes the result trustworthy.
