How Much Traffic Do You Need Before an A/B Test Means Anything?
Most A/B tests that declare a winner never had enough traffic to detect one. Here is the arithmetic, a lookup table, and what to do when the numbers say you cannot test.

A site converting at 3% needs roughly 52,000 visitors per variant — about 104,000 in total — to reliably detect a 10% relative improvement at 95% confidence and 80% power. Most businesses running A/B tests do not have that traffic in a reasonable timeframe, which is why so many “winning” tests fail to replicate.
This is the least popular finding in conversion optimisation and the most useful one. Working out the number before you build the test tells you whether the test is worth building at all.
The formula
For a standard two-variant test at 95% confidence and 80% power:
n per variant ≈ 16 × p × (1 − p) ÷ d²
Where p is your current conversion rate as a decimal and d is the absolute improvement you want to detect, also as a decimal.
Worked example. You convert at 3% and want to detect a 10% relative lift — a move to 3.3%, so an absolute difference of 0.003:
n = 16 × 0.03 × 0.97 ÷ 0.003²
n = 0.4656 ÷ 0.000009
n ≈ 51,733 per variant
The 16 is a rounded constant encoding 95% confidence and 80% power — the exact value is about 15.7. Being rounded up from 15.7 while ignoring the slight difference in variance between the two groups, the approximation lands between 0.3% and 7% below an exact calculation, with the gap widening as the effect size grows. Use it to decide whether a test is feasible; confirm the final number in a proper calculator before you commit to an end date.
Two things fall straight out of the arithmetic, and both are counter-intuitive:
- The requirement scales with the square of the effect. Halving the effect you want to detect quadruples the traffic you need. Detecting a 5% lift instead of a 10% one costs four times the sample.
- Lower conversion rates are far more expensive to test. A site at 1% needs roughly three times the sample of a site at 5% to detect the same relative change.
Lookup table: visitors needed per variant
At 95% confidence and 80% power. Double these figures for the total across both variants.
| Current conversion rate | To detect +5% relative | +10% relative | +20% relative |
|---|---|---|---|
| 1% | 633,600 | 158,400 | 39,600 |
| 2% | 313,600 | 78,400 | 19,600 |
| 3% | 206,900 | 51,700 | 12,900 |
| 5% | 121,600 | 30,400 | 7,600 |
| 10% | 57,600 | 14,400 | 3,600 |
Find your row. If your site gets 15,000 visitors a month and converts at 2%, detecting a 10% lift requires 78,400 per variant — 156,800 in total, or about ten and a half months on one test. That test does not exist. It is a plan to run one experiment a year and probably get it wrong anyway.
This is not an argument against testing. It is an argument for testing the right things, which is the rest of this article.
Why underpowered tests produce winners anyway
If the traffic is not there, why does the tool keep reporting significance?
Three mechanisms, all of them ordinary and all of them wrong:
Peeking. Checking the result daily and stopping when it crosses 95% is not a 95% confidence test. Every look is another chance to catch a random fluctuation at its peak. Repeatedly peeking at a test with no real effect can produce a “significant” result well over a third of the time. Fixed-horizon tests must be given a fixed horizon — sample size and end date set in advance, results read once.
Small samples produce large swings. Early in a test, one or two conversions move the rate visibly. A variant sitting at +40% after 300 visitors is noise with a confident-looking chart attached.
Winner’s curse. Even a correctly significant result from an underpowered test overstates the effect. Only unusually large observed differences can clear significance at small samples, so the ones that clear it are systematically the exaggerated ones. This is the specific reason a “23% lift” so often turns into nothing when it ships.
What to do when you do not have the traffic
Most businesses are in this position. The response is not to abandon evidence — it is to change the kind of evidence you gather.
Test bigger things. A button colour might move conversion 2%. A different offer, a restructured page, removing a required form field, or replacing flat product photography with 3D imagery might move it 30%. The lookup table shows a 20%+ effect is testable at a fraction of the sample. Stop testing tweaks you could never detect and start testing changes large enough to show up.
Move the test upstream. Testing the checkout of a site with 400 checkout starts a month is hopeless. Testing the landing page that 40,000 people see is feasible — and if those are paid landing pages, check whether they should be indexed at all before you start experimenting on them. Run experiments at the widest point of the funnel where the traffic is.
Use a proxy metric with more volume. Add-to-cart, form-start, scroll-to-pricing and demo-page-view all occur far more often than purchases. Higher event volume means smaller sample requirements. The trade-off is real — the proxy has to be genuinely predictive of revenue, and you should verify that relationship before relying on it. If the answer feeds a bidding decision, tie it back to your break-even ROAS rather than to the proxy alone.
Accept a lower confidence level, deliberately. 95% is a scientific convention, not a business requirement. For a low-risk, low-cost change, deciding at 85% confidence is a perfectly rational trade-off between speed and certainty — as long as it is a stated decision rather than an accident, and you do not describe the result as proven.
Do qualitative research instead. Session recordings, on-page exit surveys, five user tests and a proper form-analytics review will find the actual friction faster than a year-long test on a site with no traffic. Diagnosis first. Most low-traffic sites do not have a testing problem — they have an unexamined problem that testing was going to take a year to locate.
How long should a test run?
Run for a minimum of two full weeks, in whole seven-day increments, regardless of what the sample size calculation says.
The sample size tells you the minimum. The calendar imposes its own floor for reasons the arithmetic does not capture:
- Behaviour varies by day of week. A test that runs Monday to Friday has measured weekday visitors, not your audience.
- Purchase cycles lag. A visitor who converts on day nine was assigned on day two, and cutting at day seven credits their visit to nobody.
- One anomalous day — an outage, a press mention, a competitor’s promotion — distorts a short test far more than a long one.
Ceiling as well as floor: past about four to six weeks, cookie deletion, cross-device switching and seasonal drift degrade the assignment quality enough to undermine the comparison. If your calculated sample needs more than six weeks, that is the calculator telling you to pick a bigger change to test.
Frequently asked questions
How many conversions do I need for an A/B test?
As a rough working floor, around 250–400 conversions per variant, in addition to meeting the visitor requirement. Below that, the confidence interval on the conversion rate is too wide for the comparison to say much.
Can I A/B test with low traffic?
Yes, but only for large effects. Use the lookup table in reverse: with the traffic you have, calculate the smallest effect you could detect, and only test changes plausibly capable of producing it. If that number exceeds about 20% relative, switch to qualitative research and ship well-reasoned changes without testing them.
What does 95% statistical significance actually mean?
It means that if there were genuinely no difference between the variants, you would see a result this extreme or more extreme less than 5% of the time. It does not mean there is a 95% chance the variant is better, and it says nothing about how large the improvement is.
Why did my A/B test winner not work after launch?
The three usual causes are peeking (stopping the test at a favourable moment), the winner’s curse (underpowered tests only detect exaggerated effects), and novelty — where returning visitors respond to the change simply because it is new, an effect that decays within weeks.
Should I run A/A tests?
Once, when you first set up a testing tool. An A/A test compares two identical variants and should show no significant difference. If it reports a winner, your assignment, tracking or analysis is broken, and every test you run afterwards inherits that fault.
ThynqAi runs conversion programmes that start with diagnosis rather than a testing backlog. If you want to know which experiments your traffic can actually support, ask us for an audit.