Sample size calculator.

How many visitors your A/B test needs before the result means anything, and how long that takes at your traffic.

What the control converts at today. Enter 3 for 3%.

The smallest relative improvement worth detecting. 20% relative on a 3% baseline means 3.6%.

How often you are willing to call a winner that is not real. 95% is the convention.

Your chance of spotting a real effect of this size. 80% is the convention.

Optional. Total across both variants, used only to estimate how long this runs.

Per variant
13,914
visitors in the control, and the same again in the variant
Total sample
27,828
to detect 3.00% moving to 3.60%
Time to run
-
Add your daily traffic to estimate this.

Assumes a two-sided test comparing two independent proportions, a 50/50 split between control and variant, one variant against one control, and a single analysis at the end rather than peeking as results come in. It uses the standard normal approximation, stated in full on this page. Everything is computed in your browser.

How many visitors does an A/B test need?

It depends on your baseline conversion rate and the smallest improvement worth detecting. At a 3% baseline, detecting a 20% relative lift with 95% significance and 80% power takes about 13,900 visitors per variant, so roughly 27,800 in total. Smaller effects cost far more traffic: halving the effect roughly quadruples the sample.

Which test does this assume?

A two-sided z-test comparing two independent proportions, using the standard normal approximation. Per variant:

n = ( z1-α/2× √(2 × p̄ × (1 - p̄)) + z1-β × √(p1(1 - p1) + p2(1 - p2)) )2 ÷ (p2 - p1)2

p1 is the baseline rate, p2is the rate you want to be able to detect, and p̄ is the average of the two. The first term uses the pooled variance under the null hypothesis, the second the unpooled variance under the alternative. That pairing is what published sample size calculators use, and it is slightly more conservative than the familiar shortcut of 16 × p × (1 - p) ÷ delta².

The z values are exact standard-normal quantiles for the significance and power levels offered, so nothing is rounded there. The approximation lives in the normal model itself. It holds well at ordinary web conversion rates and stops being trustworthy when a variant would produce only a handful of conversions in total, which is why the tool warns you below about ten expected conversions per arm.

Four things it assumes and you should check: the test is two-sided, so you are asking whether the variant differs rather than whether it is better; traffic is split evenly between the two arms; there is one control and one variant; and you read the result once, at the end, rather than watching it daily and stopping when it looks good.

What each input actually costs you.

Every knob on the calculator trades traffic against certainty. This is what each one buys, holding everything else at a 3% baseline, 95% significance, 80% power.

ChangeVisitors per variantWhat it buys you
20% relative effect (the default)13,914The reference point
10% relative effect53,211Catches smaller wins, at nearly four times the traffic
30% relative effect6,455Cheap, but only detects large changes
99% significance instead of 95%20,704Fewer false winners, at about 1.5x the traffic
90% power instead of 80%18,626Misses a real effect 1 time in 10 rather than 1 in 5
6% baseline instead of 3%6,719Higher baselines need less traffic for the same relative lift

Figures computed with the formula above, rounded. Reproduce any row by entering it into the calculator.

You still need the baseline.

Every number this calculator produces depends on the rate your control converts at today. Guess that wrong and the sample size is wrong with it, usually in the direction that makes the test end too early.

Take the baseline from the same funnel step, the same traffic mix, and a period long enough to cover a full weekly cycle. If your rate swings between 2% and 5% depending on the week, size the test on the low end and expect it to run longer than the estimate.

Mrkr custom events list with counts and conversion rates.
Custom events and their conversion rates in Mrkr, the numbers a baseline comes from. Open it in the live demo.

Where tests go wrong.

Peeking and stopping early

Stopping at the first significant reading pushes your real false positive rate past 20%.

Running less than a week

Tuesday traffic is not Sunday traffic. Run whole weeks so both arms see the full cycle.

Sizing for a lift you will not get

Most copy changes move conversion a few percent, not twenty. Size for the realistic ceiling.

Counting bots in both arms

Crawlers land in control and variant and convert in neither. On one live Mrkr site they ran at 4.1x human pageviews in September 2026.

Testing many variants at once

Four variants is four comparisons, and the odds one wins by luck climb fast.

Treating a null result as failure

A powered test that finds nothing has told you the change is not worth shipping.

Where your baseline comes from

Sizing a test starts with knowing the rate you have now, per step and per source. The Mrkr demo shows funnels, goals and events on a fully populated dashboard, with no signup.

Open the live demo
A Mrkr funnel showing conversion and drop-off at each step.

What if you do not have the traffic?

If the calculator says 14,000 per variant and you get 3,000 visitors a month, an A/B test on that page is not a tool you can use. Three honest alternatives.

Test earlier in the funnel

A step with more traffic and a higher base rate reaches significance far sooner.

Test bigger changes

Sample needed collapses as the effect grows. Two offers, not two headlines.

Use qualitative evidence

Twenty session replays of an abandoned form beat a test you cannot power.

Questions, answered.

Keep reading

Your first visitor is already here.

Drop in the script and watch them land. It takes about a minute.