A/B Test Significance Calculator
Two ads, two landing pages or two headlines — and one looks better. Enter the visitors and conversions for each and the calculator tells you whether the gap is a real difference or the kind of swing random traffic produces on its own, then how much data you would need to know for sure.
A/B Test Significance Calculator
A — control
B — variant
+2.00% points vs A.
How much higher (or lower) B's rate is than A's.
z = 1.883. Significant at 95% when p < 0.05.
Absolute percentage points. The range includes zero, so “no difference” is still plausible.
Not significant yet at 95%
The gap between A and B is within what random variation could produce at this sample size. That is not proof the variants perform the same — only that the data cannot tell them apart yet. If the observed lift is real, detecting it reliably (80% power) takes about 2,213 visitors per variant.
Two-proportion z-test: z = (pB − pA) ÷ √(p̂(1 − p̂)(1/nA + 1/nB)), where p̂ is the pooled rate. The confidence interval uses the unpooled standard error.
Formula
z = (p_B − p_A) ÷ √( p̂(1 − p̂) × (1/n_A + 1/n_B) )
p_A and p_B are each variant's conversion rate (conversions ÷ visitors), n is visitors per variant, and p̂ is the pooled rate — all conversions ÷ all visitors — which is what the rates would be if A and B were the same. The two-sided p-value is the chance of a |z| at least that large under that assumption. The confidence interval on B − A uses each variant's own rate instead of the pooled one.
Worked example
Control A converts 50 of 1,000 visitors (5.0%); variant B converts 70 of 1,000 (7.0%) — a 40% relative lift. The pooled rate is 6.0%, the standard error 1.06 points, so z = 2.0 ÷ 1.06 = 1.88 and the two-sided p-value is 0.060. That clears 90% confidence but not 95%: at 95% the verdict is “not significant yet”. To reliably detect a 5% → 6% move instead (a 20% lift) at 95% confidence and 80% power, plan on 8,158 visitors per variant.
What this tells you
Every split test produces a winner on the dashboard, because two numbers are almost never identical. The question this calculator answers is narrower and more useful: if the two variants were really the same, how often would random traffic alone hand you a gap this big? When the answer is “rarely”, the result is significant. When it is “fairly often”, the honest reading is that you do not know yet — which is different from knowing there is no difference.
To read a test with it:
- Count each variant’s visitors and conversions. Use the unit the test randomised — people or sessions for a landing page, impressions for a click-through test — and count converting visitors, not total orders, so every visitor is a yes or a no.
- Pick the confidence level before you look. 95% is the usual default; 90% accepts more false winners in exchange for faster calls, 99% the reverse. Choosing it after seeing the p-value defeats the point.
- Read the interval, not just the verdict. The confidence interval shows the range of lifts the data supports. A significant result whose interval runs from +0.1 to +4 points is a different decision from one that runs from +3 to +4.
Plan the sample before the test starts — the Sample size tab does that from your baseline rate and the smallest lift worth detecting. Our ad creative testing framework covers how to structure the variants so a significant result actually tells you something, and the A/B testing glossary entry has the vocabulary.
Common mistakes
Checking every day and stopping the first time p dips under 0.05. Repeated looks inflate the false-positive rate well past 5%; decide the sample size up front and read the result once it is reached.
Testing five variants against control and celebrating the one that hits significance. With several comparisons, one is likely to clear 95% by chance — tighten the threshold (for example, divide 0.05 by the number of comparisons) or retest the winner.
Treating “not significant” as “no difference”. A small test can miss a real lift; check the confidence interval and the sample-size tab before concluding the change did nothing.
Counting orders instead of converting visitors. The test assumes each visitor converts or does not; repeat purchases from one person make the variance look smaller than it is.
When to use it
- Creative A/B on Meta: comparing two ads' link clicks per impression or purchases per click. Meta's own A/B test tool splits audiences for you — use this calculator to sanity-check its readout, or to read an informal test where both ads ran in the same ad set, keeping in mind delivery was not randomised there.
- Landing-page split test: sending paid traffic to two page versions and comparing form fills or purchases per visitor. This is where most tests are underpowered, so run the sample-size tab before launch.
- Ad-copy test on Google: two responsive search ad headlines or descriptions measured on CTR, or two campaign experiment arms measured on conversion rate. Google Ads experiments report their own significance; this is a quick, transparent second read on the same counts.
FAQ
Why shouldn't I stop a test the moment it hits significance?
Because the p-value is only valid if you look once, at a sample size chosen in advance. Early in a test the rates swing widely, and if you check daily and stop the first time p falls below 0.05, you will declare winners between identical variants far more often than 5% of the time — the peeking problem. Fix the sample size (or a run length covering full weeks) before launch, let the test reach it, and then read the result. If you need to look early, use a method designed for it, such as a sequential test with adjusted thresholds.
Should I use a one-tailed or two-tailed test?
This calculator uses two-tailed, which asks whether B is different from A in either direction. A one-tailed test only asks whether B is better, which halves the p-value for the same data and so reaches significance sooner — but it cannot flag a variant that is worse, and switching to one-tailed after seeing which way the result went is a form of cheating. Unless you decided in advance that a worse B and an equal B lead to the same action, stay two-tailed.
Should I test CTR or conversion rate?
Test the metric the decision depends on. For a CTR test, enter impressions and clicks; for a conversion-rate test, enter clicks or visitors and conversions. CTR tests reach significance faster because clicks are far more frequent than purchases, which makes them good for screening creative — but a higher CTR does not guarantee more sales, since an ad can win clicks from people who never buy. Use CTR to narrow the field and conversion rate, or cost per acquisition, to pick the winner.
What does a p-value of 0.05 actually mean?
It means that if A and B truly performed the same, a gap at least as large as the one you observed would turn up about 5% of the time from random variation alone. It is not the probability that B is better, and it is not a 95% chance the lift is real. It also says nothing about size: with enough traffic, a lift too small to be worth anything can still be highly significant, which is why the calculator shows the lift and its confidence interval alongside the p-value.
How many visitors do I need for an A/B test?
It depends on your baseline conversion rate and the smallest lift you care about, and it grows fast as either gets smaller. At a 5% baseline, detecting a 20% relative lift (5% → 6%) at 95% confidence and 80% power takes about 8,158 visitors per variant; halving the lift to 10% roughly quadruples that. Enter your own numbers in the Sample size tab. If the answer is more traffic than you can buy, test bigger differences — a new concept rather than a new button colour.
Related calculators
Let Soku run the math — and the ads
Soku AI tracks ROAS, CAC, and MER live across Meta, Google, and TikTok, then optimizes spend toward your targets automatically.
Start free