What the calculator does
The A/B test calculator tells you whether the difference between two versions of an ad is likely to be a real effect or just noise from a small sample. If you have run two versions of a headline, image or landing page and want to know whether one actually pulled ahead, this is the tool that turns your raw numbers into a straight answer: yes, no, or not yet.
An A/B test, in advertising, means showing two versions of the same ad to similar audiences and comparing how each performs, usually on click-through rate (CTR, the share of people who saw the ad and clicked it) or conversion rate. The trouble is that small differences in performance can happen by chance alone, especially when the sample is modest. The calculator applies a standard statistical test to work out how likely that is, so you are not reading meaning into a result that could easily have gone the other way.
What you need before you start
The calculator needs four figures: the number of people who saw or clicked through to each version, and the number of conversions (clicks, sign-ups, sales, whatever you are measuring) each version produced. Gather these from your ad platform's reporting before you open the tool, because the quality of the answer depends entirely on the quality of what goes in.
- Impressions or visitors for version A, the total audience that saw that version
- Conversions for version A, however you have defined a conversion for this test
- Impressions or visitors for version B
- Conversions for version B
Sample size matters more than most advertisers expect. A version that converts at 4% against 3% looks like a clear winner on a spreadsheet, but with only a few hundred visitors each, that gap could easily be chance. The calculator accounts for this, which is why the same percentage gap can come back as significant with one data set and inconclusive with another.
Reading the result
The calculator will return a significance result, usually expressed as a confidence level such as 90% or 95%. In plain terms, a 95% confidence level means that if you ran this exact test many times with new audiences, you would expect to see a difference this large or larger by chance alone only about one time in twenty. It is a statement about how likely the pattern is to be real, not a guarantee.
A result that clears the significance threshold tells you the difference is unlikely to be chance. It does not tell you the difference is large, that it will hold up next month, or that the same ad will perform the same way with a different audience or on a different platform. Statistical significance and practical significance are separate questions, and an advertiser working to a tight budget cares about both.
What "not significant" means
A result that falls short of significance is not proof the two versions perform identically. It usually means the sample is not yet large enough to tell the difference from noise, so the honest next step is to keep the test running.
Mistakes that quietly skew the numbers
A few habits creep into most A/B testing and they distort the read without the advertiser noticing.
- Stopping the test the moment it looks significant. Checking results daily and stopping as soon as one version edges ahead inflates the chance of a false positive, because you are effectively giving yourself many attempts to catch a lucky run.
- Testing more than one change at once. If version B has a different headline, image and call to action, a win tells you B beat A, not which change did the work.
- Running the test across mismatched time periods. Seasonal effects, day-of-week patterns and one-off events can all move CTR independently of the ad itself, so a test split across very different periods is comparing two things at once.
- Treating a small percentage gap as meaningful without checking the sample size. A jump from 2% to 2.5% conversion sounds like a 25% improvement, but with a few hundred impressions per version it may sit well inside the range chance would produce anyway.
What the calculator cannot tell you
The calculator works from the numbers you give it, so it can only be as reliable as those inputs. Wrong figures, a conversion definition that shifted mid-test, or a sample that is technically large enough but drawn from an unrepresentative audience will all produce a result that looks precise and is not. Treat the output as a steer to weigh against what you know about the test.
It also cannot explain why one version won. A stronger headline might have worked because of the wording, the length, or simply because it appeared first in the rotation and caught an early burst of engaged users. Understanding the mechanism behind a result is a separate piece of work from establishing that the result is real, and the calculator only does the second one.
Using the result
Once you have a significant result, the practical move is usually to shift budget or ad space toward the winning version and retire the other, then keep an eye on performance over the following weeks in case the gap narrows. Where the result is inconclusive, the more useful response is often to let the test run longer or to widen the difference being tested.
For advertisers working out where that reallocated budget should go next, the Ad Budget Calculator and its accompanying guide cover how to split spend across channels once you know which version of an ad is doing the work. If any of the terms here, CTR, conversion rate, confidence level, need a plainer definition, the Advertising Glossary covers them in the same terms used across the site.