What is an A/B split test?

An A/B split test compares two versions of a page or popup, A (the control) and B (the variation), by splitting traffic randomly between them at the same time and measuring one metric, so a difference can be attributed to the change. "Split test" and "A/B test" mean the same; split-URL tests put each version on its own URL.

Contents6 min read
  1. How do you run an A/B split test?
  2. How do you read A/B split test results?
    1. A gap can look big and still not be reliable
    2. An effect can reverse when you compare like with like
    3. What a reliable gap looks like
  3. How is A/B testing different from multivariate and split-URL tests?
  4. What should you A/B test on a popup?
  5. A/B split test examples by industry
  6. A/B split test tools
  7. Does A/B testing affect SEO?
  8. Related Blog Post
  9. A/B Split Test Related Terms
Shoppers splitting between two doors, one lit blue

It is also called bucket testing. The change may be a headline, a button, or a complete redesign of the webpage.

Most guides stop at "test your button colour". The harder part is reading the result: across the popups we power, between January 2024 and September 2026, we measured gaps that look big but could be chance, and one that reverses when you compare like with like.

How do you run an A/B split test?

The A/B testing process is the same for a page, an email or a popup:

I decide the sample size and the stop date before a test starts, and I don't read the result early: the significance calculation assumes the sample size was fixed in advance, as Evan Miller showed in How Not To Run an A/B Test (2010).

A/B Split Test Process Representation

How do you read A/B split test results?

We compared popups that ask for an email or phone number by leads per popup display, median campaign (the 2026 popup conversion benchmark explains the method). These are comparisons between campaigns, not A/B tests, which makes them a useful warning.

A gap can look big and still not be reliable

CTA buttons that promise value ("Get", "Claim", "Unlock") look stronger than "Subscribe" / "Join" buttons (1.06% vs 0.77%, English popups, median campaign), but the gap is not statistically reliable. Exit intent tells the same story: popups on the most sensitive exit-intent setting converted 1.10% vs 0.78% without exit intent (median campaign), a gap that could be chance.

When a gap looks this big but isn't reliable, I treat it as a hypothesis for the next test, not a rule.

An effect can reverse when you compare like with like

Info
By the Numbers: Adding a name field alongside the email shows no reliable effect: the overall gap in its favour (1.25% vs 0.88% for email only) reverses among campaigns whose form rarely changed (0.74% vs 0.79%). Popupsmart 2026 benchmark, January 2024 to September 2026.

Because our benchmark compares different campaigns, a gap can reflect who uses a setting as well as the setting itself. That is confounding: the name field arrives bundled with whatever else differs between those campaigns. A randomized A/B test removes it. Only the form changes, and chance decides who sees which version, which makes it a randomized controlled trial (RCT) on a website. It's why I test a setting against the current version instead of copying it from campaigns that convert better.

What a reliable gap looks like

First-person buttons ("Send my code") convert 1.18% vs 0.80% for neutral wording (English popups, median campaign). That gap is reliable in our data and a strong candidate for variant B, but it's still a comparison between campaigns: test it on your own traffic first.

How is A/B testing different from multivariate and split-URL tests?

An A/B test changes one thing. Each extra version or combination splits traffic further, so it needs more visitors:

Test type What changes Use it for
A/B (split) test One change, two versions One hypothesis to check
A/B/n test (A/B/C with three versions) One element, three or more versions Several candidates, such as three headlines
Multivariate test Several elements, in combinations High-traffic pages; shows how changes interact
Split-URL test A whole page, each version on its own URL Redesigns; follow the SEO rules below

What should you A/B test on a popup?

A popup has four parts worth testing:

I test these before the button colour: in our data, button colour showed no measurable effect on conversion, though the data are too noisy to rule out modest effects. For set-ups, see A/B testing for popups and how to A/B test Shopify popups.

A/B split test examples by industry

These examples are adapted from Optimizely's A/B testing glossary.

1. A media company might want to increase readership, increase the amount of time readers spend on their site, and amplify their articles with social sharing. To achieve these goals, they might test variations on:

2. A travel company may want to increase the number of successful bookings completed on their website or mobile app, or may want to increase revenue from ancillary purchases. To improve these metrics, they may test variations of:

3. An e-commerce company might want to increase the number of completed checkouts, the average order value, or increase holiday sales. To accomplish this, they may A/B test:

4. A technology company might want to increase the number of high-quality leads for their sales team, increase the number of free trial users, or attract a specific type of buyer. They might test:

A/B split test tools

Does A/B testing affect SEO?

A/B testing does NOT hurt rankings when it's set up the way Google asks: Google's guidance on A/B testing says small changes, such as a button's colour or its call-to-action text, often have little or no impact on a page's search snippet or ranking. Popup tests on the same URL need none of the URL rules below.

If you run a split test with multiple URLs, use rel="canonical" on each variation to prevent Googlebot from getting confused by similar versions of the same page. Google recommends it over a noindex tag.

If you run a split test that redirects the original URL, use 302 (Temporary) Redirects rather than 301s (Permanent) to enable Google to keep the original URL.

Warning
Red Flag: Showing Googlebot one set of URLs and people another is cloaking, which breaks Google's spam policies even during a test. Googlebot generally doesn't support cookies, so a cookie-based test shows it the version for browsers without cookies.

You need to avoid conducting long and unnecessary experiments: Google may treat an unnecessarily long one as an attempt to deceive search engines. When the test ends, remove the test URLs and scripts.

Related Blog Post

Why Emails Go to Spam Instead of Inbox

Best WordPress Popup Plugins: Comparison Guide

Increase Email Open Rates with Catchy Email Subject Lines

Increase Email Open Rate: Email Marketing Guide For Beginners

Email List Building: Proven Methods to Grow Your Email List (for Beginners)

Grow your email list

Email Marketing