It is also called bucket testing. The change may be a headline, a button, or a complete redesign of the webpage.
Most guides stop at "test your button colour". The harder part is reading the result: across the popups we power, between January 2024 and September 2026, we measured gaps that look big but could be chance, and one that reverses when you compare like with like.
How do you run an A/B split test?
The A/B testing process is the same for a page, an email or a popup:
- Data collection: Pick a page or popup with steady traffic and a conversion rate you want to improve.
- Goal identification: Point out your conversion goal (such as email signups or clicks to product purchases) and one primary metric to judge it by.
- Hypothesis generation: Write down the change and why you think it would be better than the current version.
- Variation creation: Make one change, so a difference can be traced back to it.
- Sample size and duration: Fix the sample size and the end date before launch, using a sample size calculator fed with your baseline rate and the smallest lift worth detecting.
- Conduct experiment: Split traffic randomly between the versions at the same time and run to that sample size.
- Results analysis: Read the primary metric. No statistically significant difference means no reliable difference, not a win.
I decide the sample size and the stop date before a test starts, and I don't read the result early: the significance calculation assumes the sample size was fixed in advance, as Evan Miller showed in How Not To Run an A/B Test (2010).

How do you read A/B split test results?
We compared popups that ask for an email or phone number by leads per popup display, median campaign (the 2026 popup conversion benchmark explains the method). These are comparisons between campaigns, not A/B tests, which makes them a useful warning.
A gap can look big and still not be reliable
CTA buttons that promise value ("Get", "Claim", "Unlock") look stronger than "Subscribe" / "Join" buttons (1.06% vs 0.77%, English popups, median campaign), but the gap is not statistically reliable. Exit intent tells the same story: popups on the most sensitive exit-intent setting converted 1.10% vs 0.78% without exit intent (median campaign), a gap that could be chance.
When a gap looks this big but isn't reliable, I treat it as a hypothesis for the next test, not a rule.
An effect can reverse when you compare like with like
Because our benchmark compares different campaigns, a gap can reflect who uses a setting as well as the setting itself. That is confounding: the name field arrives bundled with whatever else differs between those campaigns. A randomized A/B test removes it. Only the form changes, and chance decides who sees which version, which makes it a randomized controlled trial (RCT) on a website. It's why I test a setting against the current version instead of copying it from campaigns that convert better.
What a reliable gap looks like
First-person buttons ("Send my code") convert 1.18% vs 0.80% for neutral wording (English popups, median campaign). That gap is reliable in our data and a strong candidate for variant B, but it's still a comparison between campaigns: test it on your own traffic first.
How is A/B testing different from multivariate and split-URL tests?
An A/B test changes one thing. Each extra version or combination splits traffic further, so it needs more visitors:
| Test type | What changes | Use it for |
|---|---|---|
| A/B (split) test | One change, two versions | One hypothesis to check |
| A/B/n test (A/B/C with three versions) | One element, three or more versions | Several candidates, such as three headlines |
| Multivariate test | Several elements, in combinations | High-traffic pages; shows how changes interact |
| Split-URL test | A whole page, each version on its own URL | Redesigns; follow the SEO rules below |
What should you A/B test on a popup?
A popup has four parts worth testing:
- Offer: a discount against free shipping or early access.
- Headline and button: the promise and the call to action, such as first-person wording.
- Form fields: email only against email plus name, where our data shows no reliable effect.
- Trigger: a time delay, scroll depth or exit intent.
I test these before the button colour: in our data, button colour showed no measurable effect on conversion, though the data are too noisy to rule out modest effects. For set-ups, see A/B testing for popups and how to A/B test Shopify popups.
A/B split test examples by industry
These examples are adapted from Optimizely's A/B testing glossary.
1. A media company might want to increase readership, increase the amount of time readers spend on their site, and amplify their articles with social sharing. To achieve these goals, they might test variations on:
- Email sign-up modals
- Recommended content
- Social sharing buttons
2. A travel company may want to increase the number of successful bookings completed on their website or mobile app, or may want to increase revenue from ancillary purchases. To improve these metrics, they may test variations of:
- Homepage search modals
- Search results page
- Ancillary product presentation
3. An e-commerce company might want to increase the number of completed checkouts, the average order value, or increase holiday sales. To accomplish this, they may A/B test:
- Homepage promotions
- Navigation elements
- Checkout funnel components
4. A technology company might want to increase the number of high-quality leads for their sales team, increase the number of free trial users, or attract a specific type of buyer. They might test:
- Lead form components
- Free trial signup flow
- Homepage messaging and call-to-action
A/B split test tools
- Google Analytics measures results but no longer runs tests: Google Optimize shut down on September 30, 2023, and Google now integrates Analytics with third-party A/B testing tools.
- Optimizely is a dedicated experimentation platform with a visual editor, server-side testing and multivariate tests.
- Popupsmart runs A/B split tests on popups: design, audience and trigger variations with adjustable traffic ratios, from the Advanced plan (not on Free or Basic; see Popupsmart pricing). The help center shows how to set up a popup A/B test.
Does A/B testing affect SEO?
A/B testing does NOT hurt rankings when it's set up the way Google asks: Google's guidance on A/B testing says small changes, such as a button's colour or its call-to-action text, often have little or no impact on a page's search snippet or ranking. Popup tests on the same URL need none of the URL rules below.
If you run a split test with multiple URLs, use rel="canonical" on each variation to prevent Googlebot from getting confused by similar versions of the same page. Google recommends it over a noindex tag.
If you run a split test that redirects the original URL, use 302 (Temporary) Redirects rather than 301s (Permanent) to enable Google to keep the original URL.
You need to avoid conducting long and unnecessary experiments: Google may treat an unnecessarily long one as an attempt to deceive search engines. When the test ends, remove the test URLs and scripts.
Related Blog Post
Why Emails Go to Spam Instead of Inbox
Best WordPress Popup Plugins: Comparison Guide
Increase Email Open Rates with Catchy Email Subject Lines
Increase Email Open Rate: Email Marketing Guide For Beginners
Email List Building: Proven Methods to Grow Your Email List (for Beginners)
