Skip to content
A/b test
Websites Business Growth Shopify

What Is A/B Testing? And Why You Need It For Your Ecommerce Business

Zachary Kohler
Zachary Kohler

A/B testing is simple in concept: you split your audience between two versions of your website, and you find out which one actually makes more money.

Half your traffic sees the version you already have. The other half sees a version with a specific change, That could be a new headline, a different navigation, whatever the hypothesis is.

Run the test long enough, and the data tells you which version made you the most money. Not which one you like more, not which one your team likes more, which one your customers actually voted for with their wallets.

Most ecommerce owners make website decisions based on opinion. A/B testing replaces the opinion with data and gives you an answer based on revenue.

TEST A

How you know a test is won.

Say 100 people visit each version of your site test and one version does 2% better. That's a difference of about two orders. It's very easy for that gap to come from noise, not from your change in that test.

One visitor maybe was already primed to buy from retargeting ads and they happen to land on the "worse" version and buys, suddenly your data is not telling the full story. This is why statistical significance exists to solve: how confident can you actually be that a result is real, rather than random variation.

Statistical significance is a measure of confidence. It tells you how likely it is that the difference between two versions is real, not just random noise. Some visitors were always going to buy, no matter what they saw. Some never would have. Statistical significance filters that noise out and tells you whether what's left is a genuine pattern. As a rough guide, we look for around 300 orders per version before trusting a result and even then, it's a matter of confidence, not certainty. An 80% statistical significance means there's roughly an 80% chance the change is genuinely better, not a guarantee. Getting this right takes real traffic and real discipline. Read too much into a result too early, and you'll roll out a change that never actually helped or worse, one that actually hurts revenue.

What it actually costs you to skip testing

If your store is doing $1M+ a year you are in a good position to start running tests. You have the traffic and you are likely to see better results from small changes.

At this level making changes based on a gut feeling is silly. Say you make a change and it makes things 5% worse. If you're doing $1M a year, that's $50,000 gone. If you're doing $50M a year, that's $2.5M. And here's the part that makes it dangerous: you will probably never know it happened. You don't see the revenue you didn't make.

You just see the number that came in, and you assume it's the economy, or your ads underperforming, or seasonality, when the real cause was a change made to the site months earlier that nobody ever tested.

A/B testing isn't a nice-to-have layer on top of website changes. It's the only way to know whether a change helped, hurt, or did nothing at all, instead of finding out a year later in your revenue numbers, with no way to trace it back.

Why testing matters even when you trust the people making the changes

This isn't really about whether you trust your team or your agency. It's about the fact that trust isn't a metric.

A lot of businesses and a lot of agencies measure activity instead of results. "We made 40 changes to the site this month" sounds productive. It means nothing on its own. Anyone can make a hundred changes to a website. The question that actually matters is whether any of those changes moved revenue per visitor. Testing is what turns "look how much we did" into "here's what actually worked," and it's the difference between activity and outcomes.

It also solves a problem that has nothing to do with anyone acting in bad faith: business owners know their own site too well to see it the way a new customer does. A layout that feels obvious to you, because you've stared at it for three years, might be genuinely confusing to someone landing on it for the first time. Testing catches that gap between how well you know your product and how a stranger actually experiences your site.

And it ends a cycle a lot of growing stores get stuck in: a new ecommerce manager comes in, doesn't love something on the site, and changes it back to how they'd prefer it with no data behind the swap either way. Multiply that across a few hires over a few years and you're not improving, you're just going in circles based on whoever's currently in the role. Tested results settle that argument permanently. "We tested removing this page. It's up 10%. It stays" is a very different conversation to two people trading opinions.

Why measuring the right number matters as much as testing itself

Most stores that test at all make the same mistake: they optimise for conversion rate alone, and stop there.

Conversion rate is only half the picture. If a change increases conversion rate by 5% but drops average order value by 6%, you've actually gone backwards, you just wouldn't know it if conversion rate was the only number you were watching. That's why the number that actually matters is revenue per visitor (RPV): conversion rate × average order value. It's the only metric that captures the full picture of whether a change made you more money, not just more orders.

Testing lets you take real swings without betting everything on them

One of the most underrated things about testing is what it does to risk.

Without testing, a big idea is an all-or-nothing bet: you rebuild the whole site around it, and if it's wrong, walking it back is slow, expensive, and awkward. With testing, you can send a small slice of traffic, 5% or 10%, if you've got the volume, to a genuinely uncertain idea instead of going all in. If it doesn't work, you turn it off. No expensive redesign to unwind, no months of sunk cost to justify.

This also protects you from a subtler trap: comparing "last month" to "this month" and calling it a test. Running a change live for 30 days and then eyeballing whether revenue went up isn't a real test,there are too many other variables moving in the background at the same time: seasonality, ad account changes, interest rates, tax refund timing, competitor activity. A real A/B test splits the same audience, at the same time, under the same outside conditions, and isolates the one thing that's actually different between the two groups. That's the only way to know the change caused the result, rather than something else that happened to overlap with it.

The compounding effect of testing over time

The real payoff of testing isn't any single result. It's what happens when you stack a lot of small, proven wins on top of each other, month after month.

Without testing, website changes are a coin flip, two steps forward, two steps back, with no way to tell which direction you're actually moving in over a year. With testing, a change that doesn't work gets killed fast, usually after a couple of weeks of split traffic, not years of the whole site running worse. A change that works, stays. That's the difference between a website that's slowly improving every quarter and one that's just changing shape without ever getting better.

It also protects you from a habit a lot of ecommerce brands fall into without noticing: copying whatever a competitor is doing because they seem to be doing well. Their site was built for their customer, their price point, and their positioning, not yours.

If copying competitors reliably worked, every store in a category would eventually convert equally well, and that's clearly not what happens. Testing keeps your site built around what actually works for your customers, not what looks good on someone else's.

And because a real test isolates one variable at a time, it also surfaces problems fast. If a change breaks something , a broken link, a form that stops submitting, a payment method that quietly fails, a proper test shows up as an immediate, visible gap between the two versions, instead of two weeks of "why are sales down" with no clear cause.

Testing works at the audience level too

A/B testing isn't limited to "everyone who visits the homepage." You can run tests targeted at a specific audience segment, a landing page built specifically for traffic coming from a Meta campaign, for example, tested against your standard page for that same traffic. That lets you validate messaging and offers for a specific audience without changing the experience for everyone else, which matters more the more channels you're running ads through.

What this means for a $1M+ Shopify brand

If you're running significant traffic and real ad spend, the question isn't whether to test, it's whether every change to your site is currently backed by proof, or just by whoever made the call that week.

This is exactly the discipline we build into every engagement: every change we make gets tested against your actual traffic before it's treated as a win, measured against revenue per visitor, not just conversion rate in isolation.

Book a free RPV Opportunity Call,  a 60-minute walkthrough where we identify 3–5 specific revenue opportunities on your store.

FAQ

What is A/B testing in ecommerce? A/B testing splits your website traffic between two versions — your current site and a version with one specific change — to see which one generates more revenue. It replaces guesswork and opinion with actual proof of what works for your customers.

How much traffic do I need for a reliable A/B test? It depends on the size of the effect you're testing for, but as a general guide, we look for somewhere around 300 orders per version before treating a result as statistically reliable. Smaller sample sizes make it much easier for random variation to look like a real result.

What's the difference between conversion rate and revenue per visitor as a testing metric?Conversion rate only measures how many visitors bought. Revenue per visitor (conversion rate × average order value) measures how much money each visitor generated. A change can increase conversion rate while decreasing revenue per visitor overall — which is why RPV is the more reliable number to test against.

Why is testing better than just launching a big website change? A full redesign with no testing is an all-or-nothing bet — if it underperforms, walking it back is slow and expensive. Testing lets you validate an idea on a slice of traffic first, so you can kill what doesn't work in weeks instead of discovering the impact a year later in your revenue numbers.

Share this post