Most "A/B testing best practices" articles were written for SaaS signup flows, then pasted onto ecommerce. That's why so many of them don't survive contact with a scaling Shopify store.
Your traffic isn't spread evenly, it's stacked on a handful of templates, so your PDP has sample and your collection pages don't. Most of your new visitors come from paid, which means a creative change upstream can shift your audience in the middle of a test. And seasonality is violent: a test that runs across a promo window is measuring the promotion, not the variant.
None of that is in the SaaS playbook, and it's the difference between a testing program that shows up in revenue and one that only stays busy.
Shopify's own guidance says the same thing in plainer terms: best practices are a starting point, not a universal playbook, and research is what should decide the test. So that's where we'll start, because without customer insight you don't have a reason to run the test in the first place.
Diving into minor design tweaks without a clearly defined problem teaches you almost nothing, whether the variant wins or loses. Baymard's research team has argued that for years, and it holds on Shopify.
Define the problem first (for example, high abandonment at a specific checkout step), then state why you believe it happens, then design substantial variations that would only work if that hypothesis is true.
At SplitBase we structure that work through the 3Ps:
A weak hypothesis: "Change CTA from blue to green."
A strong hypothesis: "Cold Meta traffic doesn't trust our clinical claim above the fold, so they bounce before scroll. If we move third-party validation and a plain-language mechanism explanation into the first viewport for that segment, add-to-cart rate will rise without increasing discount reliance."
Only the second hypothesis produces tests worth your traffic. It names an audience, a cause, an intervention, a metric, and a guardrail, and it can be proven wrong, which is the whole point.
Not every idea should be a full theme split, so match the surface to the risk and the learning goal. Get this wrong and you'll spend structural-test QA on a change that could have shipped as a landing page test in a week.
Use these when the change is structural: PDP information architecture, collection filters, homepage hierarchy, navigation. Theme-level tests need careful QA across templates, apps, and markets, and the usual failure mode is flicker or an app conflict, not "bad creative."
Use these when paid creative makes a specific promise. The page should complete that promise, not dump traffic on a generic PDP. For scaling brands, landing page tests often pay faster than homepage tests because intent is clearer and sample accrues on the campaign.
Use these for offer architecture: bundles, subscriptions, shade finders, comparison blocks, UGC placement. Measure purchase rate and AOV together, because a module that lifts ATC but tanks AOV isn't a win.
Cart abandonment sits around 70% (Baymard), so checkout is where the money leaks, which is exactly why you test it with more discipline than anywhere else, not less.
Baymard's benchmarks show most large sites still leave substantial conversion on the table in checkout UX alone. On Shopify that means fewer simultaneous changes, stronger instrumentation, and explicit guardrails on payment error rates and express checkout behavior.
Upsells, thank-you page offers, and onboarding emails affect LTV, so score them on more than first-order revenue. Include return rate and support contacts when you decide what "won," or you'll optimize for one-time revenue that comes back as refunds.
This is where most roadmaps quietly break. Teams plan a quarter of checkout experiments, then discover their plan and stack can't run them.
The practical rule is to pick the tool after you know what you need to test, not the reverse. A tool decision made first tends to quietly define your roadmap for a year.
The fastest way sophisticated brands still ship noise is watching results daily and stopping the moment the dashboard turns green. You don't fix that with more statistics knowledge. You fix it by deciding the rules of the test before a single visitor sees it.
Lock these six things up front:
You don't need to run the math by hand. Drop your real traffic and conversion rate into a sample size calculator and you'll know in about ten minutes whether the test you want is even possible on that template.
Nielsen Norman Group makes the same point from the research side: small lifts need much larger samples, and a tiny effect may not be worth the traffic it costs to prove.
And if the test isn't possible, change the test, not your standards. If the lift you care about needs six weeks of traffic you don't have on that URL, run a bigger change, move to a higher-traffic template, or roll it out in stages with clear decision criteria.
"Let's just run it for a few days and see" is how you end a quarter with nothing but inconclusive tests.
Winning Shopify programs treat research as the input, not a postmortem. Here's a practical stack for DTC teams:
If your backlog is full of "test ideas" and empty of customer problems, pause launches until the 3Ps backlog is rebuilt.
CRO and brand aren't a tradeoff when the work is done correctly.
Here are the brand-safe rules we hold with premium clients:
Those rules aren't a tax on the work, they are the work. We call it Full-Business CRO™: conversion gains that still look and feel like the brand customers already chose.
A clean "conversion rate up 8%" slide can hide a worse business. Here's how that happens:
To catch that, run the same reporting pack on every Shopify test:
Ship only when primary and guardrails agree, and archive the learning either way. Losers that invalidate a hypothesis are assets, and undocumented winners are myths.
At scale, one lonely PDP test per quarter won't move the year. Build a portfolio instead:
Sequence tests so they don't contaminate each other on the same template, and keep a shared decision log with hypothesis, design, result, ship or no-ship, and follow-up. That log is how teams stop retesting the same button every peak season.
If you spend seriously on paid media, your landing pages are the highest-return testing surface you own, and usually the most neglected.
The reason is structural. Paid traffic arrives with a specific promise made by a specific ad, and a generic PDP only answers a general question.
A dedicated landing page can answer the exact objection the ad created, and because campaign traffic is concentrated, the sample you need shows up in days rather than months.
Here's what to test there, in rough order of impact:
Two cautions before you get excited.
First, a landing page win doesn't automatically transfer to the PDP, because it's a different audience with different intent. Second, if page-building is slow, testing velocity dies, and brands that build landing pages in days rather than sprints run several times as many experiments per quarter on the same budget.
Most brands at this stage run digital with two or three people who are already fully booked. A research-first testing program doesn't require a new hire, but it does require specific things, and it's worth being honest about them before you start.
Access. Analytics, your Shopify admin at a level that allows theme and app changes, your ad accounts for segment analysis, and read access to support tickets and reviews. Gathering these is usually the single biggest source of delay at kickoff.
Decisions. Someone with the authority to approve a variant that looks different from the current site. Programs stall when every variant has to go through a brand approval chain that was designed for campaigns.
Time. A recurring working session plus asynchronous review of hypotheses and results. That's a few hours a week of your team's attention, not a full-time role, but it can't be zero, because the person who knows the customer best usually works for you, not for your agency.
Patience through the first cycle. The first weeks are research and instrumentation. If someone needs a win in week two, say so at the start so the roadmap can front-load a faster surface.
If you're evaluating outside help, ask for these specifics from anyone you talk to, including us.
Timeline. Research and instrumentation first, then a prioritized backlog, then tests running continuously. Ask for a week-by-week picture of the first cycle, and ask when the first test goes live.
Cost and model. CRO at this level is usually a monthly retainer covering research, design, build, QA, analysis, and reporting, rather than per-test pricing.
Ask what's included, what's billed separately, and what happens when you want more landing pages in a given month. Current pricing is published on our site, so ask us to confirm the number in writing rather than working from a page you found in search.
Traffic requirements. Don't accept a universal number. Ask them to calculate expected runtime for a realistic MDE on your actual baseline traffic and conversion rate, on the specific templates they intend to test. If they can't do that on a first call, that's informative.
Team. Who does the research, who designs, who builds, who analyses, and whether those are senior people or a junior working from a template.
Deliverables. Research documentation, a maintained backlog, test briefs, variant designs, QA notes, and a readout per test with the decision recorded.
Run this before every launch:
We don't borrow a competitor's test list and paste it onto your theme. SplitBase is a research-first CRO and landing page partner for scaling DTC brands, and a Shopify Plus partner.
We build a brand-specific backlog from Patterns, Perception, and Proof, then design experiments that can win commercially without eroding what makes the brand valuable.
That's the standard behind more than a decade of work with premium skincare, haircare, men's grooming, supplements, eyewear, and wellness-tech brands, including household beauty names, where the mandate has consistently been the same: conversion programs that still look like the brand on the other side of the test.
If you want a second set of eyes on your testing roadmap, instrumentation, or a research-first backlog for peak, book a free discovery call, and we'll tell you plainly whether your program is set up to learn, or only set up to launch.
Partly, and it depends on your plan and stack. Most of what people call "checkout testing" is really cart, cart drawer, and pre-checkout testing, which is fully available. True checkout-step testing is more constrained than older articles suggest. Confirm what your plan supports before committing a roadmap to it.
You need a tool that can run your planned tests without flicker and report the metrics you've committed to. Intelligems and Shoplift are the common Shopify-native choices, and other platforms make sense for more complex statistical or cross-platform needs. Choose after the roadmap, not before.
In a full engagement, the agency should design, build, and QA the variant, and hand shipped winners to your developers in maintainable form. If you're told to have your own developer build every variant, testing velocity will be capped by your dev backlog.
It gets documented and it informs the next hypothesis. Roughly speaking, a meaningful share of well-designed tests won't win, and that's the nature of experimentation. Any agency promising otherwise is either running trivial tests or reporting them loosely. The value is in the rate of learning and the compounding of implemented winners.
Brand guardrails go into the scorecard alongside conversion, variants stay inside the brand system, and discounting and false urgency are off the table as default levers. If the only winning path is a bigger discount, that's a pricing decision for the business, not a test result.