Search for an A/B testing framework and you’ll find the same five steps everywhere.
Identify a problem. Write a hypothesis. Build a variation. Split traffic. Check for statistical significance.
That’s a workflow for running a single test, and it’s fine in that use case, but it has real downsides once you’re testing month after month.
Writing a hypothesis means predicting an outcome, and predicting an outcome creates bias. People start rooting for a variation. That’s how tests get stopped early.
Measuring only the final conversion rate tells you whether a test won or lost, but never why. And because nothing links one test to the next, month six of your program looks exactly like month one: someone suggests an idea, you test it, you move on.
This article lays out the framework we use instead, across hundreds of A/B tests for ecommerce brands. Here’s what’s in it:
- The 8 Purposes. Classify every test as Brand, Discovery, Product Appeal, Product Detail, Price & Value, Usability, Quantity, or Scarcity, so you can start to notice themes around which types of changes actually move the needle for your store.
- The Question Mentality. Replace the hypothesis with a list of questions, which removes prediction bias and increases how much you learn per test.
- A standard goal set on every test. Transactions, revenue, checkout and cart pageviews, and add to cart clicks, so you can see where in the funnel a change did its work.
- Analysis by segment and by purpose. Update what you believe about each purpose, then generate the follow up test.
- ROI checkpoints. The traffic and revenue levels where this framework is worth building.
What an A/B Testing Framework Actually Is (and What It Isn’t)
A/B testing (also called split testing) itself is simple. You split traffic between a control (your baseline) and a variation, and decide based on actual user behavior rather than what users said they’d do in a survey or a conference room.
An A/B testing framework is the structure that decides what you test, how you phrase the test, what you measure, and how each result feeds the next one.
Three things routinely get called a framework, and only one of them is:
- The software. Optimizely, VWO, and Adobe Target are platforms. They give you a visual editor, randomize traffic, deliver variations for A/B and multivariate testing, and calculate significance.
- The five step build-measure-analyze loop. That’s a workflow. It runs one test.
- The layer above both, which organizes individual tests into a body of knowledge about your customers. That’s the framework.
Without that third layer, you can run a technically perfect test every two weeks for a year and still be unable to answer the only question that matters: what do our customers actually care about?
Why Most A/B Testing Frameworks Fail: Tunnel Vision Testing
We call the status quo tunnel vision testing: testing each thing on your site in a silo without connecting them into larger learnings about what your customers fundamentally want. Someone raises a concern in a Monday meeting, you build a test for it, you get a result, you move on. No connected learning accumulates.
We’ve been running conversion rate optimization (CRO) operations for ecommerce brands for 10 years, trust us when we say this: this is how 99.9999% of organizations run AB tests.
It’s understandable, because optimizing a site is a daunting task with a million possible changes, big and small. With no framework to sort them, the loudest people’s opinions get tested first.
Copying best practices doesn’t fix this. We’ve found repeatedly that many widely accepted best practices don’t improve conversion rate for every store. Best practices are someone else’s test results applied to your customers.
Reporting results as “won” or “lost” doesn’t fix it either. A year of it and you’ll have a spreadsheet with 24 rows of wins and losses, but what are the takeaways? You need an actual system to analyze results and find connected themes. It’s non-trivial, critical thinking-required, human work.
The real cost of all of this wasted AB testing is time (and, yes, ultimately revenue). Guessing wrong about which category of change matters can burn months. A store selling to twenty-somethings may have essentially zero usability issues, so six months of bigger CTAs and shorter forms gets you nowhere while the actual barrier was that they didn’t want the product enough or didn’t trust the brand yet.

That’s the problem that our AB testing framework, The Purpose Framework, was built to solve.
The Purpose Framework: Classify Every Test Into One of 8 Purposes
Every A/B test on an ecommerce site can be categorized as having one of 8 purposes. That’s a sweeping claim, so let me explain what they are. (The full write-up, with client case studies and win-rate charts, is in our original Purpose Framework article.)
1. Brand. Tests that increase trust, credibility, or appeal of the overall brand. Adding press mentions, founder story content, guarantees, or review counts sitewide.
2. Discovery. Tests that make it easier to find or discover the right product. Navigation changes, filtering, category exposure, search prominence, quiz flows.
3. Product Appeal. Tests that make individual products more appealing through messaging, positioning, or imagery. New hero images, lifestyle photography, reworked headlines on the PDP. (Our review star rating test is an example, and it carried a second tag: Brand.)
4. Product Detail. Tests that highlight or specify details that help customers choose. Ingredients in a lotion. Specs on a car part. Fit and sizing on apparel.
5. Price & Value. Tests that improve the price to value ratio. Free shipping messaging, bundle pricing, subscription discounts, savings displays, financing.
6. Usability. Tests that reduce UX friction. Fewer form fields, larger call-to-action tap targets, fewer checkout steps. A classic example: replacing city, state, and zip fields with a zip code lookup that fills in the rest automatically.
7. Quantity. Tests that increase average order value or cart size. Upsells and cross-sells, multi-pack defaults, cart-page recommendations.
8. Scarcity. Tests that create urgency by highlighting limited time or limited quantity. Low stock indicators, sale end dates, cart timers.
The taxonomy isn’t the point. You can even create your own purpose buckets that are specific to your store (or not use some of the ones above, for example if you never have low inventory in your business, scarcity isn’t relevant for you). Once every test carries a purpose tag, individual tests connect into a larger story about what compels your customers to buy. Tracking which purposes get tested most, and which win, tells you where the real barriers are, so you can focus on what matters and stop wasting time on what doesn’t.
A note on where this came from. The Purpose Framework grew out of a simpler split we wrote about years ago: every change either makes it easier to get from A to B (usability) or makes B more desirable (desirability). Still true, but “desirability” covers too much ground to act on. The 8 purposes are that idea made actionable for ecommerce.
How to Use the Purpose Framework to Decide What to Test Next
Once you’ve established the purpose buckets and gotten buy in with the organization on this framework, here are the steps we’ve found work well to put it into practice.
Start with an audit of tests you’ve already run. Tag your last 12 to 24 tests with purposes. This almost always reveals heavy clustering into one or two purposes (usually Usability, because it’s the easiest to build) and purposes with zero tests. The unexplored purposes are frequently where the biggest lifts hide: if you’ve never run a Price & Value or Quantity test, you don’t know they don’t work. You know you haven’t looked.
When a purpose wins repeatedly, double down. When one loses repeatedly, stop spending design and dev cycles there. A store that wins on Product Detail over and over is telling you its customers need more specification before they’ll buy. Three losses on Scarcity is information too: reallocate.
Purpose tagging also resolves internal debates. “Should we make the button bigger” is an argument between opinions. “Is usability a barrier for our customers” is a question with a test record behind it.
Weight checkout and payment steps higher. A lift at the bottom of the funnel flows straight to revenue: 25% more people completing the payment page is a 25% revenue increase, while 25% more add to carts on a store where 40% of carts check out is roughly a 10% increase. Same test effort, different leverage.
Use user research to point you at purposes. On-page polls, post-purchase surveys, heatmaps, session recordings, and live user testing are cheap relative to a testing program. First figure out why users aren’t converting. Then give them what they want.
Note: if you’d like us to run a purpose audit on your existing test history and tell you which purposes you’ve never explored, you can learn about working with us here.
The Question Mentality: A Better Alternative to Hypothesis-Based Testing
Another more subtle, and unusual, framework we use is around the team psychology around AB testing. We wrote the full argument here, but this is the short version.

Every framework you’ll read tells you to write a hypothesis. “Adding a product video to the PDP will increase conversion rate by 8%.”
On the surface nothing is wrong with this and it’s true that on paper, every AB test has a hypothesis behind it. But in practice we’ve noticed a practical issue: Hypotheses naturally carry a prediction (“Video will increase conversion rate”) and that prediction comes from a person (on the team or in an agency) and humans have feelings, emotions, and bias. So when someone feels like their prediction and thus some of their credibility is on the line because a test was “theirs” (This is real language people use on CRO teams: “That’s Jen’s video test.”) they have a bias to it winning. That bias is counter-productive.
To avoid this, we use questions instead. For that same video test, the questions are:
- Do users care about watching a video at all?
- Will they actually watch it, and how far in?
- Will it affect add to cart rate?
- Does it change how much of the other information on the page they read?
- Does it affect returning users differently than new users?

Look at how refreshingly unbiased those questions are. Three things happen when you frame tests this way.
It reduces the risk of stopping tests early. When you’re waiting for a prediction to be confirmed, four days of favorable data feels like confirmation. When you’re waiting to answer five questions, it doesn’t answer any of them yet.
It stops people from tying their reputation to a test outcome. If the marketing manager predicted the video would win, the test is now about the marketing manager. Questions have no author to embarrass.
It increases learning per test. You set out to answer several things rather than validate one, so you get several answers.
The clearest benefit shows up on losing tests. Under hypothesis framing, a loss goes in the “lost” column of the spreadsheet and nobody revisits it. Under the Question Mentality, a loss teaches you that users don’t care about the thing you added, which redirects your view of that entire purpose.
Goal Tracking: What to Measure on Every Test
Most framework guides focus on statistical analysis and barely mention what to measure. That’s backwards: significance tells you whether a difference is real, your goal set tells you what happened. (We wrote a practical breakdown of ecommerce goal setups here.)
Use a standard set of success metrics on every test, so results are comparable across your whole program:
- Transactions
- Revenue (and revenue per visitor)
- Pageviews of each key page in the checkout flow
- Cart page views
- Add to cart clicks
Then add test-specific goals on the exact element you changed. If you added an accordion of ingredient detail, put a click goal on it. A flat result where nobody opened the accordion means something completely different than one where it got a 30% click-through rate and those users still didn’t buy. Track both click goals (intent) and pageview goals (the user actually arrived).
Integrate the testing tool with Google Analytics or Adobe Analytics, so you can validate its numbers independently and analyze metrics it never captured.
Here’s why this matters in practice. A test that raises add to cart clicks but doesn’t move transactions is a completely different lesson than a test that moves nothing: the first worked and something downstream ate the gain, the second didn’t register with anybody. Those results point at entirely different follow-up tests, and only funnel level goals can tell them apart.
Analyzing Results So Each Test Informs the Next
This is the step that turns a workflow into a framework that compounds, and it’s the step most programs skip.
Report more than won or lost. State what the result changes about your understanding of that purpose. If the Product Detail test won, does that mean detail generally, or detail about this one attribute?
Segment every result. At minimum, new vs. returning users and traffic source. A flat overall result frequently hides a real win in one segment and a real loss in another: Brand tests, for example, often do nothing for returning users who already know you and a lot for new visitors.
Examine how every goal in the funnel moved, not only your primary metric, to locate where user behavior actually changed.
Feed the result back into the purpose ledger. Keep a running score for each of the 8 purposes: tests run, won, lost, inconclusive. This document is the actual deliverable of a CRO program; the individual results are just the raw material.

End every analysis with a specific follow-up test. Not “we should explore this further.” An actual named next test. This is what keeps a program moving instead of restarting the idea hunt every two weeks.
Applying the Framework to Mobile
Most ecommerce stores now get more mobile traffic than desktop, and mobile conversion rates are typically around half of desktop rates. That gap is the single largest unexplored area in most testing programs.
Three adjustments when you apply the framework to mobile:
Run the purpose audit separately for mobile and desktop. The winning purposes frequently differ, and a combined ledger averages away the difference. Expect Usability to pay off more on mobile, where form fields, tap targets, and extra checkout steps create friction that desktop users barely notice.
Discovery is often the mobile bottleneck. Mobile collapses the entire catalog behind the hamburger menu, one hidden click away from every visitor. We’ve tested exposing category links directly on the mobile homepage (we call it a Link Bar) and saw pageviews of those category landing pages increase 10% to 12% with 99%+ significance, along with a likely lift in completed orders. Our study of the top 40 U.S. ecommerce sites’ mobile checkouts is a useful starting list of further mobile test candidates.

QA every variation across devices before launch. A variation that’s broken on one popular device produces a false loss, which enters your purpose ledger and misleads your strategy for months.
When an A/B Testing Framework Is Worth Building (ROI Prerequisites)
None of the above is worth doing at every company. Two rough criteria: $2 million or more in annual revenue and 100,000 or more monthly unique visitors. Enough traffic to reach meaningful sample sizes, and enough revenue that a percentage lift is worth the effort. The numbers are arbitrary in the way a driver’s license age is arbitrary: don’t argue the exact digits, adjust them for your business.
A reasonable planning assumption: a 10% revenue increase within about 6 months. At $5,000/month in testing costs, that’s $30,000 spent against a $200,000 annual lift on $2 million in revenue. Above $10 million, starting a program is a no-brainer. Below $1 million the math gets murky, and businesses at that level usually have bigger fruit hanging: we’ve seen client traffic double from one year to the next through SEO or paid media. If that’s still available to you, pick it first. The full ROI math is here, including agency costs versus in-house. And no one can predict your exact result, including us, but 10% in 6 months is achievable and our agency has hit it multiple times for multiple businesses.
The corollary rule. Meeting the thresholds gives you the potential for a good ROI, not a guarantee. If you’ve been testing for six months without routine lifts of 5% to 10% or more in orders, your framework isn’t working. That’s usually a strategy problem, not a traffic problem.
Common A/B Testing Framework Mistakes
Most “mistakes to avoid” lists are about statistics. These are the ones that happen at the framework level, which cost more.
Testing best practices without asking whether that purpose is even a barrier. Free shipping banners don’t help a store whose customers already know shipping is free. We removed two form fields for one client and saw no difference across nearly 100 conversion events per variation, while a messy-form cleanup we cited in the same article apparently lifted submissions 35%. Same purpose, opposite results, which is why the ledger has to be built per store.
Only testing one purpose because it’s the easiest to build. Usually Usability, because it needs no new copy, no photography, and no merchandising approval. Meanwhile Brand, Price & Value, and Quantity go untouched for a year.
Testing low-leverage pages. Blog templates, About pages, and low-traffic landing pages are safe to test and rarely worth testing. Checkout and payment steps convert lifts directly into revenue.
Building one design concept for an important test. Design details can decide the outcome, so one mediocre execution can bury a good idea, and you’ll write the idea off rather than the design.
Letting a web design agency roll out a full redesign untested. It isn’t in their interest to find out whether the new design performs worse than the original. Test it in phases and find out which elements helped and which hurt.
Aside: we once had an in-house designer at a client ask if they could put a hamburger menu on desktop because it “looked sleek.” That’s test selection without a framework.
A/B Testing Frameworks vs. A/B Testing Tools
Some people arrive at this topic looking for a tool recommendation, so briefly:
Platforms like Optimizely, VWO, and Adobe Target handle randomization, visual editors, variation delivery, goal tracking, and significance calculation. Any of them can run the framework described in this article. We’ve run it in all three, and compared them at length here (no affiliate relationship with any of them).
The tool decides how you serve a test. The framework decides which test is worth serving. The second decision has far more effect on your results, which is why switching platforms rarely fixes a program that has no framework. Whatever tool you use, the framework requirements are identical.
How Growth Rock Runs This Framework for Ecommerce Brands
Growth Rock is a CRO agency that works exclusively with ecommerce brands. We built the Purpose Framework and the Question Mentality out of running hundreds of A/B tests for ecommerce clients, because we needed a way to make test number 40 smarter than test number 4.
As a service, we run everything in this article for you: purpose tagging and analysis, variation design (multiple concepts on important tests), coding, goal and analytics setup, cross-device QA, and preview links so your team reviews every variation before it goes live. We also maintain a live database of the A/B tests we’ve run, organized by purpose, so you can look at the results behind the framework before you commit to anything.
Typical outcomes are conversion lifts of 5% to 10% or more, higher average order value, and a documented understanding of what motivates your specific customers. Best fit: ecommerce brands doing $2M+ in annual revenue with enough traffic to run meaningful tests, especially teams tired of settling site debates by opinion.
And as always: don’t assume any result in this article will apply to your store. What wins depends on your customers, and the only way to find out what they care about is to ask them with tests.
Learn more about working with us here, or join our email list to get new articles and A/B tests when we publish them.