Measurement

A before-and-after number is not a result

Turn a feature on, watch the number move, claim the difference. It is the standard way this industry reports uplift, and it cannot tell you whether the feature caused anything. Here is what we do instead — and how to interrogate anyone's claim, including ours.

What a before-and-after number actually contains

Switch a sizing tool on in March and compare March with February, and the difference is not the tool. It is the tool plus everything else that changed between those two months — and something always did.

Season
Fit confidence in a knitwear drop is not fit confidence in a swim drop. The calendar moves conversion on its own.
Promotions
A sale in the after-window lifts everything in it. So does a sale in the before-window, in the other direction.
Traffic mix
Paid, organic and returning shoppers convert at different rates. Change the mix and the blended rate moves with no change in behaviour.
Your other changes
Almost nobody ships one thing at a time. New photography, a checkout tweak and a sizing tool in the same month are indistinguishable afterwards.
Catalogue
New styles, restocks and sell-outs all change what is available to buy, which changes what gets bought.

None of these are exotic. They are ordinary retail, and every one of them is inside a before-and-after figure with no way to take it back out.

What we do instead

A randomised holdout — the design used to evaluate medicine, applied to a storefront.

  1. Split Every arriving shopper is randomly assigned to treatment or control. Assignment is at session level and holds across return visits, so a shopper does not flip between groups mid-decision.
  2. Withhold The control group sees the store exactly as it is, minus Magic Fit. They keep the brand's native size chart — the comparison is against what you already had, not against nothing.
  3. Run Both groups shop through the same window. Same promotions, same catalogue, same season, same traffic mix. Every confounder applies equally to both halves.
  4. Compare The difference between the two groups is the causal effect. Nothing else changed between them, because nothing else could.

What it produced at Motto

Brand
Contemporary womenswear, Shopify, Australia
Window
15 July – 14 August 2026
Design
Randomised 50/50 at session level
Exposed sessions
102,706
  • +11.4% New customers First-time purchasers, treatment against control.
  • +6.6% Conversion rate Share of exposed sessions that produced an order.
  • +27.5% Revenue through fit-guided sessions More of the store's revenue passes through shoppers who got fit guidance before buying.
  • +27.7% Shoppers engaging with fit guidance Control shoppers had the native size chart and used it. Treatment shoppers sought fit help far more often.

All results above are significant at p < 0.05. Read the full case study.

The finding a before-and-after could not have produced

Most of the lift did not come from Magic Fit converting fit-seeking shoppers better than a size chart does. It came from Magic Fit getting more shoppers to seek fit guidance at all.

That matters: it says the tool earns most of its return by pulling hesitant shoppers into a fit decision they would otherwise have skipped, not by winning a head-to-head against a size chart. A single before-and-after number is one number. It cannot be decomposed, so it cannot tell you which mechanism you are buying, and therefore cannot tell you where it will and will not work.

What a holdout costs

For the length of the test, half your shoppers do not get the tool. If it works, that is revenue you chose not to collect in order to find out by how much. We think a month of that is worth knowing the real number rather than a flattering one, but it is a genuine cost and you should weigh it rather than have it glossed.

Six questions for any vendor's number

Including ours. If a claim cannot survive these, it is a marketing asset rather than a measurement.

Was there a control group?
If the comparison is the same store before and after, every confounder above is inside the number. This is the question that eliminates most claims.
Were both groups measured over the same window?
Sequential periods are not comparable. Same weeks, same promotions, or it is not a controlled comparison.
How was assignment randomised, and at what level?
Session, visitor or device changes what the number means. Non-random assignment — by page, by collection, by opt-in — reintroduces selection.
How many sessions, and is significance stated?
A lift measured on a few thousand sessions can be noise. A p-value or confidence interval should be published, not available on request.
Is the denominator the whole store or the engaged subset?
Quoting conversion among shoppers who used the tool, against everyone who did not, compares high-intent shoppers with average ones. It reliably produces a large, meaningless number.
Who computed it?
If the vendor cannot show the method, the window and the arm sizes, the figure is a marketing asset rather than a measurement.

Run one with us

We will run a holdout on your store during your trial, and you keep the result whichever way it goes. If Magic Fit does not move your numbers, a month of measurement is a cheaper way to learn that than a year of assuming it did.