We built one, and then measured how much its answer depended on a number that was not in anybody's data — a number we had chosen ourselves and printed a default for.
Across 256 shop configurations, changing that one default moved "some threshold pays for itself" from 3% of shops to 95%.
Every figure below is measured against the running tool, not argued. The same applies to the correction we tried and had to throw away.
Free shipping over $___. What goes in the blank?
The advice you will find is to set it 15–30% above your average order value. That is wrong twice.
Your past distribution cannot predict your future distribution. Changing the threshold changes the shape of the orders — that is what a threshold is for. Forecasting revenue under a new rule from the histogram the old rule produced assumes the thing you are changing holds still. That is an invented elasticity, and no arithmetic on an orders export can produce one.
A threshold you already have contaminates the evidence for the next one. If you run "free over $50" today, the orders sitting just above $50 are your own policy's footprint. Reading them as customer demand for a $50 line, and concluding $50 is about right, is circular.
So we built something that refuses to predict. Instead of an uplift, a break-even. For each candidate threshold T:
cost = the shipping revenue you actually collected on orders already at or above T
gain = for each near-miss order, (T − order value) × margin, minus the shipping it was paying
break-even top-up rate p* = cost ÷ gain
The cost side is measured — it is sitting in the merchant's own export, with no assumption in it at all. The uncertainty is pushed into one number a merchant is better placed to judge than we are. A sentence of the form "at least this share of your near-miss orders would have to top up" is a claim about their customers that they can assess. "This will raise revenue by $1,240" is a claim nobody can assess, which is why no figure of that kind is printed anywhere on the page.
That felt honest. It was, apart from one word.
Look again at near-miss order. Which orders are those?
An order at $95 against a $100 threshold, obviously. An order at $12, obviously not. Somewhere between them is a line, and that line is not in the export. So we drew it: a near-miss is an order that could clear the threshold by adding at most 25% of what it already contains.
25% is not a measurement. It is a number that sounded reasonable.
256 combinations of order value, margin, carriage price and band width, each one run through the actual page and its actual controls.
| Near-miss band | Shops where some threshold gets a break-even rate |
|---|---|
| 10% | 3% |
| 25% — our default | 47% |
| 50% | 83% |
| 100% | 95% |
Across all 256, top-ups alone could cover the whole cost at some threshold in 57%, and some threshold had a positive top-up gain at all in 75%.
Our own sample data sat in a region where nearly every row came back no top-up rate pays for this. We were on the point of treating that as a finding about free-shipping thresholds. It was a finding about our default.
Worse, the failure was pointed the expensive way. A merchant reading "no rate pays" on every row concludes free shipping is not for their shop — a conclusion the data does not support, arrived at because a number we chose put them on the wrong side of a line. The cost of that lands on them, not on us.
The obvious next move was to find the statistic that predicts which side of the line a
shop falls on. order value × band × margin ÷ carriage did it reasonably well.
Then some algebra suggested it was missing a term.
Work out when the gain first turns positive. Near-miss orders sit between
T/(1+w) and T, so the gap to close averages roughly
T·w / 2(1+w), and the break-even condition becomes
V·w·m / (1+w) > 2s. The (1+w) belongs in the denominator. It
looked like a straightforward improvement.
Measured:
| Statistic | All-negative up to | First positive at | Transition width | Cells in transition |
|---|---|---|---|---|
V·w·m / s | 1.11 | 0.85 | 0.26 | 16 of 256 |
V·w·m / (s(1+w)) | 1.01 | 0.57 | 0.44 | 39 of 256 |
The "improvement" nearly doubled the width of the region where the predictor cannot call it, and more than doubled the number of shops sitting inside that region.
We do not have a complete account of why, and that is the honest state of it. Two pieces are measured. The derivation assumes the boundary sits at the median order value; measured, the gain first turns positive at 0.9x the median, with a spread from 0.7x to 2.5x — and that figure is censored from below, because the candidate ladder starts at the 30th percentile and cannot look lower. Separately, the assumption that the average gap is half the band's largest gap holds up: measured at each shop's best threshold it is 0.59, against 0.50 for a uniform distribution, slightly high exactly as it should be for a distribution whose mass piles up at the band's lower edge.
Those two do not reconcile into the observed constant, and they are measured at different rows, so combining them into one derivation is not legitimate arithmetic — attempting it is how the wrong mechanism nearly reached this page. What we can say without hedging is the part that decided the outcome: the correction is algebraically motivated and empirically worse, and we found that out by scoring it rather than by believing it.
We stopped fitting a constant. The per-order condition does not need one. A near-miss order that tops up is worth having only if what it adds, at your margin, beats the shipping it was already paying you. So the page now says it in the merchant's own two numbers — on our sample data:
At your 55% margin and median carriage of $4.95, a customer has to add more than $9.00 for converting them to be worth having.
That is exact. It is their shop, not our regression. It is also the third time on this project that the exact local condition beat the fitted global one, so it is now the default rather than a preference.
We took the band out of our hands. A customer topping up an order does not add "the gap." They add an item. And the price of a typical item is sitting in the export we are already reading. So the near-miss band is now the gap that a single median-priced line item would close, measured from the merchant's own file — in our sample, $26.00, from 2,212 line items.
The break-even is now a range across two readings, with the cautious end named on the row: what happens if they add exactly the gap, and what happens if they add one typical item. Where a file has no line-item column, the old proportional band is used and the screen says, in those words, that it is a number we chose.
Being straight about this is the point rather than a disclaimer.
Find the number you invented — every calculator in this category has one, because the arithmetic cannot be completed without it. Ours was hiding inside the phrase "near-miss order." Then do the cheap thing we should have done first: hold everything else still, move that one number across its plausible range, and see whether your headline verdict survives.
If it does not, you do not have a calculator yet. You have a default with a table attached.
I maintain a free shipping threshold calculator that reports break-evens instead of predictions, reads the near-miss band from your own file, and refuses to rank thresholds at all when it can see a rule already in your data. It runs entirely in the browser and requires no account. hyperbarrow.com/free-shipping-threshold-calculator