Skip to content

Twelve ecommerce benchmarks with no source behind them

RIVYL~11 min read
Most repeated ecommerce benchmarks survive because nobody follows the chain to its end. Several of the best known ones do not have one.

Most ecommerce benchmarks in circulation have no primary source. Not a weak one. None at all. Follow the citations on the popular ones and they loop: a listicle citing an agency post citing a listicle, with no sample size, no date and no method anywhere in the chain. Here are twelve claims worth retiring, where each actually came from and what the original research says instead.

15.8%

US retail return rate in 2025, down from 16.9% in 2024

NRF and Happy Returns, October 2025

40%

Largest reason for abandoning checkout, once browsers are excluded: extra costs too high

Baymard Institute checkout survey, accessed August 2026

2019

Year the data behind the most-quoted site-speed statistic was collected, published 2020

Deloitte with Google, Milliseconds Make Millions

Do 30% of first-time buyers really come back?

Nobody knows, and the people repeating it don't either. The claim is that about 30% of first-time buyers make a second purchase and more than 50% of second-time buyers make a third. It has no primary source. Two mutually inconsistent versions circulate — one giving roughly 18%, 35% and 50% across the first three repeats, the other 27% and 54% — and neither states a sample, a date range or a category mix.

A pair of figures that disagrees with itself across two retellings hasn't been measured. It has been remembered. The related category tables (supplements at 35-45%, beauty at 30-40%, apparel at 25-32%) travel as industry consensus with no disclosed methodology anywhere, and should be treated as unverified.

The replacement is sitting in your own database. Order sequence is one of the few things every store already owns, and counting how many customers whose first order landed in a given month placed a second within 90, 180 and 365 days takes an afternoon. There is a real gap here worth saying out loud: no vendor publishes category-level cumulative-revenue cohort curves. The companies best placed to produce them sell the tooling instead.

The only version of this number that means anything is the one built from your own orders. Illustrative shape — the point is the grid, not the values.

Does a 5% increase in retention really raise profit 25-95%?

That is real research about a different industry from a different decade. The figure comes from Frederick Reichheld's work at Bain in the 1990s, and the underlying studies were on financial services — credit cards, insurance, banking — where the relationship is contractual, the marginal cost of servicing an existing account approaches zero and the churn curve behaves nothing like a DTC repeat curve.

It was also a range across industries rather than a law of retail. Applied to a store selling a $32 consumable, it asserts something the original never tested. Retention still matters enormously. The number simply does not transfer, and quoting it as if it were measured on ecommerce cohorts is the most common laundering in this list.

Is 3:1 LTV to CAC the right target?

It is SaaS folklore. David Skok set the ratio out in SaaS Metrics 2.0 for subscription software, and the assumptions underneath it are software-scale gross margins and contractual recurring revenue that renews unless somebody cancels. A store has neither: margins are lower, and every repeat order is a fresh decision rather than a renewal.

The clearer tell is that the ratio has no time bound. Three to one over what period? A DTC lifetime value quoted across 24 months and a SaaS one quoted across a contract term are different kinds of object. The vertical benchmark tables quoted to one decimal place are worse still. Follow them and you find articles citing articles, with no sample disclosed at any hop. Break-even ROAS is the number that actually constrains a first order, and it comes out of your own margin rather than somebody else's benchmark.

Do the ad platforms overstate performance?

Not uniformly, and the direction depends on the campaign type. Haus published the aggregate of 640 Meta incrementality experiments in July 2025: Meta under-reported its own incremental contribution by about 15% on average under click-only attribution, for DTC-only brands. That is not Meta's default, which also counts one-day view. In the same dataset, automated campaigns over-reported themselves by 12 percentage points relative to manual, and 58% of brands got a higher incremental ROAS out of manual.

Both directions live in one study. Anyone telling you platforms simply inflate their numbers is describing half of a finding, and the half they dropped is the one that would have made you spend more. MER, ROAS and what to steer on works through which measure answers which question.

Does hook rate predict sales?

There is no published evidence that it does, in either direction, and that absence is the finding. Nobody has released a correlation between hook rate and sales on a sample anyone can inspect. What circulates instead is the observation that a strong hook rate beside flat sales is one of the commonest patterns in a creative report, plus a mechanism that predicts it: a three-second play is triggered by anything arresting, and arresting is cheap. Your own account can settle it in an afternoon. Nobody else's can settle it for you.

The second problem is that hold rate has two competing denominators, so the same video produces two very different numbers depending on which tool you opened.

Definition in useDenominatorHold rate
ThruPlays ÷ three-second plays25,00030.0%
ThruPlays ÷ impressions100,0007.5%
One video with 100,000 impressions, 25,000 three-second plays and 7,500 fifteen-second ThruPlays. Illustrative arithmetic rather than measured data.

Both definitions are in active use by a large minority of tools each, which makes cross-tool hold-rate comparison meaningless unless somebody states the formula. The benchmark bands that get quoted alongside — roughly 25% average, 30% strong, 40% elite — have no published sample behind them and should be read as unverified. Hook rate and hold rate covers what each one does diagnose, which is real and useful and is not conversion.

Can you see a competitor's ad spend in the ad library?

Almost never, and the two disclosures are separate. Spend and impression ranges appear only for ads about politics or social issues. What delivery into the EU or the UK adds is total reach, not spend — and the UK is its own disclosure region rather than part of the EU, since Article 39 of the Digital Services Act does not apply there. The figures that do appear come banded rather than exact. For an ordinary US commercial ad there is no spend figure at all, and TikTok's Creative Center withholds exact spend and targeting too.

What the public record does contain is longevity, variant count and placement mix. How long an ad has been running is the closest thing to a performance signal visible from outside, on the reasoning that nobody keeps paying to run a losing ad. That is an inference rather than a measurement, and it fails on any brand with untidy account hygiene, which is most of them.

Does a one-page checkout convert better?

The reasons people actually give have nothing to do with page count. In Baymard's checkout survey, 40% point to extra costs — shipping, tax and fees — being too high, and 18% to being made to create an account. Read those shares with the survey's own scope in mind: it excludes the 42% who were only browsing, so they describe shoppers who meant to buy.

The measured detail is more useful than the argument. Baymard's September 2025 figures put average documented cart abandonment at 70.22%, synthesised across 50 studies. The average US checkout carries 23.48 form elements against an achievable 12-14, the average large site carries 39 usability issues and fixing them is worth 35.26% more conversions. None of that is about how many pages you split them across.

Is email really 30% of your revenue?

It is 30% by your email platform's attribution model, which is a different claim. Klaviyo announced on 7 August 2024 that its default would become five-day click and five-day open, cooperative last-touch inside Klaviyo only. An open with no click can claim the sale. Klaviyo documented the change publicly and the default is defensible for what it is built to do. The error happens downstream, when a dashboard figure gets read as incremental contribution.

What is genuinely well-measured is the comparison inside that convention, because both sides are graded the same way. Klaviyo's 2026 benchmarks, drawn from more than 183,000 customers, put flows at 5.3% of sends and about 41% of email revenue, with flows earning roughly 18x more per recipient than campaigns. That ratio holds. The absolute share of company revenue does not.

A number with no stated sample, no date and no method is not a weak statistic. It is a sentence with a decimal point in it.

Are return rates rising?

Probably not, and the honest answer is softer than either side of the argument. NRF and Happy Returns put US retail returns at 15.8% of sales in 2025, worth $849.9B, against 16.9% and $890B in 2024 — and NRF describes that as in line with the year before rather than as a fall. Note what the figure is: retailers were surveyed in October and asked what they expect, so it is a forecast, not a measurement. By this article's own checklist that is a mark against it. Online-specific returns are put at 19.3%, and about 9% of all returns are estimated to be fraudulent.

Measure20242025
Return rate, share of US retail sales16.9%15.8%
Value of merchandise returned$890B$849.9B
The direction most coverage gets backwards. NRF and Happy Returns, October 2025

The apparel-specific rates quoted alongside this, usually 24-35%, have no disclosed sample and are unverified. If returns are a material line in your margin — and on apparel they are the largest single one — the figure to model with is your own, which every store can measure exactly.

Is every 100ms of load time worth 8.4% more conversion?

The study is real and almost always misdescribed. Deloitte, with Google, published Milliseconds Make Millions in 2020 on data collected in 2019 across 37 retail and travel sites in Europe and the US. A 0.1s mobile improvement was associated with 8.4% higher retail conversion and 9.2% higher average order value.

Associated. The design was correlational rather than randomised, so it does not establish that making a slow site faster produces that lift. Fast sites differ from slow ones in many ways, including how much has been invested in everything else. It also pre-dates the spread of mobile-wallet checkout, which removed a large share of the form-filling that made slow mobile pages so expensive. No comparable replication has been published since.

Speed is still worth having, and there are thresholds worth hitting that don't rely on that study: largest contentful paint under 2.5s, interaction to next paint under 200ms since it replaced first input delay in March 2024 and cumulative layout shift under 0.1.

Does a rising frequency mean the ad is fatigued?

Frequency is a lagging measure. By the time average frequency has climbed enough to notice, delivery has been recycling the same people for days and the cost per acquisition is already moving.

The leading indicator is the first-time-impression ratio, the share of a day's impressions going to someone who has never seen that ad. It turns over before cost does, because exhaustion shows up in who is being reached before it shows up in what they do. Healthy prospecting is often quoted at 65-80% with anything below 50% treated as exhausted, and those bands are unverified: nobody has published a sample behind them. The mechanism is sound and the metric is available, so set your own thresholds from your own history rather than borrowing someone else's.

Do longer articles rank better?

No. Google's own documentation on creating helpful content puts it as a question to ask yourself — whether you are writing to a word count because you heard Google prefers one — and answers it in brackets: no, it does not. There is no length requirement anywhere in it.

The correlation behind the myth is real and misread. Pages that rank well are often long because thorough answers to complicated questions take words, not because length earns anything. Padding a good 600-word answer out to 1,800 produces a worse answer at the same position, and the ranking systems that used to be called helpful content became part of core ranking in March 2024, so there is no separate switch to game.

The same misreading drives the anxiety about AI-written content, which is not penalised as such either. What Google acts on is a narrower and more specific set of behaviours, and they are worth knowing precisely rather than by rumour — we went through what Google actually penalises in AI-written content against the current spam policies.

What have we got wrong?

A fair question to ask of anyone writing this article, and there is a live example in our own codebase.

A note attached to one of our database migrations recorded a match rate for joining two internal tables: how many brands in one dataset could be matched into the other. It read like a count. It wasn't. It was the query planner's estimate of distinct values, a statistic Postgres keeps for choosing execution plans and which is allowed to be wrong by orders of magnitude. Somebody read it off, wrote it down as a measurement and moved on.

It didn't reproduce. Re-measuring it properly meant taking a deliberate sample, because counting distinct values across that table exceeds the database's statement timeout, and the answer came out materially different. The sampling method now sits next to the figure so the next person knows what they are reading.

The wrong number was not really the failure. The failure was that nothing recorded how it had been obtained, so it survived until an unrelated piece of work tripped over it. That is the same defect as every claim above, running at a smaller scale, and it is why three figures in this article are labelled unverified rather than quietly dropped, and why the two illustrative tables say so in their captions.

How do you check a benchmark before repeating it?

Four minutes and this list will catch most of them. The single highest-yield move is clicking the citation and then clicking that page's citation, which is where the loop usually reveals itself.

Checklist

0 / 8

Follow a widely repeated benchmark far enough and the citations close on themselves. Nothing in the loop is a study.

None of this argues for measuring less or for distrusting every number you meet. It argues for a short habit: before a figure enters a plan, know who produced it, when, on whom and how. The ones that survive that are worth a great deal more than the ones that never had to.

Sources

  1. National Retail Federation and Happy Returns, Consumers expected to return nearly $850 billion in merchandise in 2025. October 2025
  2. Baymard Institute, Cart abandonment rate statistics, synthesised across 50 studies. September 2025
  3. Haus, The Meta report: lessons from 640 incrementality experiments. July 2025
  4. Klaviyo, Attribution model updates to reporting, the five-day click and open default. August 2024
  5. Deloitte with Google, Milliseconds Make Millions, 37 retail and travel sites, data collected 2019. Published 2020
  6. Google Search Central, Creating helpful, reliable, people-first content. Accessed August 2026

RIVYL

Every ad your competitors are running, in one place

Search competitor creative across Meta, TikTok and TikTok Shop, sorted by how long each ad has run.