1
2
3
4
5
6
7
8
9
10
11
12

Why Industry Averages Mislead (and How to Use Them Anyway)

Industry averages mislead through four distortions: mix, shape, definition and time. Three jobs they are still genuinely good for, and one they are not.

The CROBenchmark Team
September 11, 2026

Spot your biggest conversion leaks in 15 minutes.

Check best practices, accessibility, data hygiene, and customer sentiment - then compare results with competitors and unlock tailored A/B testing ideas.

Why Industry Averages Mislead (and How to Use Them Anyway)
Quick Answer

Industry averages mislead in four specific ways, and knowing which one is operating tells you how much to discount a given figure. Mix: the average moves when the composition of the sample changes, so it can fall while every store in it improves. Shape: conversion data is skewed, so an arithmetic mean sits above most of the stores it describes. Definition: sources disagree about what counts as a session and a conversion, which explains more of the gap between published figures than anything else. Time: an average is historical, and the window is rarely stated. What survives all four is still useful for bounding a target, triaging whether to investigate, and explaining a number to somebody outside the team.

Key Takeaways
  • A category average can fall while every store inside it improves, purely from a change in who is in the sample.
  • Conversion data is skewed, so a mean describes a position most stores in the group sit below. Use the median.
  • Most of the gap between two published benchmarks is definitional, not real.
  • The distance between your rate and the average is not a quantity of available improvement.
  • Benchmarks are good for bounding, triage and communication, and bad for deciding what to change.

Industry averages are the most widely used and least reliable number in eCommerce reporting. They get quoted in board decks, used to set targets, and treated as a standard a store is passing or failing, and almost none of that survives contact with how the figure was built. The useful move is not to discard them. It is to know precisely which distortion is operating, because each one has a different size and a different fix. Last updated: September 2026.

Omniconvert has measured conversion behaviour across the CROBenchmark dataset of 7,000+ websites in 15+ industries, against 248+ audit criteria, over 13 years in eCommerce, and has run Omniconvert Explore across 70,000+ experiments. The finding that matters here is the same one that makes benchmark publishing awkward: the spread within a category is consistently wider than the distance between categories, which means the category label explains far less about a store than the label implies.

This article is about the statistics rather than the decision. If what you actually need is which signal to act on when a benchmark and your own history disagree, our companion piece on benchmarks versus your own baseline answers that directly, and for the category figures themselves, start with a good conversion rate.

Further detail sits in two stores with the same CVR aren't equal.

What an industry average actually is

A single number summarising a group of stores that have almost nothing in common except a category label. It is an accurate description of that group and a poor description of any member of it, and the distance between those two statements is where every benchmark mistake lives.

Nothing about an average is dishonest. The arithmetic is correct, and a competently published benchmark is measuring real stores. The problem is one of transfer: the figure describes a population, and it gets used to make a claim about an individual.

That transfer is only valid when members of the population resemble each other. For height in a national population it works well. For conversion rate in an eCommerce category it works badly, because two stores filed under the same label can differ by a factor of several while both being run competently, and the reasons have little to do with site quality.

So the question is not whether the average is right. It is how much information it carries about you, and the answer depends on four distortions that are worth naming separately, because they have different magnitudes and different remedies.

Distortion 1: mix, or the average that moves when nothing moves

An aggregate changes when the composition of the group changes, independently of any member changing. A category average can fall while every store in it improves, simply because more low-converting stores entered the sample. This is the largest distortion and the least visible.

Mix effects are counterintuitive enough that they routinely survive review by careful people. The mechanism is simple: if a benchmark's sample grows to include many small, young, paid-traffic-heavy stores, the average falls. Every existing store may have improved. The published figure still went down.

This matters in both directions. A benchmark that rose does not mean the field got better, and your position relative to it can shift by a meaningful margin without anything happening on your site at all. Year-on-year benchmark comparisons are therefore close to meaningless unless the source states that the sample is held constant, which it almost never does.

The same effect operates inside your own store. A sitewide conversion rate is itself an average across device, source and category, so it moves whenever your traffic composition moves. A successful campaign that brings in cold traffic reliably lowers a sitewide rate while raising revenue, and teams reporting on the rate rather than on the revenue have been known to switch off the campaign.

Distortion 2: shape, or the mean that describes nobody

Conversion data is skewed, not symmetric: a small number of very strong performers pull the arithmetic mean upward, so the average sits above the majority of stores it summarises. Most stores comparing themselves to a published mean are comparing themselves to a position most of the sample is also below.

When a distribution is symmetric, the mean is a fair summary. Conversion rates are not symmetric. The floor is zero and the ceiling is soft, and a handful of stores with unusual traffic, subscription mechanics or extremely narrow catalogues sit far above the bulk of the field. Those stores drag the mean.

The practical consequence is that the typical store is below average, which sounds like a paradox and is simply what skew does. A team benchmarking against a published mean and finding itself below it has learned very little, because so is most of the sample.

The fix is cheap when it is available: use the median. Where a source publishes only a mean, treat the figure as optimistic and do not build a target on it. Where a source publishes percentiles, use those instead of either, because a range tells you about the spread and a single number never can.

Distortion 3: definition, or two sources measuring different things

Most of the gap between two published benchmarks is definitional rather than real. What counted as a session, whether bots were excluded, whether a conversion meant an order or a transaction event, and which devices were included, vary source to source and are rarely stated.

This is the distortion most easily fixed and most often skipped, because checking it is tedious. Two benchmarks for the same category, published in the same year, can differ substantially while both are correct, purely on measurement choices.

The usual suspects are few and worth memorising. Session definition and timeout. Bot and internal traffic exclusion, or the absence of it. Whether a conversion counts at order placement or at payment confirmation. Whether mobile app traffic is in scope. Whether returns are netted off. Each is defensible, and each moves the number.

This is also why comparing your figure to a published one requires reconciling your own definitions first. A store measuring conversions at payment confirmation against a benchmark measuring them at order placement will read low by a margin that has nothing to do with performance.

Distortion 4: time, or the average that already expired

Every benchmark is historical, the window is often unstated, and the periods that get measured are not neutral. A figure drawn across a quarter containing a peak trading season describes a different world from one drawn across a quiet one, and neither says so.

Seasonality alone moves category conversion rates enough to swamp most of what a store could achieve through optimisation in the same period. A benchmark that spans a peak looks strong; one that avoids it looks weak. When the window is not disclosed, the reader has no way to correct for this.

There is a slower version of the same problem. Published benchmarks lag their measurement period, sometimes by many months, and the market conditions they describe may no longer hold. Using a figure without knowing its vintage is a common way to set a target against a market that has moved.

Source: Omniconvert, the four distortions by how much they move a benchmark and what to do about each
Distortion What causes it Typical size What to do
Mix The composition of the sample changes Large, and invisible Never compare a benchmark to itself year on year
Shape Skewed distribution pulls the mean up Moderate and consistent Use the median, or percentiles if offered
Definition Sources count sessions and conversions differently Often the largest single gap Reconcile definitions before comparing
Time Unstated window, seasonality, publication lag Large in seasonal categories Check the window and the vintage, or discard
All four together Compounding, in unknown directions Unquantifiable Treat the figure as a range, never a target

The bottom row is the honest summary. Because the four distortions compound in directions you cannot determine from a published figure, the responsible reading of any single benchmark number is as a rough region rather than a point.

How to use industry averages anyway

Three jobs survive all four distortions: bounding a target so it is not impossible, triaging whether a number deserves investigation time, and giving a shared external reference when explaining a figure to someone outside the team. None of the three requires the average to be precise.

Bounding. Knowing roughly where a category tops out stops a team committing to a target that nothing in the data supports. This works because it needs only the order of magnitude, which survives every distortion above. A target set at four times the category range is wrong regardless of how the benchmark was built.

Triage. A benchmark is a reasonable trigger for whether to spend a day looking. If a figure is wildly outside the plausible region, something is worth checking, and frequently what is worth checking is your own measurement rather than your performance. Implausible numbers are more often tracking faults than business events.

Communication. An external reference is genuinely valuable when explaining a result to a board, an investor or a new stakeholder who has no context for what normal looks like. Used honestly, with the caveats attached, it makes a conversation possible that would otherwise be an argument about vibes.

What none of the three do is tell you what to change. That requires information about your store, which an average does not contain by construction. The widely quoted Baymard Institute figure of roughly 70% average cart abandonment illustrates the point neatly: it is a real, carefully produced average, it is useful for knowing that abandonment is normal and large, and it cannot tell any individual store which of its own checkout steps is losing people.

What to do this week

Three changes, none of which requires new data. Stop reporting the gap to an average as an opportunity, write your own definitions down, and replace the benchmark comparison in your weekly report with your own trend. The first one saves the most money.
  • Delete the gap-to-average line from your reporting. It is the most misread number in a CRO deck, because it looks like available upside and is mostly composition.
  • Write down your own definitions. Session, conversion, exclusions, device scope. One paragraph. Without it no external comparison you make is valid.
  • Lead the weekly report with your own trend. Movement you cannot explain is the real alarm; distance from an average is not.
  • Use the benchmark for bounding only. Set the plausible range at the start of planning, then put it away.
  • Check the vintage of any figure you quote. If the window is not stated, say so when you quote it, or do not quote it.

Where the underlying difficulty is that order, traffic and cost data never sit together long enough to build a defensible baseline of your own, that is a data problem rather than a benchmarking one. Nexus by Omniconvert is an AI for eCommerce growth engine that unifies commerce data, prioritises experiments by True Profit, and generates campaigns and creative you approve before they go live. Bain and Company's retention work with Fred Reichheld makes the adjacent case that the compounding advantages are the ones worth measuring properly, and those are precisely the ones a category average cannot see.

Once you know which distortion is biting, the next question is what to fix, and that is an audit rather than a comparison. Our neighbours at croaudit.marketing cover how to audit it step by step, and Omniconvert Explore is the CRO platform for validating the change afterwards, with A/B and multivariate testing, on-site surveys, heatmaps and session recordings.

FAQ: industry averages and benchmark data

Why do industry averages vary so much between sources?

Because each source aggregates a different set of stores and defines the metric differently, and both differences are usually undisclosed. One dataset may be weighted toward large paid-traffic retailers and another toward small email-led brands, which produces genuinely different averages from equally honest measurement. Before comparing two published figures, check what counted as a session and what counted as a conversion in each, because that single definitional difference explains more of the gap than anything else.

What is a mix effect, and why does it matter for benchmarks?

A mix effect is when an aggregate number moves because the composition of the group changed, not because any member of the group changed. A category average can fall while every single store in it improves, simply because more low-converting stores joined the sample. This is why a benchmark moving year on year tells you almost nothing about whether stores are getting better, and why your own position relative to it can shift without you doing anything at all.

Should I use the median instead of the average?

Yes, wherever it is offered, because conversion data is skewed rather than symmetric. A small number of very high performers pull an arithmetic mean upward, so the average describes a position that most stores in the group are below. The median is the figure that answers the question people think they are asking, which is what a typical store in this category looks like. If a published benchmark does not say which it used, assume the mean and treat it as optimistic.

Are industry averages useful for anything at all?

Three things, and they are genuinely useful. Bounding tells you the plausible range so you do not set an impossible target. Triage tells you whether a number is unusual enough to justify investigation time. Communication gives a shared external reference when you need to explain to somebody outside the team why a figure is or is not alarming. What an average cannot do is tell you what to change, because it contains no information about your store.

How far from the average should I be before I worry?

Distance from an average is the wrong trigger, because the spread inside any category is wide enough that being well below it can be entirely normal for your traffic mix. A better trigger is movement in your own trend that you cannot explain. A store sitting far below a category average with stable, profitable unit economics has no problem. A store sitting comfortably on the average while its own rate declines month over month has a real one.

What is the most common mistake people make with benchmark data?

Treating the gap between their number and the average as a quantity of available improvement. It is not a gap that can be closed, because much of it is explained by traffic mix, price point and category composition rather than by anything on the site. Teams that plan against that gap set targets they cannot reach, and then read a normal result as a failure, which is a reliable way to abandon a programme that was working.

The bottom line

An industry average is a correct answer to a question almost nobody is asking. It describes a population accurately and an individual store badly, and the four distortions explain how badly: mix moves it when nothing moved, shape puts it above most of the field, definition makes two honest sources disagree, and time hides which world it was measured in. Knowing that is not a reason to throw benchmarks away. It is a reason to give them the three jobs they can actually do, bounding a target, triaging whether to investigate, and explaining a number to somebody who needs context, and to stop giving them the job they cannot do, which is telling you what to change. The single most valuable thing you can do this week is delete the gap-to-average line from your reporting. It is the number most likely to be read as available upside, and it is mostly a description of who else is in the sample.