Insight
Product data is the e-commerce growth lever nobody budgets for
Advertising and site design get the budget. Attribute completeness quietly determines how well both of them can possibly work, and it is measurable today.
One dataset, five dependent systems
Product data is not a catalogue chore. The same attributes drive on-site search, faceted filtering, recommendations, shopping feeds, marketplace listings and the accuracy of what a customer expects when the parcel arrives. Every one of those systems is capped by the quality of the same underlying records.
This is why the symptoms are so scattered and so rarely traced back. On-site search returns nothing for a term customers actually use. Filters miss products that should qualify. Shopping ads underperform for reasons the platform describes only as relevance. Returns run high in one category. Each gets investigated separately, by different people, and the common cause never surfaces.
It is also why the return on fixing it is unusually broad. An improvement to attribute completeness does not improve one channel; it raises the ceiling on several at once, including the paid ones you are already funding.
Measure completeness before doing anything else
The diagnostic is simple and most retailers have never run it. For each category, what share of products has every attribute a customer might filter or search by, expressed as a percentage. Not whether the field exists — whether it is populated, consistently, with values from a controlled list rather than free text.
The results are reliably worse than expected and they are not evenly distributed. Newer categories are usually better than older ones, own-brand better than supplier-supplied, high-volume better than long tail. That distribution is what makes the fix affordable, because it points at where the effort actually needs to go.
Pair it with a second measurement: the top on-site searches returning no results, or returning results the customer did not click. That list is a direct statement of what your customers want and your data cannot express, and it is usually available in an afternoon from tooling you already have.
- Attribute completeness by category
- The share of products with every filterable attribute populated from a controlled list, not free text.
- Zero-result searches
- A direct list of what customers ask for that your data cannot answer. Available today, rarely looked at.
- Feed disapprovals and warnings
- The merchant platform is already telling you which records are inadequate. Most retailers never read it in detail.
- Returns by reason and category
- Where "not as described" concentrates is where the product content is failing before purchase.
Free text is the underlying problem
Most attribute problems are not missing values; they are inconsistent ones. Colour recorded as navy, dark blue, midnight and NVY across four suppliers. Dimensions in centimetres in one record and millimetres in the next. Material described in a sentence rather than selected from a list.
A human reading a product page copes with all of this. A filter, a feed and a search index do not. They match values, so four spellings of the same colour become four values that each match a quarter of the products they should, and the customer filtering by blue is shown a fraction of your blue products and concludes you do not have what they want.
The fix is a controlled vocabulary per attribute, with supplier values mapped onto it on ingest rather than corrected later. This is boring, it is not a project anyone gets excited about, and it is the single change with the widest downstream effect in most catalogues.
Supplier data is the source of most of it
Retailers that do not manufacture are largely at the mercy of supplier feeds, and those arrive in wildly varying quality, format and completeness. The instinct is to accept them and clean up afterwards, which means cleaning forever because the next update overwrites the work.
The alternative is to treat ingest as the control point: define the attributes you require, map each supplier's values onto your vocabulary once, and reject or quarantine records that do not meet the minimum rather than publishing them incomplete. That is a harder conversation with a supplier and it is the only version that stays fixed.
Where a supplier genuinely cannot provide an attribute, the decision is whether to enrich it yourself or accept that those products will not appear in filtered results. Both are legitimate; what is not legitimate is publishing them and assuming the gap does not cost anything, because the zero-result search list will show that it does.
Where AI genuinely helps, and where it quietly hurts
Attribute extraction from supplier descriptions and images is one of the better applications of AI in retail. Pulling material, dimensions, fit or compatibility out of unstructured supplier text at scale is exactly the sort of task that is tedious for people and tractable for a model, and it can populate a large catalogue in a fraction of the time.
The condition is that extracted values are treated as provisional. They should be marked as machine-derived, sampled for accuracy by category, and corrected where the error rate is unacceptable — because an attribute that is confidently wrong is worse than one that is missing. A missing value excludes a product from a filter; a wrong value puts it in front of a customer who will return it.
Generated product descriptions are the other common application and deserve more caution. At scale they produce near-identical text across a catalogue, which adds nothing a customer or a search engine values, and where the copy makes claims the model inferred rather than read, it creates a consumer law exposure rather than a content asset.
The returns connection
Returns are frequently framed as a logistics cost and are substantially a content problem. When "not as described", "wrong size" or "not what I expected" concentrate in a category, the product content failed before the purchase rather than the product failing after it.
This is where the arithmetic gets persuasive, because return costs are large, measured and already visible to finance in a way that attribute completeness is not. Connecting the two — this category has the worst completeness and the highest not-as-described rate — turns a data quality project into a margin conversation.
The specific fixes are usually mundane: accurate dimensions rather than nominal sizes, real photographs of the actual variant, fit or compatibility stated explicitly, and materials described consistently. None of it is sophisticated and all of it is measurable against the return reason codes you already collect.
Feeds are where the gaps become expensive
Shopping and marketplace feeds are unforgiving in a useful way: they tell you exactly which records are inadequate, in detail, continuously. Most retailers look at the summary count and never at the itemised warnings, which is discarding the best free data quality audit available.
Feed performance is also where poor attributes cost money most directly, because you are paying for impressions that match badly. A product with incomplete or inconsistent attributes competes for the wrong queries, converts worse, and drags the account's performance in ways that get diagnosed as bidding problems.
The practical sequence is to fix the categories with both high spend and low completeness first. That intersection is where data work has an immediate, attributable effect on money already being spent, which also makes it the easiest phase to get funded.
A sequence that pays for itself as it goes
Measure completeness by category and pull the zero-result search list. Cross-reference with advertising spend and return rates. Fix the intersection first — high spend, high returns, low completeness — because that is where the effect is fastest and most attributable.
Then establish the controlled vocabularies and move the cleaning to ingest, so the improvement holds rather than degrading with the next supplier update. Then extend coverage down the long tail, where AI extraction earns its place because the manual cost per product stops being justifiable.
Throughout, keep measuring the same completeness figure, because it is the number that connects the work to every downstream symptom. It is unglamorous to report and it is the closest thing e-commerce has to a single leading indicator of how well everything else can possibly perform.
Questions
How do we know if we have a product data problem?
Run two measurements this week: attribute completeness by category, and the top on-site searches that return nothing. Both are available from tooling you already have, and the results are reliably worse than expected.
Can AI fill in our missing attributes?
Extraction from supplier descriptions and images is one of the better AI applications in retail, provided extracted values are marked as machine-derived and sampled for accuracy. A confidently wrong attribute is worse than a missing one, because it puts products in front of customers who will return them.
Should we generate product descriptions with AI?
With caution. At scale it produces near-identical text that adds nothing a customer values, and where the copy states things the model inferred rather than read, it creates a consumer law exposure. Attributes are a much better use of the same effort.
Our supplier data is poor. What can we do?
Move the control point to ingest: define required attributes, map supplier values onto your vocabulary once, and quarantine records that fall short rather than publishing them incomplete. Cleaning after the fact means cleaning again after every update.
Where does this show up in returns?
In the "not as described" and "wrong size" reason codes, concentrated in the categories with the worst attribute quality. That connection turns a data project into a margin conversation, because return costs are already visible to finance.
Where should we start if the catalogue is large?
At the intersection of high advertising spend, high return rate and low completeness. That is where the effect is fastest and most attributable to money you are already spending, which is also what makes the next phase straightforward to fund.
How often should attribute completeness be measured?
Monthly, by category, reported alongside commercial metrics. It degrades continuously as new products arrive from suppliers, so a one-off cleanup with no ongoing measurement returns to roughly where it started within a year.
Where this sits in what we do
This article covers one decision inside a wider engagement. The solution page sets out how that engagement runs, what it includes and what it costs to find out.
- E-commerce Growth System — Store, brand, advertising, lifecycle communication and order automation as one system — with local payment methods and the margin arithmetic checked before launch.
- AI for custom customer care: what it can answer and what it must not
- When to build software instead of buying it
- Why more traffic rarely fixes a pipeline problem
- Lead scoring that sales will actually use
- All insight articles
Suspect your catalogue is capping your channels?
We will measure attribute completeness against your spend and return rates, and show you which categories are costing you most before proposing any work.
Get in touch