Menu
Request a storefront

LinkToLooks · Guides

Research desk · Data study

Every listing has a price. Only 72% say which product it is

The fields that sell an item are always present. The fields that identify it are optional — and a third of the catalogue leaves them out.

Some links here are affiliate links — if you buy, the retailer may pay us a commission at no extra cost to you. We have earned $0 from them so far. Here’s how it works.

Part of: Research

Across 5,502 live affiliate listings pulled on 3 September 2026, every row carried a price, a link, an availability flag and a description — and only 72.0% carried a product identifier, 50.3% a manufacturer part number and 43.5% a variant group id. Method: 14 unrelated product keywords read at five catalogue depths through the CJ shopping-product API, de-duplicated on retailer, title and destination URL.

The pattern is not random missingness. The fields that help you sell an item are always there. The fields that let a machine work out which item it is are optional, and roughly a third of the catalogue leaves them out.

14 carat gold diamond ring photographed against white for a retailer product listing
This listing has a price, a photograph and a description. Whether it also says which exact ring it is depends entirely on which retailer published it. Image: Poshmark product listing
Zero-earnings disclosure: LinkToLooks holds no affiliate network approvals and has earned $0 in commission to date. We publish no income screenshots, no creator case studies and no testimonials, because we have none. Every figure here is either computed from live data during this run and dated, or quoted from a named source, or flagged as something we could not verify.

Field completeness, all 20 fields

Every field the API will return for a shopping product, ordered by how often it is populated. The middle column is the question a shopper or a machine would be asking.

FieldThe question it answersPresent in 5,502 listings
availabilityIs it in stock?100.0%
conditionNew or used?100.0%
lastUpdatedWhen was this row written?100.0%
targetCountryWhich market is this for?100.0%
priceWhat does it cost?100.0%
linkWhere do I buy it?100.0%
descriptionWhat is it?100.0%
brandWho makes it?99.9%
imageLinkWhat does it look like?99.8%
productTypeWhat kind of thing is it?98.9%
gtinWhich exact product is it?72.0%
genderWho is it for?64.2%
colorWhat colour?58.5%
sizeWhat size?55.3%
mpnThe manufacturer's own code50.3%
ageGroupAdult or child?49.3%
itemGroupIdWhich listings are the same product?43.5%
salePriceIs it reduced?40.8%
materialWhat is it made of?34.1%
additionalImageLinkAny other photographs?8.4%

The cliff, and where it falls

Seven fields are at 100%. Three more are above 98%. Then it falls off, and the place it falls is the interesting part.

Everything above the cliff is commercial. Price, link, availability, description, brand, image: the fields without which a listing cannot function as an offer. A retailer that omitted any of them would be publishing something unsellable, so nobody does.

Everything below the cliff is identifying. The GTIN — the barcode number that says which exact product this is — is present on 72.0% of listings. The manufacturer part number, 50.3%. The variant group id that says which rows are the same product in different sizes, 43.5%. What it is made of, 34.1%.

None of that is anyone behaving badly. A feed exists to be matched against a shopper’s search, and a search does not use a GTIN. The fields are optional because for the retailer’s own purpose they genuinely are.

The consequence lands on everyone downstream. When roughly one listing in four carries no identifier, matching a photograph or a product name to a specific SKU stops being a lookup and becomes an inference. That is precisely why a tool like ours — which reads a photograph and tries to name the product in it — sometimes returns “similar” rather than an exact match. The number in that table is the honest explanation for our own product behaviour, and we would rather publish it than describe the limitation vaguely.

Keep reading

Completeness is a retailer decision, not a spectrum

Here is the result we did not expect. Split the same sample by retailer and the averages dissolve: almost every cell is near 0% or near 100%. There is barely any middle.

RetailerListingsHas GTINHas MPNHas sizeHas material
Wayfair North America1,80478%100%93%83%
OnBuy.com1,065100%2%2%0%
OnBuy EU935100%0%0%0%
Poshmark5131%0%100%0%
Wayfair UK31381%100%99%66%
SHEIN2230%100%82%9%
JD.com1270%0%0%0%
Factcool Europe80100%100%100%0%
Perigold6442%100%91%61%

OnBuy publishes a GTIN on 99.7% of its listings and a part number on 2.0%. Poshmark publishes a size on 100% and a GTIN on 1.0%. SHEIN publishes a part number on 100% and a GTIN on 0.0%. Wayfair North America publishes almost everything.

These are not different levels of diligence. They are different export configurations, decided once, applied to an entire catalogue. Which means “how complete is affiliate feed data” has no useful average answer — it is a question about which retailers happen to be in your sample, and any single figure (including our 72.0%) is really a statement about catalogue mix.

The practical form of that, for anyone building on feeds: test per retailer, not per field. A retailer that gives you one identifying field will usually give you all of them, and one that gives you none will never start.

Cultured freshwater pearl necklace photographed for a marketplace product listing
One retailer's whole catalogue carries identifiers and another's carries none. Nothing in the listing tells you which you are looking at. Image: Poshmark product listing

The fields that exist but do not agree with themselves

Presence is not the only way a field fails. Two of the fields at 100% coverage are not normalised, which is a quieter problem because nothing looks missing.

FieldDistinct values seenThe values
availability4in stock (5,262), out of stock (135), in_stock (104), out_of_stock (1)
condition5New (5,056), Used (423), new (21), refurbished (1), NEW (1)

The same fact is written four ways in one field and five in the other, differing only in case and in whether the space is an underscore. Anyone filtering on availability == "in stock" silently drops 104 in-stock listings. Anyone counting new versus used by exact match on New misses 22 rows.

A hundred per cent coverage of a field nobody normalised is a different kind of incompleteness, and it is the kind that produces wrong numbers rather than missing ones.

How old the data behind a link is

Every row carries a lastUpdated timestamp, and reading it gives the shelf life of the whole catalogue.

Age of the feed rowListingsShare of 5,502
Median7 days
Middle half of the distribution3 to 25 days
Older than 30 days1,28123.3%
Older than 90 days4237.7%
Older than a year2835.1%
Oldest row in the sample2,665 days — about 7 years

The median is a week, which is better than the reputation feeds have. The tail is not: 283 listings — 5.1% — were last written more than a year ago and are still being served as live in-stock inventory today.

One row in the sample carried a timestamp one day in the future. We have not tried to explain it and we have not excluded it; it is one row in 5,502 and it belongs in the record rather than in a footnote we quietly dropped.

What this data cannot show

Method, and a correction to our own earlier note

Frequently asked

How complete is affiliate product feed data?

Commercially complete and structurally patchy. Across 5,502 live listings on 3 September 2026, price, link, availability, condition and description were present on 100%, while a product identifier appeared on 72.0%, a manufacturer part number on 50.3% and a variant group id on 43.5%.

Which fields are most often missing?

Additional images (8.4%), material (34.1%), variant group id (43.5%) and sale price (40.8%). The identifying fields, in other words, not the selling ones.

Why does a missing GTIN matter?

Without an identifier, matching a product photograph or name to a specific listing is inference rather than lookup. It is the difference between naming an exact item and offering a similar one.

Is feed data stale?

Less than its reputation at the median and worse at the tail. The median row was 7 days old, the middle half spanned 3 to 25 days, and 5.1% of rows had not been updated in over a year while still being served as live in-stock inventory.

Does completeness vary by retailer?

Almost entirely. Every retailer in the sample published each field on close to 0% or close to 100% of its listings, so an average across retailers describes catalogue mix rather than data quality.

More from the research desk

Last verified 3 September 2026 against 5,502 live affiliate listings pulled across 14 product categories and read field by field

How this guide was made. We research and draft these guides with AI, then a person checks every price, link and factual claim against the source before it publishes. We work this way because it lets us re-verify prices across hundreds of guides in a day, which is what keeps the numbers here current; it does not decide what we recommend. Anything we could not verify is labelled as unverified rather than filled in.

LinkToLooks identification desk — we identify objects out of real photographs and check every listing by eye before we link it. Empty beats wrong. About the desk · published 3 September 2026.