Book a demo
The Notebook

The Notebook · Essay 02

We read 20 million products. Here is how catalogs are built.

A structural census of the apparel catalog — the vantage point that only exists after you’ve normalized twenty million products into a single vocabulary.

We don’t have an opinion on fashion. We have a very large, very specific view of its plumbing. Our engine has ingested roughly 20 million products across 82 catalogs, and normalized every one of them under a single taxonomy: around 400 attributes, resolving to more than 16,000 distinct attribute values. That last number is the interesting one. It’s not a count of clothes. It’s a count of the distinctions the apparel world actually makes — every neckline it names, every closure it recognizes, every weight class it distinguishes — collapsed into one consistent vocabulary.

~20M
products ingested
82
catalogs normalized
~400
attributes in the taxonomy
16,000+
distinct attribute values

One clarification, because the numbers invite it: these 82 are catalogs the engine has read and normalized to learn the structure — the study corpus, not a customer list. Which brands run Bunsar in production — on a live retail catalogue, since 2019 — is a separate and smaller matter. Read nothing here as a count of customers. And nothing here is scraped from a competitor.

Normalize at that scale and the catalog stops being a collection of products. It becomes a structure. Here is some of what that structure looks like, stated the only way we’re willing to state it: brand-blind, in aggregate, as engineering — never “brand X did Y,” and never a guess about tomorrow.

A catalog is a taxonomy before it’s a collection

The first thing you learn normalizing at this scale is that the products vary far more than the shape of the catalog does. The 16,000-plus values sound like enormous diversity, and at the item level they are. But the skeleton — the set of attributes every garment has to declare — barely moves from catalog to catalog. Two catalogs from completely different corners of the market, once flattened into attributes, turn out to declare the same things: the same silhouette families, the same closures, the same layers. That stability is what makes 82 different catalogs comparable at all.

Catalogs breathe, and they breathe on a schedule

Here is the structural regularity we find most telling, because it’s reproducible across multiple brands and multiple years, which is the only reason we’ll publish it.

Schematic — no brand, no scale, no data points An abstract schematic of catalog respiration: intake climbs to a peak, then falls to a trough falls to a trough climbs to a peak the year → share of the catalog
The shape of the regularity, not a chart of anyone’s catalog: as a share of the assortment, outerwear intake climbs toward a peak and falls to a trough, on a cadence you only see after averaging many catalogs across years. Drawn as a schematic on purpose — no axes, no numbers, no brand.

An apparel catalog is not a static shelf. It breathes — its composition expands and contracts across the year in a pattern dictated by how the category is built, not by anyone’s taste. Outerwear is the clearest example. As a share of the catalog, its intake climbs into autumn and falls to its low in summer. It does that with a regularity no single catalog reveals — you only see it after averaging many. This isn’t a trend and it isn’t a forecast — it’s the mechanical cadence of how apparel assortments are constructed against the seasons. Coats are built for cold; the catalog fills with them before the cold and empties of them after. Read one catalog and it looks like a choice. Read 82 across years and it’s plainly a structure — the industry’s respiration, visible only in aggregate.

We’re deliberately not telling you which catalogs, or how much, or what it means for next season. The first two would be a leak. The last would be a forecast. We do neither. What we’re showing you is that the cadence is real, reproducible, and structural.

Why only a census can see this

Any single brand can see its own catalog breathe. What a single brand cannot see is that the breathing is a general property of apparel catalogs rather than a fact about itself. That distinction — self-fact versus industry-structure — only becomes visible after you’ve normalized many catalogs into one vocabulary and looked at them together. It’s the difference between watching one tide and understanding the moon.

That’s the whole argument for a census. Not to rank anyone, not to expose anyone, and emphatically not to predict anyone — but to describe the shared machinery underneath, in language every catalog would recognize as true of itself.

We read twenty million products so that the structure would hold still long enough to be described. This is the beginning of that description. There’s a great deal more of it in that corpus, and all of it obeys the same rule: what the engine measures, and how the industry is built — never what’s in style.