Search & Discovery

B2B Search Is a Different Problem

Most search relevance advice is written for consumer retail. Applied to a wholesale catalogue it produces a search box that is technically impressive and commercially wrong.

A consumer opens a retail app to browse. They are open to persuasion, they may not know the exact product, and discovery is genuinely valuable to them and to the retailer.

A restaurant owner opening a wholesale app at 6am is not browsing. They are restocking the same forty lines they order every week, under time pressure, before service. Every feature that widens their consideration set is friction.

In B2B, the best search result is frequently the one the customer has already bought eleven times.

What changes

Repeat purchase dominates. Prior order history is the strongest relevance signal available, and it is stronger than almost anything you can infer from the query text. A search ranking that ignores what this specific buyer has bought before is discarding its best feature.

Pack size and unit are not attributes, they are the product. "Rice 5kg" and "Rice 25kg" are not variants to be collapsed for a tidy results page. Ordering the wrong one is a real operational problem for the customer — wrong stock, wrong storage, wrong cost. Collapsing them is not a UX simplification; it is a correctness bug.

Substitution is commercially loaded. Consumer search can freely suggest alternatives. In wholesale, an out-of-stock line means the buyer has a gap in service tonight. Substitutions must be genuinely equivalent in unit and grade, or they waste the buyer's time and erode trust in the whole result set.

Zero results are worse. A consumer who finds nothing shrugs. A business buyer who cannot find a line they order weekly assumes you have stopped carrying it and calls a competitor.

Where language models genuinely help

Not, mostly, in ranking. In understanding the query.

Business buyers search the way they talk in their own operation: trade names, abbreviations, local terms, misspellings under time pressure, sometimes in mixed script. A model that can map that to catalogue vocabulary — grounded in your actual catalogue, not general knowledge — closes a gap that synonym lists never fully close, because the tail is genuinely endless.

The second place they help is catalogue quality. High-SKU wholesale catalogues have inconsistent, sparse, sometimes contradictory product text, and search quality is capped by it. Improving attribute coverage lifts relevance more reliably than tuning the ranker on top of poor data.

Order the work correctly

Query understanding and catalogue data quality beat ranking sophistication in this domain, and they beat it consistently. A well-tuned ranker over inconsistent product data is an expensive way to be wrong faster.

Measuring it

Click-through rate is a poor primary metric here. A buyer clicking three results before finding the right pack size is not engagement, it is failure with extra steps.

Better signals: time from search to basket, reorder completion rate, zero-result rate on terms with prior purchase history — that last one is close to a pure defect count — and searches per completed order, where lower is better and the opposite of what consumer dashboards celebrate.

The strategic point generalises past search. Patterns imported from consumer commerce mostly do not survive contact with business buying, because the buyer's job is different. They are not shopping. They are restocking, against a clock, with real consequences for getting it wrong.

← All writing Next: Performance Is an Org Chart Problem →