Engagement snapshot
- Mandate
- Increase catalog throughput without letting noisy supplier data degrade product quality.
- Timeline
- 9 weeks to launch with top supplier groups, then incremental expansion.
- Team shape
- Merchandising lead, 2 data engineers, 1 ML engineer, and 4 catalog specialists.
The problem
Adding new suppliers required staff to map inconsistent manufacturer naming, PDF specifications, and spreadsheet data into the distributor internal taxonomy by hand.
What we built
Built a catalog matching engine with attribute normalization, semantic candidate ranking, and low-confidence approval queues for merchandisers.
Operating context
The distributor growth plan depended on faster supplier onboarding, but its product data model was significantly more structured than the incoming supplier material. Manual catalog normalization had become the bottleneck.
Key constraints
- Attribute mapping had to respect internal taxonomy rules and incompatible product classes.
- Supplier PDFs, spreadsheets, and feeds contained conflicting data that needed provenance tracking.
- Merchandisers needed review queues that showed why two items were considered a likely match.
What we built
Attribute normalization layer
Converted units, standardized manufacturer naming, and extracted structured fields before ranking candidates.
Semantic candidate ranking
Used embeddings and metadata weighting to surface likely matches even when naming conventions diverged.
Human review queue
Presented candidates with confidence, conflicting attributes, and source evidence so merchandisers could resolve the hard cases quickly.
Delivery path
Taxonomy alignment
Mapped which product families behaved well with semantic matching and which required heavier rules.
Supplier pilot
Started with a narrow supplier set to measure match confidence, false positives, and review time.
Operational rollout
Integrated approval decisions back into the ranking loop so match quality improved with real use.
Why it mattered
The distributor did not chase full automation. It made the matching system good at easy and medium-difficulty cases and gave specialists a faster review surface for the rest, which is where the throughput gain came from.
Implementation notes
- Normalizing attributes before semantic ranking was mandatory for precision.
- Review UX had to show conflicts clearly or specialists ignored the suggestions.
- Category-by-category rollout prevented one bad product family from poisoning trust in the whole system.