How We Clean Company Data
At ImportYeti, we process millions of Bills of Lading (BOLs). Company names often appear in multiple variations, with typos, codes, or formatting differences. We use machine learning to unify them into a single clean identity.
From Dirty Names to One Clean Company
Raw shipping records contain messy, inconsistent company names — we treat them as a black box of variations. Our machine learning models learn patterns, group related names, and output one standardized company page.
100 Com 10 Ikea Supply Ag
460 Ikea Wholesale
5100 Com 10 Ikea 5100 Com 15 Ikeasupply Ag
5100 Com 18 Ikea Supply
Ikea Suppl
Ikea Suppl Ag
Ikea Supplu
Name joining, standardization, noise removal, and deduplication — trained on millions of BOL records.
What Our Models Do
Name joining — A single company like IKEA can appear under dozens of different names. We group all variations into one unified company page instead of fragmented data.
Standardization — We remove noise such as internal codes, suffixes (Inc, LLC, Ltd), and formatting inconsistencies. Example: IKEA Supply AG, IKEA Supply Inc., and Ikea Supply all become IKEA Supply.
Personal data removal — Models also detect and remove personal information that may appear in shipping records (names or sensitive identifiers that should not be publicly displayed), keeping the dataset compliant and clean.
Easy to Visualize — Dirty vs. Clean
Without cleaning, the same company is scattered across many pages with different names — hard to grasp at a glance. After machine learning, you get one visual story: one page, one name, one view of global shipping activity.
It is much easier to visually show a single IKEA page than to explain 100 separate pages for IKEA with slightly different names. That clarity makes search, analysis, and lead generation significantly more powerful.
Same company, many names — fragmented, confusing, impossible to scan in one view.
IKEA Supply — one clean name, one accurate view of all related shipments.