Questions · Deduplication

How do I find the duplicates in a database, even when they are written differently?

First every address is put in its correct form, then the names are compared allowing for initials, typing errors and surname and first name in the wrong order. The result is not a yes-or-no duplicate: it is a level of certainty, with the reason written beside it.

Why an exact search does not find them

Two records of the same person are hardly ever written the same way. Whoever entered them used an abbreviation or an initial, swapped surname and first name, wrote “de Vries” once with a capital and once without, or wrote the address with the handy short form of the moment. A search that compares strings recognizes none of this: for it “Kalverstr. 92” and “Kalverstraat 92” are two different addresses, even though a postman would read them as the same place.

This is why the comparison has to be done in two stages, not one. First every address is brought into its correct form with the same engine RadarAddress uses for normalization: street in full, postcode in the country's format and confirmed, city written properly. Only when two addresses are in the same form does it make sense to ask whether they coincide — comparing first and fixing afterwards produces only false negatives, because the comparison keeps seeing two spellings and not one address.

The same goes for the name: someone who searches for “de Vries” expecting to find the record already in the database, and the record is there but filed as “Vries, J. de” or “Jan de Vries”, simply does not find it. The most common consequence is not a visible error: it is a new record opened in good faith, sitting next to the one that already existed.

What the comparison recognizes

  • Initials and abbreviated first names. “J. de Vries” and “Jan de Vries” are compared as the same person when the rest of the record coincides.
  • Surname and first name swapped. “De Vries Jan” and “Jan de Vries” almost always come from an import with the columns exchanged.
  • A typo in the surname. A small typing error does not prevent the match, if the rest of the record coincides.
  • Prefixes written differently. “de Vries”, “De Vries” and “DE VRIES” are the same surname, with or without the capital.
  • The title in front of the name. “Dr. M. Jansen” and “M. Jansen” are compared as the same person: the title does not count in the verdict.
  • The abbreviations of the street, because we fix the addresses before comparing them.
  • The same address in two countries' spellings: records that declare different countries are never merged, while accented and unaccented letters, or ü and ue, are read as the same.

The verdict: certain, probable, ambiguous

Not all duplicates are alike. When two records coincide on everything after the fixing, the verdict is certain. When an initial, an abbreviation or a small typo comes into play, it stays probable: almost always right, but worth a look. When the link is weak — for example similar names only — the group is ambiguous, and the decision to merge the records stays with you. Every group carries its reason: you always know why two records were put together.

The width of the search is yours too, not only the reading of the result: you can choose how broad the comparison should be, from strict (only name and address that coincide) to the widest (weak links too, which you then look at yourself). The default setting, the one most jobs run with, already recognizes initials, abbreviations and small typos.

The main record: a duplicate is not thrown away, it is merged

Discarding one of the two records is the most convenient choice and the most expensive one: if the email was only on one and the phone only on the other, keeping only one loses the other contact detail. RadarAddress builds a main record taking the best field from each — the address from the more complete record, the name in its proper form, email and phone from wherever they are — and for every field it stays written which record it comes from, so you can check where each piece of information came from.

Where you do it: from the site, from a file, from the API

On the site you see the verdict on an example pair: the quickest way to understand how it works before uploading a whole database. For a whole database you upload a file: batch jobs go up to a hundred thousand records, and at the end you find the column with the record to keep and the group every row belongs to. If you work from code, the endpoint /api/v1/dedupe compares up to five hundred records per call, with the same certain/probable/ambiguous verdict and the same merged main record. The price list is on the pricing page: before launching a batch job you always see what it involves, and the operations are taken only when you confirm.

Two records, the same person

Before
Record A: Jan de Vries, Kalverstr. 92, 1012PH Amsterdam
Record B: De Vries J., Kalverstraat 92, 1012 PH AMSTERDAM
After
Same person — level certain
Main record: De Vries Jan, Kalverstraat 92, 1012 PH AMSTERDAM

The abbreviated first name (J. and Jan) and the address written in two different ways are recognized only after both addresses have been fixed in their correct form and verified in the national register. The names are invented for demonstration purposes; any reference to real persons is purely coincidental.

The questions that follow

Do I have to prepare the database before uploading it?

No. You upload the file as it is: fixing the addresses is the first step of the comparison, not a job to do by hand beforehand.

Are the duplicates deleted automatically?

No. Every group of duplicates stays at your disposal with the verdict and the reason: you decide whether and how to merge it, RadarAddress deletes nothing on its own.

What happens if two different people have the same name in the same city?

The comparison flags it as ambiguous instead of declaring it certain: a weak link stays under your decision, it is not forced.

How many records can I compare together?

From a file uploaded, up to a hundred thousand records in a single job. From the API, up to five hundred per call.

Does the record I get lose any data compared with the two I started from?

No: the main record takes the best field from each, email and phone included, and it stays written where every piece of information comes from.

Does it work even if the surname has a typing error?

Yes, within a reasonable margin: a small typo does not prevent the match if the rest of the record coincides.

Do I have to register to compare my records?

Yes: the public example on the site compares a pair chosen by us, with no data entered, so you can see how it works. To work on your own records you need a free account.

See the deduplication

Two example records, compared by the engine: the verdict, the reason and the merged main record.

See the deduplication

Read also: Why Excel does not find the duplicates · Verifying the addresses in an Excel file · Deduplication pricing