Questions · Addresses

How do you verify and correct a customer database?

Correcting a customer database is a job in three steps, in the right order: first the addresses are put in their correct form, then the duplicate records are recognized, finally the email addresses and phone numbers of every record are checked. Every change is declared, field by field, and the decisions that call for judgement stay with the people who know the customers.

What correcting a database means

A customer database ages on its own: whoever enters the data is in a hurry, whoever imports a list swaps the columns, whoever takes a call writes the name by ear. After a while the same database contains addresses abbreviated in ten different ways, the same person on two or three records, email addresses with a typo in the domain and phone numbers with half a country code. None of these is a serious error on its own. All together they send letters back, mail the same communication twice and inflate the numbers of every campaign.

Correcting does not mean deleting the imperfect records. It means bringing every piece of data back to its correct form when possible, flagging what cannot be rebuilt with certainty, and merging the records that speak of the same person without losing anything they contained. The result is a database on which a campaign, an invoice or a search by name gives the expected result.

The order of the three steps, and why it matters

First: the addresses. Street in full, house number in its place, postcode in the country's format and confirmed for that house number, city in its official form. You start here because everything else rests on the address: two records with “Kalverstr. 92” and “Kalverstraat 92” look different until the address is written the same way.

Second: the duplicates. With the addresses fixed, the comparison between records recognizes the same person even with an abbreviated first name, surname and first name swapped, a typo or a “de” written with a capital. Every group of duplicates carries a level of certainty and the reason for the match, and proposes the main record with the best field taken from each.

Third: the contact details. Email addresses with the domain corrected, phone numbers in the international form with the type recognized, websites that really answer. These are checks on the single record: they are done at the end, when the records are already the right ones, so the same contact is not verified twice.

What a correction must not do

  • Invent. If a street is not found in the municipality or the house number is missing, the outcome says so and leaves the decision to you. A plausible address built at a desk is worse than an address declared uncertain.
  • Delete on its own. Duplicates are flagged with the verdict and the proposed main record. Merging or keeping two records apart is a choice for the people who know the customers.
  • Discard for a typo. An email address at “gmial.com” is not a lost contact: it is a contact with the correction proposed beside it.
  • Work in silence. Every modified field must have next to it the value before, the value after and the reason. Without this, the correction can neither be checked nor explained to whoever pays for it.

How it is done in practice with RadarAddress

You export the database to an Excel or CSV file with the columns you have: surname, first name, street, postcode, city, country, email, phone. No need to prepare it: the upload recognizes the separator, the encoding and the headers. Before the job starts you see a preview of the rows and how many operations it needs; the operations are taken only when you confirm.

The first job is the contact verification, which fixes the address and, in the same row, tidies surname and first name. The second, on the same file, is the deduplication, which returns every row with the group it belongs to and the record to keep. Email addresses, phone numbers and websites are verified with the dedicated jobs, one item at a time. In every result file you find, for every row, the outcome, the correct form and what changed; the end-of-job report sums up the outcomes by category, so you know at once how many rows were fine, how many were corrected and how many need a look by hand.

If you work from code, the same three steps are in the API: one endpoint per service, single or in batches, and the asynchronous job for large databases. To prepare the file well there is the file guide; the price list is on the pricing page.

How the result is measured

A correction is judged on the numbers before and after, not on a feeling. The useful numbers are few: how many rows had an address the engine had to correct, how many are left with an outcome to review, how many groups of duplicates were found and with what level of certainty, how many email addresses and phone numbers are unusable. The report of every job gives them by outcome category, and your account keeps the history of everything you have processed.

The number that really counts comes later, and it is yours: how many letters come back at the next campaign, and how many duplicate communications go out. If you were not measuring it before the correction, this is the moment to start: it is the only way to know what it was worth.

A record before and after the correction

Before
de vries jan, kalverstr 92, 1012 PH amsterdam
jan.devries@gmial.com · 0031 6 12345678
After
De Vries Jan, Kalverstraat 92, 1012 PH AMSTERDAM
jan.devries@gmail.com (typo corrected) · +31 6 12345678 mobile

Every row of the result carries the outcome and the list of changes with the reason. If the database also holds “J. de Vries, Kalverstraat 92, Amsterdam”, the deduplication matches it to this record with its level of certainty. The name is invented for demonstration purposes; any reference to real persons is purely coincidental.

The questions that follow

Do I have to clean the file before uploading it?

No. The file is uploaded as it is: recognizing the separator, the encoding and the headers is part of the job. The columns can have the names you use.

Can I do only part of the correction, for example only the duplicates?

Yes. Each step is a job of its own. The deduplication fixes the addresses before the comparison anyway, because without that step duplicates written differently are not recognized.

Are the records changed in my own software?

No. RadarAddress works on the file you upload and gives you back a result file with the correct form next to the original one. What to carry over into your software is up to you.

How big a database can I upload?

Address verification works up to a million rows in a single job, deduplication up to a hundred thousand records. Above that, you split the file.

What about the contacts that cannot be fixed?

They stay in the result file with the outcome that explains why: street not found in the municipality, house number missing, email address with a domain that does not exist. They are the list of what needs a look by hand or a question to the customer.

How often should it be repeated?

It depends on how much enters the database. Those who load external lists or have a lot of manual entries repeat it before every campaign; those with little movement do it once and then verify new contacts from the form or the API at the moment they are entered.

Prepare the file and upload it

The guide explains which columns each step needs and how to read the result file.

Prepare the file and upload it

Read also: Verifying the addresses in an Excel file · Finding duplicates in a database · Cleaning an email list