Two records, one person
A duplicate is not an extra record to delete: it is one person's history split in two. Putting it back together has four concrete effects.
Reply once, and with full knowledge
Whoever replies looks at one record: if the order is registered on the other, that person receives nothing — and notices. With a single record the confirmation always goes out, and the communication arrives once: not twice at the same address, perhaps with the name spelled two different ways.
Every extra envelope is money thrown away
A duplicate is a print, an envelope and a stamp paid for a person your communication has already reached. In a large campaign it is the easiest cost to cut: not one recipient fewer, only the shipments that are needed.
Recognise who is worth more
Someone who bought three times under one record and twice under the other shows up in reports as two lukewarm contacts instead of one loyal one. Merging the records makes the history whole again, and changes who is worth calling, who deserves different treatment, who to offer more to.
The numbers add up
How many customers do I really have? What is a contact worth on average? With a database full of duplicates every answer is inflated, and the decisions that follow are flawed from the start. Before analysing, you have to count right.
They are not careless mistakes
Duplicates come from three ordinary situations, repeated in every database that is meant to grow.
The order comes in and there is no time
Whoever enters the data has a pile of forms to get through (or of payment slips, in a non-profit) and doesn't stop to check whether that person is already on file. Opening a new record takes thirty seconds; really searching for it, among the variants of the name and the address, takes much longer.
The record is there, but the search doesn't find it
It is the most frequent case and the hardest to avoid: the record exists, but is written just slightly differently from how it is searched — “Via Giovanni Pascoli 4” against “Pascoli”, “V.le” against “Viale”. The search returns nothing, and the record is created again in good faith.
A whole list comes in
Sign-ups from the website, a list collected at an event, a file from a supplier: they come in as a block without being compared with the existing database. It happens less often than the other two cases, but when it does, duplicates come in by the hundreds.
All three are solved the same way: a check of the database at regular intervals, and one before every important campaign.
First we fix, then we compare
Comparing two records word by word doesn't get you far: Daniela and Dany are the same person, V. Roma 12 and Via Roma, 12 are the same address, and no literal comparison notices. That is why we work in two steps: first we fix the two addresses — street name in full, correct postcode, the right city — and only then do we compare them, together with the name, the email and the phone. Two spellings of the same street become the same street before the comparison begins.
The result is not a yes or a no. When the verdict is not clear-cut we say so — certain, probable or ambiguous, with the reason next to it: two different people can have the same name in the same town, and merging two records remains your decision, to be taken knowing what it rests on.
Normalisation, duplicate detection, deduplication, enrichment
Four different operations, usually all filed under the word “cleaning”. Telling them apart matters, because each one is the premise of the next.
Normalisation
Reducing every address to a single, correct spelling, completing what is missing: the postcode, the province, the street name in full.
“Via S. Caterina 20” and “Strada Santa Caterina 20” become the same line.
Duplicate detection
Recognising that two records are the same subject even when written differently: “V. Roma 12” and “Via Roma 12”, Dany and Daniela, surname and first name swapped.
It comes after normalisation, and not by chance: without it, you compare strings instead of addresses.
Deduplication
Deciding what to do with the duplicates found: which record to keep, which fields to take from one and which from the other.
The duplicate is not thrown away: it is merged in, and where each piece of information comes from stays on record.
Enrichment
Adding what is missing or can be deduced from what is there: the salutation from the name, the title, the house number suffix, the tax code.
Data nobody gave you, but that your database already contains implicitly.
Why it matters, beyond cleaning: whoever receives the same letter twice notices, and whoever doesn't receive the thank-you because the donation was registered on one record while you looked at the other notices even more.
You choose how strict the comparison is, we tell you how sure we are
We don't ask you to decide which records to merge: we do that, and for every group we say how certain we are. The only knob is the strictness of the comparison, that is how different two records can be and still be the same subject. Three values: the website and file jobs use the default, the API lets you choose with livello_minimo. Not to be confused with the precision of address normalisation, which is a different thing and has five levels.
Strict
Name and address coincide after normalisation. No group you wouldn't have made by hand.
The default
Variants of the same name (Dany and Daniela), abbreviations and small typos come in. It is the setting most jobs run with.
Loose
Weak links too: same address and similar names. It finds more, and you have to look at more.
The verdict returned on every group — certain, probable, ambiguous — is our classification, not your choice: it says how solid the group is, and helps you decide what you can accept blindly and what is worth reviewing.
The duplicate is not thrown away: it is merged
This is the difference that counts. Whoever deduplicates usually keeps one record and discards the other — and with it discards the email that was only there. We build a new record taking the best field from each.
| Field | Record A | Record B | Record that remains |
|---|---|---|---|
| First name | Dany Rossi | Rossi Daniela | Daniela Rossi canonical form, fields put back in place |
| Address | v. leopardi 4, milano | Via G. Leopardi 4, Milano | Via Giacomo Leopardi 4 from the more complete record |
| Postcode and city | — | Milano | 20123 MILANO MI completed during normalisation |
| dany.rossi@gmail.com | — | dany.rossi@gmail.com was only on A | |
| Phone | — | 347 032 8959 | 3470328959 was only on B |
Keeping A you would have lost the phone, keeping B the email. The merged record has both, and for every field it stays on record which one it comes from: if one day the merge doesn't convince you, you know exactly what to touch.
The same duplicates that escape an exact search
- Dany Rossi=Daniela Rossi equivalent name
- Rossi Mario=Mario Rossi surname and first name swapped
- V. Roma 12=Via Roma 12 abbreviation expanded
- Rossi Daniela, via leopardi 4, milano=Daniela Rossi, Via Giacomo Leopardi 4, 20123 MILANO MI address fixed before the comparison
We also recognise a typo in the surname; the title before the name, which does not count in the comparison (Dott.ssa Maria Bianchi is the same person as Maria Bianchi); and the locality brought back to its municipality, because Sampierdarena is a district of Genoa.
Related questions
- How do you verify and correct a customer database?
- How do I find the duplicates in a database, even when they are written differently?
- Excel does not find the duplicates: why, and what to do?
All the questions, with the answer, in Questions and answers.