Two records,
one person.

An example with addresses in Netherlands: two records written differently, and why they are the same person.

With a free account you work on your own addresses, from any country: from the website, from a file or through the API. Sign up or log in.

Record A

Record B

The examples use addresses of public buildings and city-centre streets; the names, where they appear, are invented for demonstration purposes: any reference to real persons is purely coincidental.

Two records, one person

A duplicate is not an extra record to delete: it is one person's history split in two. Putting it back together has four concrete effects.

Reply once, and with full knowledge

Whoever replies looks at one record: if the order is registered on the other, that person receives nothing — and notices. With a single record the confirmation always goes out, and the communication arrives once: not twice at the same address, perhaps with the name spelled two different ways.

Every extra envelope is money thrown away

A duplicate is a print, an envelope and a stamp paid for a person your communication has already reached. In a large campaign it is the easiest cost to cut: not one recipient fewer, only the shipments that are needed.

Recognise who is worth more

Someone who bought three times under one record and twice under the other shows up in reports as two lukewarm contacts instead of one loyal one. Merging the records makes the history whole again, and changes who is worth calling, who deserves different treatment, who to offer more to.

The numbers add up

How many customers do I really have? What is a contact worth on average? With a database full of duplicates every answer is inflated, and the decisions that follow are flawed from the start. Before analysing, you have to count right.

They are not careless mistakes

Duplicates come from three ordinary situations, repeated in every database that is meant to grow.

1

The order comes in and there is no time

Whoever enters the data has a pile of forms to get through (or of payment slips, in a non-profit) and doesn't stop to check whether that person is already on file. Opening a new record takes thirty seconds; really searching for it, among the variants of the name and the address, takes much longer.

2

The record is there, but the search doesn't find it

It is the most frequent case and the hardest to avoid: the record exists, but is written just slightly differently from how it is searched — “Via Giovanni Pascoli 4” against “Pascoli”, “V.le” against “Viale”. The search returns nothing, and the record is created again in good faith.

3

A whole list comes in

Sign-ups from the website, a list collected at an event, a file from a supplier: they come in as a block without being compared with the existing database. It happens less often than the other two cases, but when it does, duplicates come in by the hundreds.

All three are solved the same way: a check of the database at regular intervals, and one before every important campaign.

First we fix, then we compare

Comparing two records word by word doesn't get you far: Daniela and Dany are the same person, V. Roma 12 and Via Roma, 12 are the same address, and no literal comparison notices. That is why we work in two steps: first we fix the two addresses — street name in full, correct postcode, the right city — and only then do we compare them, together with the name, the email and the phone. Two spellings of the same street become the same street before the comparison begins.

The result is not a yes or a no. When the verdict is not clear-cut we say so — certain, probable or ambiguous, with the reason next to it: two different people can have the same name in the same town, and merging two records remains your decision, to be taken knowing what it rests on.

Normalisation, duplicate detection, deduplication, enrichment

Four different operations, usually all filed under the word “cleaning”. Telling them apart matters, because each one is the premise of the next.

Normalisation

Reducing every address to a single, correct spelling, completing what is missing: the postcode, the province, the street name in full.

“Via S. Caterina 20” and “Strada Santa Caterina 20” become the same line.

Duplicate detection

Recognising that two records are the same subject even when written differently: “V. Roma 12” and “Via Roma 12”, Dany and Daniela, surname and first name swapped.

It comes after normalisation, and not by chance: without it, you compare strings instead of addresses.

Deduplication

Deciding what to do with the duplicates found: which record to keep, which fields to take from one and which from the other.

The duplicate is not thrown away: it is merged in, and where each piece of information comes from stays on record.

Enrichment

Adding what is missing or can be deduced from what is there: the salutation from the name, the title, the house number suffix, the tax code.

Data nobody gave you, but that your database already contains implicitly.

Why it matters, beyond cleaning: whoever receives the same letter twice notices, and whoever doesn't receive the thank-you because the donation was registered on one record while you looked at the other notices even more.

You choose how strict the comparison is, we tell you how sure we are

We don't ask you to decide which records to merge: we do that, and for every group we say how certain we are. The only knob is the strictness of the comparison, that is how different two records can be and still be the same subject. Three values: the website and file jobs use the default, the API lets you choose with livello_minimo. Not to be confused with the precision of address normalisation, which is a different thing and has five levels.

1 · uguale

Strict

Name and address coincide after normalisation. No group you wouldn't have made by hand.

2 · poco diverso

The default

Variants of the same name (Dany and Daniela), abbreviations and small typos come in. It is the setting most jobs run with.

3 · diverso

Loose

Weak links too: same address and similar names. It finds more, and you have to look at more.

The verdict returned on every group — certain, probable, ambiguous — is our classification, not your choice: it says how solid the group is, and helps you decide what you can accept blindly and what is worth reviewing.

The duplicate is not thrown away: it is merged

This is the difference that counts. Whoever deduplicates usually keeps one record and discards the other — and with it discards the email that was only there. We build a new record taking the best field from each.

Example of merging two records: which field remains and where it comes from
FieldRecord ARecord BRecord that remains
First nameDany RossiRossi DanielaDaniela Rossi canonical form, fields put back in place
Addressv. leopardi 4, milanoVia G. Leopardi 4, MilanoVia Giacomo Leopardi 4 from the more complete record
Postcode and city—Milano20123 MILANO MI completed during normalisation
Emaildany.rossi@gmail.com—dany.rossi@gmail.com was only on A
Phone—347 032 89593470328959 was only on B

Keeping A you would have lost the phone, keeping B the email. The merged record has both, and for every field it stays on record which one it comes from: if one day the merge doesn't convince you, you know exactly what to touch.

The same duplicates that escape an exact search

  • Dany Rossi=Daniela Rossi equivalent name
  • Rossi Mario=Mario Rossi surname and first name swapped
  • V. Roma 12=Via Roma 12 abbreviation expanded
  • Rossi Daniela, via leopardi 4, milano=Daniela Rossi, Via Giacomo Leopardi 4, 20123 MILANO MI address fixed before the comparison

We also recognise a typo in the surname; the title before the name, which does not count in the comparison (Dott.ssa Maria Bianchi is the same person as Maria Bianchi); and the locality brought back to its municipality, because Sampierdarena is a district of Genoa.