Data Cleaning and Normalization in Large Archives
Mapping tables unify changing email addresses and multiple domain names into single canonical identifiers, preventing entity fragmentation in email archive analysis.
The Fragmentation Problem
Large email archives can contain several addresses used by the same person at different times. They can also contain multiple DNS names that belong to one organization, such as different campuses of a university system. If these variations are analyzed as unrelated identities, the archive becomes fragmented. Data cleaning and normalization address this problem by connecting variations to canonical identifiers before analysis begins.
One Person, Several Addresses
The Mapping table tracks individual identity changes. Its purpose is to map multiple email addresses to one unified address or canonical user identifier. The observed addresses remain part of the archive's history, but analysis can use the chosen unified identifier when it needs to follow the individual across those changes.
Following an identity change
An archive contains an earlier email address and a later email address that belong to the same individual. How should the addresses be represented for analysis?
Collect the variations: Treat each observed email address as an identity variation that appears in the archive.
Choose one canonical identifier: Select a single unified address or user identifier to represent the individual during analysis.
Create the Mapping entries: Map each observed address to the chosen unified identifier.
Analyze after mapping: Use the Mapping table so records associated with either variation can be followed as belonging to the unified identity.
The individual is represented by one canonical identity rather than by separate fragments for each changing email address.
From Record to Canonical Entity
Applying a mapping table is a two-level process. First, the archive record contains an address as it was observed. Next, the analysis consults the Mapping table. The result is the unified identifier selected for that address. This lets the analyst track an entity through an identity change without discarding the original variation.
Choose the unified identifier before analysis begins, then map all relevant variations consistently to it. Mapping only some addresses leaves the same identity split across normalized and unnormalized records.
Unifying Organizational DNS Names
Email identity changes are handled through the Mapping table. Organizational domain variations are handled through the DNSMapping table. DNSMapping consolidates multiple DNS names into a single DNS identifier. This is useful when different campuses or other organizational parts should be represented as one underlying organization during archive analysis.
| Mapping table | Variation being unified | Canonical result |
|---|---|---|
| Mapping | Multiple email addresses belonging to one individual | One unified address or user identifier |
| DNSMapping | Multiple DNS names belonging to one organization | One DNS identifier |
When to Add a Mapping
Create a mapping entry when a newly observed email address is another variation of an individual already represented by a canonical identifier, or when a DNS name is another organizational variation represented by a canonical DNS identifier. The purpose is data unification: related variations should resolve to the same chosen entity during analysis.
- Identify the newly observed email address or DNS name.
- Decide whether it belongs with an entity that already has a canonical identifier.
- Choose the existing unified identifier when the variation belongs to that entity.
- Create the Mapping or DNSMapping entry before analysis begins.
- Apply the mapping consistently to the relevant archive records.
An archive contains a newly observed email address that appears to be another address used by an individual already represented in the Mapping table. Describe the normalization steps you would take before analyzing the archive.
Hints
- Start with the relationship between the new address and the existing canonical identity.
- Use the Mapping table rather than creating a separate canonical identity.
- Make the entry before analysis and apply it consistently.
Common Normalization Mistakes
Treating every changing email address as a separate person
The Mapping table is specifically intended to track individual identity changes by mapping multiple email addresses to one unified address.
Fix:
Map the relevant address variations to one chosen canonical user identifier before analysis.Treating related organizational DNS names as unrelated organizations
DNSMapping is intended to consolidate organizational domains, including different campuses of a university system, into one DNS identifier.
Fix:
Create DNSMapping entries that connect the relevant DNS names to one canonical DNS identifier.Beginning analysis before creating mappings
Effective mapping requires mappings to be created before analysis begins.
Fix:
Choose the canonical identifiers and create the required Mapping and DNSMapping entries in advance.Mapping only some variations
All variations need to be mapped consistently to the selected unified identifier.
Fix:
Review the relevant variations and map each one consistently.
Normalization Checklist
- Use the Mapping table to connect multiple email addresses belonging to one individual across time.
- Use the DNSMapping table to consolidate multiple organizational DNS names into one DNS identifier.
- Select one canonical identifier for each unified entity.
- Create mappings before analysis begins and apply them consistently.
- Normalization prevents related people or organizations from becoming fragmented archive entities.
Key Takeaways
- A Mapping table unifies changing email addresses that belong to one individual.
- A DNSMapping table consolidates multiple organizational DNS names into one DNS identifier.
- Canonical identifiers must be chosen before analysis, with all relevant variations mapped consistently.
- Applying mappings prevents entity fragmentation in large email archive analysis.