Concepts / Data Cleaning and Normalization in Large Archives

Data Cleaning and Normalization in Large Archives

Mapping tables unify changing email addresses and multiple domain names into single canonical identifiers, preventing entity fragmentation in email archive analysis.

  • Programming

The Fragmentation Problem

Large email archives can contain several addresses used by the same person at different times. They can also contain multiple DNS names that belong to one organization, such as different campuses of a university system. If these variations are analyzed as unrelated identities, the archive becomes fragmented. Data cleaning and normalization address this problem by connecting variations to canonical identifiers before analysis begins.

One Person, Several Addresses

The Mapping table tracks individual identity changes. Its purpose is to map multiple email addresses to one unified address or canonical user identifier. The observed addresses remain part of the archive's history, but analysis can use the chosen unified identifier when it needs to follow the individual across those changes.

maps tomaps toEarlier addressobserved email addressUnified usercanonical identifierLater addressobserved email address
How do several email addresses used by the same person across time map to a single user identifier?

Following an identity change

An archive contains an earlier email address and a later email address that belong to the same individual. How should the addresses be represented for analysis?

Collect the variations: Treat each observed email address as an identity variation that appears in the archive.

Choose one canonical identifier: Select a single unified address or user identifier to represent the individual during analysis.

Create the Mapping entries: Map each observed address to the chosen unified identifier.

Analyze after mapping: Use the Mapping table so records associated with either variation can be followed as belonging to the unified identity.

The individual is represented by one canonical identity rather than by separate fragments for each changing email address.

From Record to Canonical Entity

Applying a mapping table is a two-level process. First, the archive record contains an address as it was observed. Next, the analysis consults the Mapping table. The result is the unified identifier selected for that address. This lets the analyst track an entity through an identity change without discarding the original variation.

look up addressreturn unified identifierArchive recordoriginal email addressMapping tableaddress-to-identifier entryCanonical entityunified user identifier
How does an email record move from its original address through a mapping table to a canonical entity identifier?

Choose the unified identifier before analysis begins, then map all relevant variations consistently to it. Mapping only some addresses leaves the same identity split across normalized and unnormalized records.

Unifying Organizational DNS Names

Email identity changes are handled through the Mapping table. Organizational domain variations are handled through the DNSMapping table. DNSMapping consolidates multiple DNS names into a single DNS identifier. This is useful when different campuses or other organizational parts should be represented as one underlying organization during archive analysis.

maps tomaps toCampus DNS nameobserved DNS nameUnified DNScanonical DNS identifierOrganization DNSnameobserved DNS name
What happens when several DNS names refer to one underlying organization, and how are they represented by one DNS identifier?
Mapping tableVariation being unifiedCanonical result
MappingMultiple email addresses belonging to one individualOne unified address or user identifier
DNSMappingMultiple DNS names belonging to one organizationOne DNS identifier

When to Add a Mapping

Create a mapping entry when a newly observed email address is another variation of an individual already represented by a canonical identifier, or when a DNS name is another organizational variation represented by a canonical DNS identifier. The purpose is data unification: related variations should resolve to the same chosen entity during analysis.

  1. Identify the newly observed email address or DNS name.
  2. Decide whether it belongs with an entity that already has a canonical identifier.
  3. Choose the existing unified identifier when the variation belongs to that entity.
  4. Create the Mapping or DNSMapping entry before analysis begins.
  5. Apply the mapping consistently to the relevant archive records.
MEDIUM

An archive contains a newly observed email address that appears to be another address used by an individual already represented in the Mapping table. Describe the normalization steps you would take before analyzing the archive.

Hints
  • Start with the relationship between the new address and the existing canonical identity.
  • Use the Mapping table rather than creating a separate canonical identity.
  • Make the entry before analysis and apply it consistently.

Common Normalization Mistakes

  • Treating every changing email address as a separate person

    The Mapping table is specifically intended to track individual identity changes by mapping multiple email addresses to one unified address.

    Fix: Map the relevant address variations to one chosen canonical user identifier before analysis.

  • Treating related organizational DNS names as unrelated organizations

    DNSMapping is intended to consolidate organizational domains, including different campuses of a university system, into one DNS identifier.

    Fix: Create DNSMapping entries that connect the relevant DNS names to one canonical DNS identifier.

  • Beginning analysis before creating mappings

    Effective mapping requires mappings to be created before analysis begins.

    Fix: Choose the canonical identifiers and create the required Mapping and DNSMapping entries in advance.

  • Mapping only some variations

    All variations need to be mapped consistently to the selected unified identifier.

    Fix: Review the relevant variations and map each one consistently.

Normalization Checklist

  1. Use the Mapping table to connect multiple email addresses belonging to one individual across time.
  2. Use the DNSMapping table to consolidate multiple organizational DNS names into one DNS identifier.
  3. Select one canonical identifier for each unified entity.
  4. Create mappings before analysis begins and apply them consistently.
  5. Normalization prevents related people or organizations from becoming fragmented archive entities.

Key Takeaways

  • A Mapping table unifies changing email addresses that belong to one individual.
  • A DNSMapping table consolidates multiple organizational DNS names into one DNS identifier.
  • Canonical identifiers must be chosen before analysis, with all relevant variations mapped consistently.
  • Applying mappings prevents entity fragmentation in large email archive analysis.