Understanding Email Archive Structure
Mapping tables unify changing email addresses and multiple domain names into single canonical identifiers, preventing entity fragmentation in email archive analysis.
Why Identity Unification Matters
An email archive may contain several email addresses used by the same individual at different times. It may also contain several domain names that belong to one organizational structure. If analysis treats every address or domain name as unrelated, one real entity can appear as several separate entities. Mapping tables solve this fragmentation by connecting variations to one canonical identifier.
The central rule is simple: choose one unified identifier, then map every known variation to it before analysis begins.
Following One User Across Time
The Mapping table is designed to track individual identity changes. Its purpose is to map multiple email addresses to one unified address, so an analysis can follow the same user even when the address appearing in messages changes over time.
Combining Two Addresses for One User
An archive contains two email addresses used by one individual at different times. How should the addresses be represented for unified analysis?
Identify the variations: Treat the addresses as identity variations that may belong to the same individual rather than automatically counting them as two unrelated users.
Choose the unified identifier: Select one address as the unified address used to represent the individual.
Create the mapping: Map both email addresses to the selected unified address in the Mapping table.
Analyze after mapping: Use the completed mapping before analysis so messages associated with either address can be considered under the unified identity.
The two changing email addresses are represented by one canonical user identifier instead of fragmenting the individual into separate entities.
Canonical Choices in Mapping Tables
A mapping table is not merely a list of aliases. It establishes a consistent relationship between variations and a chosen canonical identifier. For email identities, the Mapping table connects multiple addresses to one unified address. For organizational domains, the DNSMapping table connects multiple DNS names to one DNS identifier.
| Table | Raw variations | Canonical result | Purpose |
|---|---|---|---|
| Mapping | Multiple email addresses | One unified address | Track an individual through identity changes |
| DNSMapping | Multiple DNS names | One DNS identifier | Consolidate organizational domains |
Preparing the Archive Before Analysis
Mapping entries should be created before analysis begins. This ordering matters because the analysis should operate on unified identities rather than first producing fragmented results and attempting to repair them afterward.
- Identify the email addresses or DNS names that represent variations in the archive.
- Choose one unified identifier for the entity being analyzed.
- Map every relevant variation to that identifier.
- Complete the mappings before starting archive analysis.
- Use the canonical identifiers when interpreting the analysis results.
Mistakes That Preserve Fragmentation
Treating every changing email address as a separate user
The Mapping table exists to connect multiple email addresses belonging to one individual to one unified address.
Fix:
Create a mapping that assigns the relevant addresses to one chosen canonical identifier before analysis.Using several canonical identifiers for variations of the same entity
The purpose of mapping is to prevent entity fragmentation through consistent unification.
Fix:
Choose one unified identifier and map all relevant variations to it.Leaving mapping until after analysis
The source guidance specifies that mappings should be created before analysis begins.
Fix:
Prepare the Mapping and DNSMapping entries before running the analysis.Applying the email Mapping table to DNS names
The source distinguishes the Mapping table for email addresses from DNSMapping for DNS names.
Fix:
Use Mapping for individual email identities and DNSMapping for organizational DNS consolidation.
Practice the Mapping Decision
An email archive contains several addresses associated with one individual across different periods and several DNS names associated with different campuses of one university system. Decide which mapping table applies to each identity type and describe what the canonical result should be.
Hints
- Separate individual email identities from organizational DNS identities.
- The Mapping table produces one unified address.
- The DNSMapping table produces one DNS identifier.
- Mappings should exist before analysis begins.
Practice Solution
Choose the correct table and canonical result for changing email addresses and multiple organizational DNS names.
Individual addresses: Use the Mapping table because it tracks individual identity changes by mapping multiple email addresses to one unified address.
Organizational domains: Use the DNSMapping table because it consolidates multiple DNS names into one DNS identifier.
Timing: Create both kinds of mappings before analysis begins so the archive is analyzed using unified identities.
Email-address variations resolve to one unified address, while DNS-name variations resolve to one DNS identifier.
Key Takeaways
- The Mapping table unifies multiple email addresses belonging to one individual across identity changes.
- The DNSMapping table consolidates multiple organizational DNS names into one DNS identifier.
- A canonical identifier must be chosen before variations are mapped.
- All relevant variations should be mapped consistently to that identifier.
- Mapping entries should be created before archive analysis to prevent entity fragmentation.
Key Takeaways
- Mapping tables prevent one real entity from appearing as several fragmented entities.
- The Mapping table unifies changing email addresses for an individual.
- The DNSMapping table unifies multiple DNS names into one canonical DNS identifier.
- Choose one canonical identifier and map every relevant variation to it.
- Create mappings before beginning archive analysis.