Querying the mapping.sqlite Database
Mapping tables unify changing email addresses and multiple domain names into single canonical identifiers, preventing entity fragmentation in email archive analysis.
Why Identity Fragments
Large email archives can contain different email addresses used by the same person over time. They can also contain several DNS names that belong to different parts of one organization. If each variation is treated as a separate entity, analysis can split one identity into multiple pieces. The mapping.sqlite database addresses this problem by using mapping tables to connect variations to a single canonical identifier.
Following One User Through Change
The Mapping table is designed for individual identity changes. Instead of analyzing each historical email address independently, the table maps multiple addresses to one unified address. A query or analysis that uses this mapping can therefore follow the same person across address changes without treating every address as a new person.
Combining a user's historical addresses
A hypothetical archive contains two email addresses that represent one person's identity at different times. How should the Mapping table support analysis?
Choose the unified identifier: Select one address as the single identifier that will represent the individual in the analysis.
Map each variation: Create mapping entries that connect both historical addresses to the selected unified address.
Analyze after mapping: Run the archive analysis with the mappings already available so activity associated with either variation can be considered under the unified identity.
The two historical addresses are treated as variations of one mapped identity rather than as unrelated entities.
Connecting Address Variations
When querying or using the Mapping table, think in terms of direction: each known email-address variation points to the one unified address selected for analysis. The table does not exist merely to list addresses. Its purpose is to express that several values belong to one identity, allowing archive analysis to follow that identity through change.
A mapping entry is valuable when it makes identity continuity explicit: a changing address is connected to the single identifier chosen for the analysis.
Consolidating Organizational Domains
The DNSMapping table handles a different level of identity. It consolidates multiple DNS names into a single DNS identifier. This is useful when an organization has multiple domains, such as different campuses within one university system, but the analysis needs to treat them as one organizational identity.
| Mapping table | Identity level | Purpose |
|---|---|---|
| Mapping | Individual | Maps multiple email addresses to one unified address. |
| DNSMapping | Organization or domain | Consolidates multiple DNS names into one DNS identifier. |
Preventing Entity Fragmentation
Without mapping, an analysis may encounter several email addresses or DNS names and count them as separate entities. With mapping, those variations can be connected to their chosen canonical identifiers. The result is not a new identity; it is a more unified representation of an identity that was already represented by multiple values in the archive.
Creating Mappings Before Analysis
Mapping entries should be created before analysis begins. First choose the single unified identifier for the entity. Next identify the email-address or DNS-name variations that should resolve to it. Then create the corresponding Mapping or DNSMapping entries and use the same choices consistently throughout the analysis.
- Identify the entity variation in the archive.
- Decide whether the variation belongs to an existing individual or organization.
- Choose one unified identifier for that entity.
- Create the appropriate Mapping or DNSMapping entry before analysis.
- Apply the mapping consistently whenever the variation appears.
Common Mapping Mistakes
Analyzing before creating mappings
The source guidance requires mappings to be created before analysis begins, so the initial analysis can treat variations as fragmented entities.
Fix:
Prepare the Mapping and DNSMapping entries before running the analysis.Using several canonical identifiers for one entity
Effective mapping requires choosing a single unified identifier.
Fix:
Select one unified identifier and map all supported variations to it.Confusing individual and organizational mappings
The Mapping table addresses email-address identity changes, while DNSMapping consolidates organizational DNS names.
Fix:
Use Mapping for individual email-address variations and DNSMapping for multiple DNS names belonging to one organizational identity.Treating every variation as a confirmed match
A mapping represents an intentional unification, not merely the existence of similar-looking values.
Fix:
Create a mapping entry only when the variation should be represented by the chosen canonical identifier.
Practice the Decision
A large email archive contains several historical email addresses associated with one individual and several DNS names associated with different campuses of one university system. Describe which mapping table applies to each case, what must be selected before creating entries, and when the entries should exist relative to the analysis.
Hints
- Separate the individual identity problem from the organizational DNS problem.
- For each case, identify the one unified identifier.
- Remember the required timing for creating mappings.
Practice solution
Apply the mapping approach to the hypothetical archive described above.
Individual addresses: Use the Mapping table to connect the multiple email addresses to one selected unified address.
Organizational DNS names: Use the DNSMapping table to consolidate the campus DNS names into one selected DNS identifier.
Timing and consistency: Create both types of mapping entries before analysis begins and apply the chosen identifiers consistently.
The archive can represent the individual and the organization as unified entities instead of fragmented collections of address or DNS-name variations.
Key Takeaways
- The Mapping table connects multiple email addresses belonging to one individual identity across time.
- The DNSMapping table consolidates multiple organizational DNS names into one DNS identifier.
- Mapping tables prevent entity fragmentation by directing variations to selected canonical identifiers.
- Choose one unified identifier for each entity and create mappings before analysis begins.
- Apply mappings consistently so the archive analysis treats identity variations as intended.
Key Takeaways
- The Mapping table unifies changing email addresses for one individual.
- The DNSMapping table unifies multiple DNS names for one organizational identity.
- Canonical identifiers prevent people and organizations from being split into separate entities during archive analysis.
- Mappings should be created before analysis and applied consistently.
- A mapping entry represents a deliberate identity decision, so variations should be unified only when supported by the analysis context.