Concepts / Querying the mapping.sqlite Database

Querying the mapping.sqlite Database

Mapping tables unify changing email addresses and multiple domain names into single canonical identifiers, preventing entity fragmentation in email archive analysis.

  • Programming

Why Identity Fragments

Large email archives can contain different email addresses used by the same person over time. They can also contain several DNS names that belong to different parts of one organization. If each variation is treated as a separate entity, analysis can split one identity into multiple pieces. The mapping.sqlite database addresses this problem by using mapping tables to connect variations to a single canonical identifier.

Following One User Through Change

The Mapping table is designed for individual identity changes. Instead of analyzing each historical email address independently, the table maps multiple addresses to one unified address. A query or analysis that uses this mapping can therefore follow the same person across address changes without treating every address as a new person.

maps tomaps toold-address@examplehistorical addressunified addresscanonical identifiernew-address@examplelater address
How do several historical email addresses map to one canonical user identifier as the user's identity changes?

Combining a user's historical addresses

A hypothetical archive contains two email addresses that represent one person's identity at different times. How should the Mapping table support analysis?

Choose the unified identifier: Select one address as the single identifier that will represent the individual in the analysis.

Map each variation: Create mapping entries that connect both historical addresses to the selected unified address.

Analyze after mapping: Run the archive analysis with the mappings already available so activity associated with either variation can be considered under the unified identity.

The two historical addresses are treated as variations of one mapped identity rather than as unrelated entities.

Connecting Address Variations

When querying or using the Mapping table, think in terms of direction: each known email-address variation points to the one unified address selected for analysis. The table does not exist merely to list addresses. Its purpose is to express that several values belong to one identity, allowing archive analysis to follow that identity through change.

lookuplookupresolve toEmail variation 1historical valueMappingidentity connectionsUnified addresscanonical identifierEmail variation 2historical value
How does a query use Mapping table entries to connect different email addresses to the same user?

A mapping entry is valuable when it makes identity continuity explicit: a changing address is connected to the single identifier chosen for the analysis.

Consolidating Organizational Domains

The DNSMapping table handles a different level of identity. It consolidates multiple DNS names into a single DNS identifier. This is useful when an organization has multiple domains, such as different campuses within one university system, but the analysis needs to treat them as one organizational identity.

maps throughmaps throughconsolidates toCampus DNS 1organizational domainDNSMappingdomain connectionsUnified DNSidentifiercanonical organizationCampus DNS 2organizational domain
How are several DNS names associated with a single canonical DNS identifier in the DNSMapping table?
Mapping tableIdentity levelPurpose
MappingIndividualMaps multiple email addresses to one unified address.
DNSMappingOrganization or domainConsolidates multiple DNS names into one DNS identifier.

Preventing Entity Fragmentation

Without mapping, an analysis may encounter several email addresses or DNS names and count them as separate entities. With mapping, those variations can be connected to their chosen canonical identifiers. The result is not a new identity; it is a more unified representation of an identity that was already represented by multiple values in the archive.

MappingMappingDNSMappingEmail identity Abefore mappingUnified userafter MappingEmail identity Bbefore mappingUnified organizationafter DNSMappingDNS identity Abefore mapping
What separate entities would appear without mapping tables, and how do the mappings merge them into unified identities?

Creating Mappings Before Analysis

Mapping entries should be created before analysis begins. First choose the single unified identifier for the entity. Next identify the email-address or DNS-name variations that should resolve to it. Then create the corresponding Mapping or DNSMapping entries and use the same choices consistently throughout the analysis.

evaluateyesnothenNew email or DNSvaluevariation foundIdentity evidencesupports unificationChoose canonicalidentifierone unified valueCreate mapping entrybefore analysisKeep value separateno supported unification
What evidence or identity change indicates that a new email or DNS value should be mapped to an existing canonical identifier?
  1. Identify the entity variation in the archive.
  2. Decide whether the variation belongs to an existing individual or organization.
  3. Choose one unified identifier for that entity.
  4. Create the appropriate Mapping or DNSMapping entry before analysis.
  5. Apply the mapping consistently whenever the variation appears.

Common Mapping Mistakes

  • Analyzing before creating mappings

    The source guidance requires mappings to be created before analysis begins, so the initial analysis can treat variations as fragmented entities.

    Fix: Prepare the Mapping and DNSMapping entries before running the analysis.

  • Using several canonical identifiers for one entity

    Effective mapping requires choosing a single unified identifier.

    Fix: Select one unified identifier and map all supported variations to it.

  • Confusing individual and organizational mappings

    The Mapping table addresses email-address identity changes, while DNSMapping consolidates organizational DNS names.

    Fix: Use Mapping for individual email-address variations and DNSMapping for multiple DNS names belonging to one organizational identity.

  • Treating every variation as a confirmed match

    A mapping represents an intentional unification, not merely the existence of similar-looking values.

    Fix: Create a mapping entry only when the variation should be represented by the chosen canonical identifier.

Practice the Decision

MEDIUM

A large email archive contains several historical email addresses associated with one individual and several DNS names associated with different campuses of one university system. Describe which mapping table applies to each case, what must be selected before creating entries, and when the entries should exist relative to the analysis.

Hints
  • Separate the individual identity problem from the organizational DNS problem.
  • For each case, identify the one unified identifier.
  • Remember the required timing for creating mappings.

Practice solution

Apply the mapping approach to the hypothetical archive described above.

Individual addresses: Use the Mapping table to connect the multiple email addresses to one selected unified address.

Organizational DNS names: Use the DNSMapping table to consolidate the campus DNS names into one selected DNS identifier.

Timing and consistency: Create both types of mapping entries before analysis begins and apply the chosen identifiers consistently.

The archive can represent the individual and the organization as unified entities instead of fragmented collections of address or DNS-name variations.

Key Takeaways

  1. The Mapping table connects multiple email addresses belonging to one individual identity across time.
  2. The DNSMapping table consolidates multiple organizational DNS names into one DNS identifier.
  3. Mapping tables prevent entity fragmentation by directing variations to selected canonical identifiers.
  4. Choose one unified identifier for each entity and create mappings before analysis begins.
  5. Apply mappings consistently so the archive analysis treats identity variations as intended.

Key Takeaways

  • The Mapping table unifies changing email addresses for one individual.
  • The DNSMapping table unifies multiple DNS names for one organizational identity.
  • Canonical identifiers prevent people and organizations from being split into separate entities during archive analysis.
  • Mappings should be created before analysis and applied consistently.
  • A mapping entry represents a deliberate identity decision, so variations should be unified only when supported by the analysis context.