Analyzing Contributor Networks and Participation Patterns
Email addresses and domain names change over time as people move between organizations, requiring a mapping system to consolidate multiple variants into single canonical addresses
Why Identity Fragments
A contributor can appear under different email addresses after moving between organizations. If an archive treats every address as a separate identity, one person's messages can be split across several apparent contributors. The mapping system addresses this problem by transforming multiple historical variants into a single canonical address during processing.
The goal is not to rewrite the original archive. The goal is to give gmodel.py rules for treating related addresses as equivalent while it processes the archive.
Following a Career Across Addresses
A canonical address is the destination chosen to represent a contributor's identity. In the source example, Steve Githens has three email addresses. The most recent address, swgithen@mtu.edu, is selected as the canonical destination, while the Northwestern and Cambridge addresses are treated as historical variants.
Building Individual Address Rules
The Mapping table in mapping.sqlite stores individual email mappings using arrow notation: source address to destination address. For the Steve Githens example, two entries are required. One entry sends the Northwestern address to swgithen@mtu.edu, and another sends the Cambridge address to swgithen@mtu.edu. No third entry is needed for the canonical address.
Consolidating Three Addresses
Steve Githens appears under three email addresses and the most recent address is swgithen@mtu.edu. How many Mapping entries are needed?
Select the destination: Use swgithen@mtu.edu as the canonical destination.
Map the first variant: Create an entry from the Northwestern address to swgithen@mtu.edu.
Map the second variant: Create an entry from the Cambridge address to swgithen@mtu.edu.
Leave the destination alone: Do not create a self-mapping entry for swgithen@mtu.edu.
Two Mapping entries are needed: one for each non-canonical address.
Applying Rules During Processing
When gmodel.py processes the archive, it reads each message and extracts sender and recipient addresses. For every address, it checks the Mapping table. If a mapping exists, gmodel.py uses the destination address. If no mapping exists, it keeps the address as-is. Because this check happens for every email, one mapping can affect many messages when a contributor participated frequently.
What do you think happens?
Suppose an address has no Mapping entry. What address does gmodel.py use?
Reveal answer
Answer: It uses the address as-is.
gmodel.py substitutes the destination only when a mapping exists. Without a mapping, the encountered address remains unchanged during processing.
Normalizing Domains First
DNSMapping works at the domain level rather than the individual address level. When gmodel.py encounters someone@iupui.edu, it first checks whether iupui.edu has a DNSMapping entry. If iupui.edu maps to indiana.edu, the address becomes someone@indiana.edu before individual email address mappings are checked.
| Mapping system | Level | Purpose | Processing order |
|---|---|---|---|
| DNSMapping | Domain | Consolidates domain names | First |
| Mapping | Individual email address | Consolidates address variants | After DNSMapping |
Avoiding Identity Errors
Creating a mapping entry for the canonical address itself.
The canonical address is already in its final form and does not require a self-mapping entry.
Fix:
Create entries only for the non-canonical addresses that should point to swgithen@mtu.edu.Treating DNSMapping as an individual-address table.
DNSMapping operates on domain names, while the Mapping table handles individual email addresses.
Fix:
Use DNSMapping for domain-level consolidation and Mapping for complete address variants.Assuming mappings rewrite the original archive.
Mappings are applied during gmodel.py processing and do not modify the original email data.
Fix:
Treat mappings as adjustable transformation rules that can be modified or removed.Applying individual address mapping before domain mapping.
Domain mapping happens first, followed by individual email mapping.
Fix:
Follow the processing order: DNSMapping first, then Mapping.
When resolving archive data quality issues, first decide whether the inconsistency is at the domain level or the complete-address level. Then select DNSMapping or Mapping accordingly, choose a clear canonical destination, and remember that the original archive remains available if the rule needs correction.
Practice the Decision
An archive contains an address whose domain should be consolidated with another domain, and it also contains a historical address variant for a contributor. Decide which mapping system should handle each issue and state the processing order.
Hints
- Ask whether the rule concerns a domain or a complete email address.
- DNSMapping operates before individual address mappings.
- Only non-canonical address variants need entries in Mapping.
- A strong answer identifies DNSMapping as the tool for the domain-level issue and Mapping as the tool for the historical complete-address issue. The domain transformation is applied first. The resulting address is then checked against individual address mappings. The rules affect processing while leaving the original archive unchanged.
Key Takeaways
- The Mapping table in mapping.sqlite sends historical email addresses to one canonical destination.
- For three addresses, only the two non-canonical addresses need Mapping entries.
- DNSMapping consolidates domains before individual email address mappings are applied.
- gmodel.py applies these rules while processing each message; it uses the original address when no mapping exists.
- Mappings are non-destructive, so they can be revised without changing the original email archive.