Concepts / Analyzing Communication Patterns in Email Networks

Analyzing Communication Patterns in Email Networks

Raw email data contains inconsistencies (mixed-case addresses, subdomain variations, gateway artifacts) that must be standardized before analysis.

  • Programming

Why Raw Addresses Mislead

An email network is built from identities such as senders, recipients, and domains. Raw email storage records messages as they arrive, including full headers, complete message bodies, and addresses in the exact format used by senders. That fidelity is useful for preservation, but it creates problems for analysis. The same person may appear under mixed-case versions of an address, while related institutional addresses may include different subdomains or gateway-related forms.

For example, John.Doe@EXAMPLE.COM and john.doe@example.com can refer to the same person even though a simple comparison treats them as different strings. Similarly, si.umich.edu may need to be represented at the institution level as umich.edu rather than as a separate subdomain identity.

lowercasealready standardizedtruncate subdomainJohn.Doe@EXAMPLE.COMraw formjohn.doe@example.comstandardized identityjohn.doe@example.comraw formumich.eduinstitution-level domainsi.umich.edusubdomain form
What misleading differences can appear when inconsistent raw addresses are analyzed without cleaning?

Without normalization, analysis can overcount unique senders, split related domains into separate entities, and distort conclusions about collaboration networks. The goal is not merely to make the data look tidy; it is to make communication relationships analytically meaningful.

The gmodel.py Pipeline

gmodel.py implements a pipeline with a clear direction: it reads raw email data from content.sqlite, applies cleaning and normalization rules, and writes a new compressed, normalized database to index.sqlite. The output is rebuilt on each run. This allows the cleaning process to be refined by adjusting parameters or editing mapping tables in the source database, then running the transformation again.

raw recordsmessagesaddressesnormalized recordscontent.sqliteraw email storageRead messagesheaders and bodiesClean addressescase and gateway rulesNormalize domainstruncate by TLDindex.sqlitecompressed normalizedoutput
How does an email record move from raw storage through cleaning and normalization into the final compressed output?

Tracing One Record Through the Pipeline

Trace the conceptual transformation of a message whose address includes mixed case and a subdomain.

Raw storage: The message is retained in content.sqlite with the address as it arrived, along with the full headers and message body.

Address cleaning: The address is converted to lowercase so a mixed-case spelling does not create a separate identity.

Domain normalization: The domain is reduced according to its top-level-domain rule. For a common TLD such as .edu, the normalized domain keeps two levels.

Output storage: The cleaned result is written into the rebuilt index.sqlite output rather than being analyzed as a separate raw spelling.

The record moves from faithful but inconsistent raw storage to a smaller representation designed for consistent communication analysis.

Address Standardization

gmodel.py applies lowercase conversion to email addresses. This causes mixed-case spellings to converge when the underlying address is the same. It also performs gmane address replacement, which helps unify fragmented sender identities across the corpus. The replacement is conditional: it occurs when a matching real address is available, rather than treating every gmane-related form as automatically interchangeable.

lowercase conversionreplace when matchedJohn.Doe@EXAMPLE.COMraw addressjohn.doe@example.comlowercase addressgmane addressgateway-related formmatching real addressreplacement target
What changes in an email address when gmodel.py standardizes case and handles a gateway-related identity?

Do not assume that every gateway-related address should be replaced blindly. The normalization rule for a gmane address depends on a matching real address. If that match is not present, the replacement condition described for gmodel.py has not been satisfied.

What do you think happens?

What should happen to John.Doe@EXAMPLE.COM when the lowercase rule is applied?

  • It remains a separate mixed-case identity
  • It becomes john.doe@example.com
  • Only the domain is removed
  • It is replaced by a gmane address
Reveal answer

Answer: It becomes john.doe@example.com.

Lowercase conversion standardizes the address spelling. It does not remove the local part or automatically invoke gmane replacement.

Domain Truncation Rules

Domain truncation controls the level at which an organization is represented. Common top-level domains, specifically .com, .org, .edu, and .net, are truncated to two levels. Other top-level domains are truncated to three levels. The asymmetric rule preserves more structure for domains where three levels are needed to retain institutional identity.

extract domaininspect TLDtruncate to two levelsperson@si.umich.edufull addressumich.edutwo domain levelssi.umich.edusubdomain included.educommon TLD
When is a domain truncated, and how does a subdomain-containing address map to the normalized domain?
Top-level domain categoryTruncation levelAnalytical purpose
.com, .org, .edu, .netTwo levelsRepresent the broader organization or domain
Other top-level domainsThree levelsPreserve institutional identity

The domain rule is asymmetric: the top-level-domain category determines how much of the domain is retained.

Applying the TLD Decision

A university address appears as si.umich.edu. Determine the normalized domain under the common-TLD rule.

Identify the TLD: The domain ends in .edu, which is one of the common TLDs listed by the normalization rule.

Count retained levels: Common TLDs are truncated to two levels.

Remove the subdomain: The si subdomain is removed while umich.edu is retained.

si.umich.edu maps to umich.edu.

Identity Mapping

Normalization maps several raw representations onto the identity level needed by the analysis. Mixed-case addresses can converge through lowercase conversion. Subdomain forms can converge at an institution-level domain through truncation. Gateway-related sender forms can converge through gmane replacement when a matching real address exists. The result is not that the original data never existed; it is that the network analysis receives standardized identities instead of treating every spelling or routing artifact as a distinct participant.

lowercasetruncate by TLDreplace if matchedMixed-case addressraw spellingStandardized senderlowercase or matchedreplacementSubdomain addressraw organization formStandardizedorganizationtruncated domainGmane addressgateway-related form
How can multiple raw email or domain forms represent the same normalized sender, recipient, or organization?

The cleaning rules encode an analytical choice about granularity. For university communication analysis, the institution may matter more than the specific subdepartment mail server. Domain truncation therefore changes the identity used by the network model to match the level of organization being studied.

Operational Feedback

gmodel.py reports information that helps you inspect a run. It reports the number of unique sender addresses loaded and the number of active mapping rules, then prints progress for every 250 messages. The progress lines include timestamps and sender addresses. In the described run, the program loaded 1588 unique senders and 28 mapping rules. The output also showed normalized lowercase addresses and two-level .edu domains, providing a practical check that the rules were being applied.

Because each run deletes and rebuilds index.sqlite, the pipeline supports an iterative cleaning workflow. If a truncation rule is too aggressive or a sender needs special handling, you can revise the relevant parameters or mapping tables and regenerate the output. The rebuilt output should then be checked again through the reported counts, progress lines, and normalized addresses.

Common Analysis Mistakes

  • Treating every raw spelling as a separate sender

    Lowercase conversion is intended to unify mixed-case spellings that represent the same address.

    Fix: Apply the address normalization rules before counting senders or building communication links.

  • Keeping every subdomain as a separate organization

    For a common TLD such as .edu, gmodel.py truncates to two levels to represent the institution.

    Fix: Inspect the TLD and apply the corresponding two-level or three-level truncation rule.

  • Assuming all domains use the same truncation depth

    The rules are asymmetric. Other TLDs are truncated to three levels to preserve institutional identity.

    Fix: Classify the TLD before deciding how many domain levels to retain.

  • Replacing every gmane-related address without checking for a match

    The replacement rule applies when a matching real address exists.

    Fix: Use the replacement only when the required match is present.

  • Judging the transformation only by compression

    Compression is a side benefit; the main purpose is to improve the validity of communication analysis.

    Fix: Check whether equivalent identities and related domains have been standardized correctly.

Apply the Rules

MEDIUM

For each case, state what gmodel.py should do and explain why: a mixed-case email address; si.umich.edu; a domain ending in .com with an extra subdomain; a domain with a non-common TLD; and a gmane address with no matching real address.

Hints
  • Start with lowercase conversion for the email address.
  • For domains, identify whether the TLD is one of .com, .org, .edu, or .net.
  • Remember that other TLDs retain three levels.
  • For gmane replacement, check whether a matching real address exists before replacing anything.

What do you think happens?

Before checking your reasoning, which two checks are most important when normalizing a domain?

  • Check only the message timestamp and body length
  • Check the TLD category and the required truncation depth
  • Check only whether the local part contains uppercase letters
  • Check whether the archive is compressed before reading the domain
Reveal answer

Answer: Check the TLD category and the required truncation depth.

The truncation rule is asymmetric: common TLDs use two levels, while other TLDs use three levels.

What to Remember

  1. Raw email records preserve exactly what arrived, but mixed case, subdomains, and gateway artifacts can split one communication identity into several raw forms.
  2. gmodel.py reads raw data from content.sqlite, applies cleaning rules, and rebuilds compressed normalized data in index.sqlite.
  3. Email normalization includes lowercase conversion and conditional gmane address replacement when a matching real address exists.
  4. Common TLDs .com, .org, .edu, and .net are truncated to two levels; other TLDs are truncated to three levels.
  5. Normalization improves the validity of sender, organization, and communication-pattern analysis; compression is a secondary benefit.

Key Takeaways

  • Raw email data must be standardized because equivalent senders and organizations can appear under different spellings or domain forms.
  • gmodel.py transforms data from content.sqlite into a rebuilt, compressed, normalized index.sqlite database.
  • Lowercase conversion, conditional gmane replacement, and TLD-dependent domain truncation are distinct normalization rules.
  • Common TLDs are reduced to two domain levels, while other TLDs retain three levels to preserve institutional identity.
  • The main success criterion is more accurate communication analysis, not simply a smaller database.