Concepts / Dictionary Basics: Creating and Accessing

Dictionary Basics: Creating and Accessing

A histogram is a frequency count stored in a dictionary, where keys are unique values (like email addresses) and values are counts of how many times each key appears.

  • Programming

From Mail Log to Useful Summary

Imagine a mail log containing thousands of messages. The raw messages are difficult to compare directly, but a dictionary can transform them into a frequency summary. In a histogram, each unique value becomes a dictionary key and the value stored for that key is the number of times it appears.

The same counting pattern can answer several related questions. First count messages by full email address, then find the sender with the largest count, and finally count by domain to analyze organizations rather than individual senders.

Building the Email Histogram

A histogram uses a dictionary as a frequency counter. Each email address is a key. When an address appears for the first time, its count starts at zero and is immediately increased to one. When the address appears again, the existing count is retrieved and increased by one. The pattern counts[key] = counts.get(key, 0) + 1 handles both cases in one line.

creates countincrements countcreates countana@example.orgfirst appearanceana@example.orgcount: 2ana@example.orgsecond appearancebo@sample.netcount: 1bo@sample.netone appearance
How does each email address become a unique dictionary key, and how does its count change when the address appears again?

emails = [ "ana@example.org", "bo@sample.net", "ana@example.org", "ana@example.org", "bo@sample.net" ] counts = {} for email in emails: counts[email] = counts.get(email, 0) + 1 print(counts)

lookupreturns valueana@example.orgdictionary keycountsemail histogram3stored frequency
Given an email address, how does the program locate its corresponding message count?

Finding the Largest Sender Count

After building the histogram, the most prolific sender can be found with a maximum loop. The loop examines each dictionary entry and keeps two pieces of information: the key belonging to the best entry so far and that entry's count. When a larger count is found, both tracked values are replaced.

examine entryexamine next entryretain larger countstartbest sender: none; bestcount: 0ana@example.orgcount: 3bo@sample.netcount: 2ana@example.orgwinning sender; count: 3
As the loop examines each dictionary entry, how do the current maximum count and winning sender change?
python
Output (expected)
ana@example.org
3

The two tracking variables must be updated together. The count identifies whether the current entry is better, while the sender records which key produced that count. At the end of the loop, the tracked sender and tracked count describe the highest-frequency dictionary entry.

From Sender Names to Domains

A full email address answers an individual-sender question. A domain answers an organization-level question. To extract the domain, split the email address at the @ symbol and use the second part: email.split('@')[1]. The find method can also locate the @ symbol, after which string slicing can take the text from the position after that symbol.

applypart 1part 2ana@example.orgoriginal emailsplit('@')separate at @anafirst partexample.orgsecond part
How does splitting an email address at the @ symbol produce a username and a domain?

domain_counts = {} for email in emails: domain = email.split('@')[1] domain_counts[domain] = domain_counts.get(domain, 0) + 1 print(domain_counts)

same domainsame domainsame domainana@example.orgcount: 2example.orgcount: 3lee@example.orgcount: 1sample.netcount: 1bo@sample.netcount: 1
How do multiple full email addresses become one domain key, and how does that change the frequency counts?

Mistakes in Histogram Analysis

  • Using a new count of one every time instead of increasing the existing count.

    A repeated key is overwritten with one, so the dictionary does not retain the frequency of earlier appearances.

    Fix: Use counts[email] = counts.get(email, 0) + 1.

  • Tracking only the largest count during the maximum loop.

    The final number identifies the frequency but not which dictionary key produced it.

    Fix: Track both the sender key and its count whenever a larger count is found.

  • Counting full email addresses when the question is about organizations.

    Different people at the same domain remain separate keys.

    Fix: Extract domain = email.split('@')[1] and use domain as the histogram key.

  • Selecting the first part after splitting an email address.

    The first part is the text before the @ symbol, not the domain.

    Fix: Use the second part, email.split('@')[1], for the domain.

Before writing the loop, decide what one dictionary key should represent. If the key is a full email address, the result measures individual senders. If the key is a domain, the result measures combined activity from organizations. The histogram pattern stays the same; only the value used as the key changes.

Practice the Three Transformations

MEDIUM

Given the email sequence ana@example.org, bo@sample.net, lee@example.org, ana@example.org, and bo@sample.net, write a program that creates a full-address histogram, finds the sender with the highest count, and creates a domain histogram.

Hints
  • Start with an empty dictionary for full email counts.
  • Use counts[email] = counts.get(email, 0) + 1.
  • For the maximum loop, keep both the winning key and its count.
  • Extract each domain with email.split('@')[1], then apply the same histogram pattern.

What do you think happens?

Before running the domain histogram, predict the count for example.org when ana@example.org appears twice and lee@example.org appears once.

  • 1
  • 2
  • 3
  • The domain is not counted
Reveal answer

Answer: 3

Changing the key from the full email address to the domain combines both senders under example.org, so their appearances contribute to one domain count.

Key Takeaways

  1. A dictionary histogram maps each unique key to the number of times that key appears.
  2. The pattern counts[key] = counts.get(key, 0) + 1 handles both new and existing keys.
  3. A maximum loop must track both the highest count and the key associated with it.
  4. email.split('@')[1] extracts the domain for organization-level counting.
  5. The choice of key controls the granularity and meaning of the analysis.

Key Takeaways

  • Use dictionaries as frequency counters for repeated values.
  • Count full email addresses to compare individual senders.
  • Track both key and value in a maximum loop to identify the most frequent sender.
  • Extract domains with email.split('@')[1] when the analysis should group senders by organization.
  • Changing the dictionary key changes the question answered by the histogram.