Maximum and Minimum Loops
A histogram is a frequency count stored in a dictionary, where keys are unique values (like email addresses) and values are counts of how many times each key appears.
From Mail Log to Answer
Suppose a mail log contains thousands of messages. Questions such as which sender emailed most often or which domain sent the most mail cannot be answered reliably by looking at one message at a time. The raw messages must first be transformed into a structured summary. A dictionary histogram provides that summary by storing each unique key with the number of times it appears.
The overall process has three stages: count messages by full email address, scan the counts to find the highest frequency, and then repeat the counting at domain level when the question is about organizations rather than individual senders.
Building the Histogram
A histogram is a frequency count stored in a dictionary. Each key is a unique value, such as a full email address, and each value is the number of messages associated with that key. As each message is processed, its sender becomes the dictionary key and its count increases.
The expression handles both situations in one line. If key is not yet present, get returns 0, so the first occurrence produces a count of 1. If key already has a count, get returns that existing count and the expression increases it by 1.
Counting Two Senders
Process three messages: alex@example.org, alex@example.org, and sam@example.net.
First message: The key alex@example.org is new, so its count starts at 0 and becomes 1.
Second message: The same key already has a count of 1, so it becomes 2.
Third message: The key sam@example.net is new, so its count becomes 1.
The histogram is alex@example.org: 2 and sam@example.net: 1.
Tracking the Highest Count
After the histogram has been built, a maximum loop examines the dictionary entries. It must track two related pieces of information: the highest count seen so far and the key belonging to that count. Tracking only the number would tell you the frequency but not which sender produced it.
counts = { "alex@example.org": 2, "sam@example.net": 5, "lee@example.org": 3 } largest = -1 winner = None for sender in counts: if counts[sender] > largest: largest = counts[sender] winner = sender print(winner, largest)
In the trace, alex@example.org becomes the first winner because its count is larger than the initial value. sam@example.net then replaces it because 5 is larger than 2. The final sender does not replace sam@example.net because 3 is smaller than 5. The winner therefore remains paired with the largest count.
Changing the Grouping Level
A full email address identifies an individual sender, while the domain portion identifies the organization or mail domain associated with that address. Counting full addresses answers which individual sender appears most often. Counting domains answers which organization or domain appears most often. The raw messages can therefore produce different winners depending on the key chosen for the histogram.
The split method divides the email at the @ character. Taking the second part selects the domain. That extracted domain can then be used as the histogram key instead of the complete email address. The source also describes locating @ with find and slicing from the position after it; split is often cleaner.
One Message Set, Two Questions
Analyze messages from alex@example.org, lee@example.org, and sam@example.net, where the first address appears twice, the second appears three times, and the third appears four times.
Full-address histogram: The keys remain individual email addresses, so sam@example.net has the largest individual count with 4 messages.
Domain extraction: The first two addresses produce example.org, while the third produces example.net.
Domain histogram: The two example.org sender counts combine to 5, while example.net remains at 4.
The most frequent individual sender is sam@example.net, but the most frequent domain is example.org. Changing the key changes the question and can change the result.
Mistakes in Frequency Analysis
Using the full email address when the question is about organizations
Different senders from the same domain remain separated, so their frequencies are not combined.
Fix:
Extract the domain with email.split('@')[1] and use the extracted value as the histogram key.Tracking only the largest count
The frequency alone does not identify which key produced it.
Fix:
Track both the current largest value and the key that owns that value.Failing to use a default count for a new key
A new dictionary key does not yet have an accumulated frequency.
Fix:
Use counts[key] = counts.get(key, 0) + 1 so new and existing keys are handled together.Treating a domain result as an individual-sender result
A domain groups messages from potentially multiple full email addresses.
Fix:
State clearly whether the histogram key represents a full email address or a domain.
Decide the question before choosing the histogram key. Use the full email address for individual-sender frequency, and use the extracted domain for organization-level frequency. The same maximum-loop idea can then be applied to whichever histogram represents the question.
Practice the Scan
A histogram contains these counts: pat@example.org: 6, kim@example.net: 2, jo@example.org: 4, and ray@example.net: 3. Identify the most frequent full email address. Then describe what changes if the keys are converted to domains before counting.
Hints
- For the full-address question, compare the values while retaining the key that owns the largest value.
- For the domain question, group the two example.org addresses together and the two example.net addresses together.
- After regrouping, scan the new histogram rather than reusing the full-address winner.
What do you think happens?
Which domain becomes most frequent after grouping the practice data by domain?
Reveal answer
Answer: example.org
The example.org counts combine to 10, while the example.net counts combine to 5. Grouping by domain changes the keys and combines messages from senders with the same domain.
Key Takeaways
- A histogram stores unique keys and their frequencies in a dictionary.
- The pattern counts[key] = counts.get(key, 0) + 1 handles both first occurrences and repeated keys.
- A maximum loop must retain both the highest count and the key associated with it.
- Splitting an email at @ and taking the second part extracts its domain.
- Changing from full email addresses to domains changes the grouping level, the analysis question, and potentially the winner.
Key Takeaways
- Build a dictionary histogram before searching for maximum or minimum frequencies.
- Track a winning key together with its selected frequency.
- Use full email addresses for individual-sender analysis.
- Extract domains when the analysis should combine senders by organization.
- The granularity of the key determines what the result means.