Working with Files in Python
Dictionaries are ideal for counting and categorizing data because they map keys to values instantly.
From Log Lines to Counts
A mail log contains many lines of structured text. Each relevant line starts with From, followed by an email address, a day name such as Sat or Fri, and other information. Instead of examining every line separately, you can summarize the file with a dictionary. The dictionary's keys identify categories such as days, senders, or domains, while its values record how many messages belong to each category.
The central transformation is simple: read one item, decide which key represents it, and increase that key's count. Repeated items use the same key, so their values accumulate. This makes dictionaries useful for counting and categorizing data because they map keys to values directly.
The Counting Pattern
The core counting pattern has three actions: check whether a key exists, initialize its value to 0 if it does not, and increment its value. The initialization gives a new category a starting point; the increment records another occurrence.
Counting Day Names
Suppose the extracted day names arrive in this order: Sat, Fri, Sat.
Read Sat: Sat is not yet a key, so create Sat with a count of 0 and then increase its count to 1.
Read Fri: Fri is not yet a key, so create Fri with a count of 0 and then increase its count to 1.
Read Sat again: Sat already exists, so do not create a new category. Increase the existing Sat count from 1 to 2.
The resulting counts are Sat: 2 and Fri: 1.
Extracting Fields from Each Line
A structured log line contains several fields in a predictable order. Use split() to separate the line into fields, then select the correct index for the information you want. For the mail log, the day name is one field after the email address, while the full email address is another field on the same From line.
Before counting, identify the question you are answering. To count messages by day, select the day field. To count messages by sender, select the full email address. The same input line can support different summaries because different fields become the keys.
Tracking the Largest Count
After building a histogram, you can ask which key has the largest value. Use a maximum loop: initialize max_key and max_value, then inspect each dictionary entry. When an entry has a larger value than the current maximum, update both variables. At the end of the loop, max_key identifies the key associated with the largest count.
Finding the Most Frequent Sender
Consider a sender histogram with alex@example.com: 2, bea@example.com: 5, and cy@example.com: 3.
Initialize: Start with a max_key and max_value representing the current best candidate before comparing the dictionary entries.
Compare alex@example.com: Its count becomes the current largest count encountered so far.
Compare bea@example.com: The count 5 is larger than the current maximum, so update both max_key and max_value.
Compare cy@example.com: The count 3 is not larger than 5, so keep bea@example.com as the current best key.
bea@example.com is the key with the highest count in this example.
Grouping by Domain
Counting individual email addresses distinguishes every sender. Sometimes the useful category is broader: the domain from which the message came. The domain is the part after the @ symbol. Extract that part from each full email address, use it as the dictionary key, and apply the same initialize-and-increment pattern.
Aggregating Senders into Domains
Transform these email addresses into domain counts: first@umich.edu, second@iupui.edu, third@umich.edu.
Extract the first domain: The part after @ in first@umich.edu is umich.edu. Create that key and count one message.
Extract the second domain: The part after @ in second@iupui.edu is iupui.edu. Create that key and count one message.
Extract the third domain: The part after @ in third@umich.edu is umich.edu again. Reuse the existing key and increase its count.
The domain counts are umich.edu: 2 and iupui.edu: 1.
| Counting choice | Dictionary key | Meaning of the value |
|---|---|---|
| By day | Day name | Messages received on that day |
| By sender | Full email address | Messages from that sender |
| By domain | Part after @ | Messages from that domain |
Mistakes to Avoid
Using the wrong field from a structured line
The dictionary will group records by email address instead of by day.
Fix:
Split the line and select the field that answers the counting question.Creating a new category every time an item repeats
Repeated items must contribute to the existing key's value.
Fix:
Check whether the key exists, initialize it when necessary, and increment the existing value.Updating only the maximum value
The largest number would no longer be connected to the key that produced it.
Fix:
Update max_key and max_value together whenever a larger count is found.Counting full addresses when the goal is domain totals
The result counts individual senders rather than the shared domain.
Fix:
Extract the part after @ and use that domain as the key.
A mail log produces these extracted email addresses: one@umich.edu, two@iupui.edu, three@umich.edu, four@umich.edu. Decide what the dictionary should contain when the goal is to count messages by domain. Then identify which domain has the larger count.
Hints
- Use the part after @ as the key.
- The repeated umich.edu domain must update one existing key.
- Compare the final dictionary values to find the largest count.
A Reusable Mental Model
- Read each structured record and identify the field that answers your question.
- Use that field as a dictionary key and apply the initialize-then-increment counting pattern.
- Reuse an existing key when the same item appears again so its value grows.
- For the most frequent item, scan the dictionary while updating max_key and max_value together.
- To aggregate by email domain, extract the part after @ before counting.
Key Takeaways
- Dictionaries summarize repeated data by mapping category keys to occurrence counts.
- The counting pattern checks for a key, initializes a missing key to 0, and increments its value.
- Structured mail lines can provide different keys, including day names and full email addresses.
- A maximum loop tracks both the largest value and the key associated with it.
- Extracting the part after @ changes sender-level data into domain-level aggregates.