Dictionary Creation and Access
Dictionaries are ideal for counting and categorizing data because they map keys to values instantly.
From Raw Records to Counts
A dataset often contains repeated items, but the useful result is a summary of how often each item appears. A mail log, for example, can be summarized by counting messages for each day, sender, or email domain. A dictionary is well suited to this task because it maps each key to a value. The key identifies the category being counted, and the value stores the running count.
The central counting pattern is simple: check whether a key exists, initialize it to 0 if it does not, and then increment its value.
The Counting Pattern
Suppose each record produces one category key. For every record, first determine the key. If the dictionary has not seen that key before, give it an initial value of 0. Then increase the value by 1. When the same key appears again, the dictionary does not create a second entry. It updates the existing value instead.
counts = {} for item in ["Sat", "Fri", "Sat"]: if item not in counts: counts[item] = 0 counts[item] = counts[item] + 1 print(counts)
Selecting the Useful Field
A mail log line contains several fields. The line starts with From, followed by an email address, then a day name such as Sat or Fri, followed by other information. The counting task determines which field becomes the dictionary key. To count messages by day, extract the day field. To count messages by sender, extract the full email address instead.
SatFrom Addresses to Domains
A sender histogram uses the full email address as its key. A domain histogram groups senders more broadly. To create it, first extract the email address from the From line, then split the address at the @ symbol and use the part after @ as the dictionary key. Addresses from the same domain therefore contribute to the same count.
lines = [ "From alice@umich.edu Sat details", "From bob@iupui.edu Fri details", "From carol@umich.edu Fri details" ] domain_counts = {} for line in lines: fields = line.split() email = fields[1] domain = email.split("@")[1] if domain not in domain_counts: domain_counts[domain] = 0 domain_counts[domain] = domain_counts[domain] + 1 print(domain_counts)
Tracking the Largest Count
After building a histogram, you can ask which key has the highest count. The maximum-loop pattern keeps two pieces of information: the key currently associated with the largest value and that largest value itself. As the loop examines each dictionary entry, it compares the entry's value with the current maximum. When the new value is larger, both tracked pieces of information are updated.
domain_counts = { "umich.edu": 2, "iupui.edu": 1, "example.edu": 4 } max_key = None max_value = None for key in domain_counts: value = domain_counts[key] if max_value is None or value > max_value: max_key = key max_value = value print(max_key) print(max_value)
Mistakes in Dictionary Counting
Incrementing a key before initializing it
A newly encountered item has no stored count yet.
Fix:
Check whether the key exists, initialize it to 0 when necessary, and then increment it.Using the wrong field from a structured line
The dictionary then groups records by a different category from the one intended.
Fix:
Split the line and select the field that represents the desired category.Counting full addresses when the task is to count domains
The two messages come from the same domain but remain separated by their individual addresses.
Fix:
Extract the portion after the @ symbol and use that domain as the key.Updating only the maximum value
The stored maximum would no longer identify which dictionary key produced it.
Fix:
Update max_key and max_value together whenever a larger count is found.
Practice the Transformation
Given these structured records, describe the dictionary produced when counting by domain: From ana@umich.edu Sat details, From ben@iupui.edu Fri details, From cy@umich.edu Fri details. Then identify which domain has the larger count.
Hints
- Select the email address after splitting each line.
- Split the email address at @ and use the portion after it as the key.
- Apply the initialization-and-increment pattern for each domain.
Domain Count and Maximum
Count the domains in three mail-log records and find the domain with the highest count.
Extract: The email addresses are ana@umich.edu, ben@iupui.edu, and cy@umich.edu.
Transform: The domains are umich.edu, iupui.edu, and umich.edu.
Count: The resulting dictionary is {'umich.edu': 2, 'iupui.edu': 1}.
Compare: A maximum loop compares the values 2 and 1 and retains umich.edu as the key with the larger value.
The domain with the highest count is umich.edu, with 2 messages.
Working Pattern
- Identify the category you need before choosing a dictionary key.
- Use split() to separate structured text into fields.
- Use the correct field index for the record format.
- Initialize unseen keys to 0 before incrementing their counts.
- Use the full email address for sender counts and the portion after @ for domain counts.
- When searching for the largest count, update the maximum key and maximum value together.
Key Takeaways
- A dictionary represents categories as keys and their occurrence counts as values.
- Counting requires checking for a key, initializing unseen keys to 0, and incrementing the value.
- Structured mail-log lines can provide different keys depending on the selected field: day, full email address, or domain.
- A domain count is created by extracting the part of an email address after the @ symbol.
- A maximum loop compares dictionary values and updates both the key and value when a larger count appears.
Key Takeaways
- Dictionaries map counting categories to occurrence totals.
- The standard counting pattern initializes unseen keys and increments every key.
- The selected field determines whether records are grouped by day, sender, or domain.
- A maximum loop remembers the key and value associated with the largest count.
- Email-domain aggregation transforms individual addresses into broader category counts.