String Methods: split, find, startswith
A histogram is a frequency count stored in a dictionary, where keys are unique values (like email addresses) and values are counts of how many times each key appears.
From Mail Log to Evidence
A mail log may contain thousands of messages. Looking at the raw lines makes it difficult to answer questions such as which sender emailed most often or which organization sent the most mail. A histogram turns those repeated values into a dictionary: each unique value becomes a key, and its frequency becomes the associated value.
The central analytical choice is the key you count. A full email address groups messages by individual sender. A domain groups messages by organization. The same raw messages can therefore produce different results when the analysis becomes more or less granular.
Building the Email Histogram
To count messages from each full email address, process one address at a time and use the address as the dictionary key. The pattern counts[key] = counts.get(key, 0) + 1 handles both cases: a key that has not appeared before starts at zero, while an existing key receives one more count. Repeating this operation makes all messages from the same sender accumulate under one dictionary key.
Counting Repeated Senders
Count these sender addresses: ana@example.com, lee@example.com, ana@example.com, sam@school.org.
First address: ana@example.com is not yet a key, so its count becomes 1.
Second address: lee@example.com is new, so its count becomes 1.
Third address: ana@example.com already has a count of 1, so it becomes 2.
Fourth address: sam@school.org is new, so its count becomes 1.
The histogram is ana@example.com: 2, lee@example.com: 1, and sam@school.org: 1.
Separating Username and Domain
An email address contains a username part before the @ character and a domain part after it. The find method locates the position of @. Once that position is known, slicing from the position after @ selects the domain. The split method provides another approach: splitting the address at @ produces parts, and the second part is the domain.
Selecting the Maximum Sender
After building the histogram, finding the most prolific sender requires a maximum loop. Keep two pieces of state: the key for the current best sender and the value for that sender's current highest count. For each dictionary entry, compare its count with the current maximum. When the new count is higher, replace both tracked values.
largest_sender = None largest_count = 0 for sender in counts: if counts[sender] > largest_count: largest_sender = sender largest_count = counts[sender]
largest_sender: ana@example.com
largest_count: 2Changing the Counting Granularity
Counting full addresses answers an individual-sender question. To ask which organizations send the most mail, extract each address's domain first and use that domain as the histogram key. Messages from different usernames then accumulate together when their domains match.
| Histogram key | Question answered | Grouping effect |
|---|---|---|
| Full email address | Which individual sender sent the most messages? | Different usernames remain separate. |
| Domain | Which organization or domain sent the most messages? | Addresses sharing a domain are combined. |
Checking Message Prefixes
The startswith method can identify lines or messages whose text begins with a required prefix. This is useful when processing a log and selecting only lines that begin with a particular marker before extracting or counting their contents.
Mistakes in Histogram Analysis
Using the full email address when the question asks for organization-level results.
The two senders have different full addresses, but their shared domain should be combined for a domain histogram.
Fix:
Extract the domain with email.split("@")[1] and use that domain as the dictionary key.Updating a histogram without handling a key's first appearance.
A new email address does not yet have a dictionary value.
Fix:
Use counts[email] = counts.get(email, 0) + 1.Tracking only the largest count during the maximum loop.
The count tells you the frequency, but not which dictionary entry produced it.
Fix:
Track both the current maximum key and its value.Treating find's result as the domain itself.
find locates the position of @. It does not extract the text after that position.
Fix:
Slice from the position after @, or use email.split("@")[1].Using startswith to find a prefix anywhere in a line.
The method checks whether the string begins with the specified prefix.
Fix:
Use startswith only when the required text must occur at the beginning.
Practice the Transformation
Suppose the messages come from ana@example.com, lee@example.com, ana@example.com, pat@school.org, and lee@example.com. Determine the full-address histogram, identify the sender with the highest count, and then determine the domain histogram.
Hints
- Use each full email address as the first histogram key.
- For the domain histogram, split each address at @ and use the second part.
- For the maximum, compare each dictionary count while tracking both the sender and the count.
- A strong solution produces a full-address histogram with ana@example.com counted twice, lee@example.com counted twice, and pat@school.org counted once. The maximum loop identifies a tie between ana@example.com and lee@example.com at two messages; the exact selected sender depends on how the loop handles equal counts. The domain histogram combines the first two addresses under example.com for a total of four messages, while school.org has one.
Key Takeaways
- A histogram stores frequencies in a dictionary: unique values are keys and occurrence counts are values.
- Use counts[key] = counts.get(key, 0) + 1 to update a histogram for new and existing keys.
- Use a maximum loop that tracks both the highest count and the key associated with it.
- find locates the @ separator, while split("@")[1] directly extracts the domain.
- Changing the key from a full email address to a domain changes the analysis from individual senders to organizations.