Dictionaries and the .get() Method
urllib.request.urlopen() opens a web URL and returns a file-like object that you can iterate over line by line.
From Web Text to Word Counts
Imagine downloading a text file from the internet and needing to find which words appear most often. Counting manually would be tedious and error-prone. A program can instead retrieve the text, read it line by line, split each line into words, and keep a running total in a dictionary.
This task has two connected parts. First, urllib.request.urlopen() opens a web URL and returns a file-like object. Second, a dictionary records how many times each word occurs. The remote location does not change the basic parsing approach: after the URL is opened, the data can be processed line by line in the same general way as file data.
A First Dictionary Trace
A dictionary is a key-value store. In a word-frequency dictionary, each unique word is a key and its total occurrence count is the value. The dictionary starts empty. Each processed word either creates a new key with count 1 or increases the value already stored for that key.
Counting Two Short Lines
Trace the dictionary while processing the source example lines: But soft what and light through yonder.
Start: The counts dictionary is empty: {}.
Process But: But is not present, so its count becomes 1.
Process soft: soft is new, so its count becomes 1.
Process what: what is new, so its count becomes 1.
Process light: light is new, so its count becomes 1.
Process through: through is new, so its count becomes 1.
Process yonder: yonder is new, so its count becomes 1.
The final dictionary contains each of the six words with a value of 1.
Reading and Updating Counts
Each line supplied by the web source arrives as bytes. Bytes must be decoded into a string before the line can be split into words. The processing sequence is therefore: obtain a line, decode it, split it into words, then process each word.
The expression reads the current count with get. If the word is already a key, counts.get(word, 0) returns its current value. If the word is not a key, the method returns the default value 0. Adding 1 then produces the correct next count in either case, and assigning the result back to counts[word] stores it.
| Word status | Value returned by get | Value stored after adding 1 |
|---|---|---|
| Word is already in the dictionary | The current count | The current count increased by 1 |
| Word is not in the dictionary | 0 | 1 |
The two cases handled by dictionary.get(word, 0).
Tracing Repeated Words
The important state change appears when a word repeats. A first occurrence creates a key with value 1. A later occurrence does not create a separate entry; it reads the existing value and replaces it with that value plus 1.
Following One Repeated Word
Trace the count for the word soft while processing the generated sequence: soft, light, soft.
Read soft first time: soft is absent, so get returns 0. Adding 1 stores soft with count 1.
Read light: light is absent, so light receives count 1. The existing soft count remains 1.
Read soft second time: soft is present with count 1. get returns 1, adding 1 produces 2, and soft is stored with count 2.
The final counts assign soft the value 2 and light the value 1.
counts = {} counts[word] = counts.get(word, 0) + 1
Debugging the Loop
When the result looks wrong, inspect the program while it runs instead of looking only at the final dictionary. Print the word being processed and the dictionary after each update. This exposes whether the input is being decoded correctly, whether the loops are iterating as expected, and whether counts are accumulating.
Mistakes in Frequency Updates
Resetting a word to 1 every time it appears
A repeated word loses its previous count, so the dictionary cannot accumulate the total number of occurrences.
Fix:
Read the existing value or use 0, then add 1 with counts[word] = counts.get(word, 0) + 1.Forgetting to store the updated value
The expression calculates a number but does not assign that number back to the dictionary key.
Fix:
Assign the result to counts[word].Processing bytes as though they were already readable text
Each line from the web source arrives as bytes and must be decoded to a string before it can be split into words.
Fix:
Decode each line first, then split the resulting string.Inspecting only the final dictionary
The final state does not show when an unexpected word or incorrect count entered the result.
Fix:
Print the current word and dictionary inside the processing loop to trace state changes.
A word is processed three times. Describe the dictionary value after the first, second, and third occurrences when the update is counts[word] = counts.get(word, 0) + 1. Then explain what would happen if the update were replaced by counts[word] = 1.
Hints
- For the first occurrence, the word is absent and get returns 0.
- For later occurrences, get returns the value already stored for the word.
- Compare accumulating the old value plus 1 with replacing the value by 1.
A Reliable Processing Pattern
- Open the remote URL with urllib.request.urlopen().
- Iterate over the returned file-like object one line at a time.
- Decode each bytes line into a string.
- Split the decoded string into individual words.
- For each word, use counts[word] = counts.get(word, 0) + 1.
- Inspect the word and dictionary during debugging when the final counts are unexpected.
Suppose the next three processed words are soft, soft, and what. Write the dictionary state after each update, starting with an empty dictionary. In a second attempt, describe which stage you would inspect first if an unexpected bytes value appeared as a word.
Hints
- The first soft creates a key with value 1.
- The second soft increases that value.
- what is a separate key.
- An unexpected bytes value points you toward the decoding stage.
Key Takeaways
- urllib.request.urlopen() returns a file-like object that can be iterated line by line.
- Web lines arrive as bytes, so they must be decoded into strings before word processing.
- A word-frequency dictionary maps each unique word to its total occurrence count.
- counts.get(word, 0) handles both a word's first occurrence and its later occurrences.
- Printing the current word and dictionary during the loop reveals decoding, iteration, and update errors.
Key Takeaways
- Remote text can be opened as a file-like object and processed one line at a time.
- Each bytes line must be decoded before it is split into words.
- The dictionary stores words as keys and total frequencies as values.
- The expression counts[word] = counts.get(word, 0) + 1 accumulates counts correctly for new and repeated words.
- Tracing each word and dictionary state is an effective way to find word-counting mistakes.