Nested Loops and Iteration
urllib.request.urlopen() opens a web URL and returns a file-like object that you can iterate over line by line.
From Web Text to Counts
Suppose a text file is stored on a remote web server and you want to discover which words appear most often. Counting manually is tedious and error-prone. A program can retrieve the text, process it line by line, examine each word, and maintain a running total in a dictionary.
This task combines two operations. First, urllib.request.urlopen() opens a web URL and returns a file-like object. Second, nested iteration processes the retrieved data: the outer loop visits each line, while the inner loop visits each word in that line. The final dictionary uses each unique word as a key and its total occurrence count as the value.
The Outer and Inner Loops
The outer loop controls the larger units of input: it receives one line at a time from the file-like response. Each line arrives as bytes, so the line must be decoded into a string. The inner loop then processes the words produced by splitting that decoded string. Only after the inner loop has handled all words in the current line does the outer loop move to the next line.
- Open the web URL with urllib.request.urlopen().
- Receive the file-like response and obtain one line through the outer iteration.
- Decode the bytes in that line into readable text.
- Split the decoded text into words.
- Use the inner iteration to update the count for each word.
- Continue until every line and every word has been processed.
Updating the Running Tally
A word-count dictionary maps each unique word to the total number of times that word has appeared across the entire processed text.
For every word, the program needs one rule that handles both cases: a word that has not appeared before and a word that already has a count. counts.get(word, 0) returns the current count when the key exists. When the key does not exist, it returns 0. Adding 1 therefore initializes a new word to 1 and increments an existing word by 1.
counts[word] = counts.get(word, 0) + 1
A Manual Trace
Use the first two source lines, But soft what and light through yonder, to trace the process. The first line is decoded and split into three words. The inner loop encounters But, soft, and what in sequence. The outer loop then advances to the second line and the inner loop processes light, through, and yonder.
Counting Two Lines
Build the dictionary while processing the words But soft what and light through yonder.
Start: The dictionary is empty because no word has been encountered yet.
First line: But, soft, and what are each new words, so each receives a count of 1.
Second line: Light, through, and yonder are also new words, so each receives a count of 1.
Finish: Every word in the two lines has been processed. No word repeats in this small example, so every stored count is 1.
The dictionary contains But: 1, soft: 1, what: 1, light: 1, through: 1, and yonder: 1.
| Word encountered | Reason for update | Count after update |
|---|---|---|
| But | New word; default count is 0 | But: 1 |
| soft | New word; default count is 0 | soft: 1 |
| what | New word; default count is 0 | what: 1 |
| light | New word; default count is 0 | light: 1 |
| through | New word; default count is 0 | through: 1 |
| yonder | New word; default count is 0 | yonder: 1 |
A step-by-step trace of the two source lines
Debugging State Changes
When the final frequencies look wrong, inspect the program while the loops are running. Print the word currently being processed and print the dictionary after each update. This reveals whether the bytes were decoded into the expected text, whether splitting produced the expected words, whether both loops are iterating, and whether counts are accumulating instead of being replaced.
Mistakes in Dictionary Updates
Processing the bytes directly as though they were already readable text.
Each line from the web source arrives as bytes, and it must be decoded to a string before it can be split into words.
Fix:
Decode each line first, then split the resulting string into words.Assigning 1 every time a word is encountered.
A repeated word loses its earlier total because the previous value is overwritten.
Fix:
Use counts[word] = counts.get(word, 0) + 1 so a missing word starts at 1 and an existing word increases by 1.Iterating through lines but not iterating through the words within each line.
The dictionary must be updated once for each word, not merely once for each line.
Fix:
Decode and split each line, then use an inner loop to process every resulting word.Inspecting only the final dictionary when counts are unexpected.
The final result does not show whether the problem occurred during decoding, splitting, iteration, or updating.
Fix:
Add print statements inside the loops to display the current word and dictionary state after each update.
| Situation | Required dictionary behavior | Result |
|---|---|---|
| Word is missing | Use a default of 0, then add 1 | The word receives count 1 |
| Word already exists | Retrieve its current count, then add 1 | The previous total is preserved and increased |
| Word is assigned 1 every time | Replace the previous value | Repeated occurrences are not accumulated |
Practice the Trace
Trace the word-count dictionary for the sequence But, soft, But. Write the dictionary after each word is processed, then explain why the final count for But is greater than the count for soft.
Hints
- Begin with an empty dictionary.
- For a new word, counts.get(word, 0) returns 0 before 1 is added.
- When But appears the second time, retrieve its existing count instead of starting over.
What do you think happens?
After processing the sequence But, soft, But with counts[word] = counts.get(word, 0) + 1, what will the dictionary contain?
Reveal answer
Answer: But: 2, soft: 1
But is encountered twice, so its first count of 1 is retrieved and increased to 2. Soft is encountered once, so its count remains 1.
Key Takeaways
- urllib.request.urlopen() returns a file-like object that can be iterated over line by line.
- Web lines arrive as bytes, so decode each line before splitting it into words.
- Nested loops use an outer iteration for lines and an inner iteration for words.
- counts.get(word, 0) + 1 handles both new words and repeated words.
- Printing the current word and dictionary during processing helps reveal decoding, splitting, iteration, and update errors.
Key Takeaways
- A remote response from urllib.request.urlopen() can be processed like a file, one line at a time.
- The outer loop processes lines, and the inner loop processes words within each decoded line.
- The dictionary stores each unique word as a key and its total occurrence count as the value.
- The update expression counts both new and repeated words without discarding earlier totals.
- Intermediate state tracing is an effective way to find errors in decoding, splitting, iteration, and dictionary updates.