Concepts / Dictionaries and the .get() Method

Dictionaries and the .get() Method

urllib.request.urlopen() opens a web URL and returns a file-like object that you can iterate over line by line.

  • Programming

From Web Text to Word Counts

Imagine downloading a text file from the internet and needing to find which words appear most often. Counting manually would be tedious and error-prone. A program can instead retrieve the text, read it line by line, split each line into words, and keep a running total in a dictionary.

This task has two connected parts. First, urllib.request.urlopen() opens a web URL and returns a file-like object. Second, a dictionary records how many times each word occurs. The remote location does not change the basic parsing approach: after the URL is opened, the data can be processed line by line in the same general way as file data.

openiteratesplitupdateWeb URLFile-like objectlines of bytesText linedecoded stringWordone item at a timeWord-count dictionaryword to total count
How does raw web text move from lines to individual words and then into dictionary entries?

A First Dictionary Trace

A dictionary is a key-value store. In a word-frequency dictionary, each unique word is a key and its total occurrence count is the value. The dictionary starts empty. Each processed word either creates a new key with count 1 or increases the value already stored for that key.

Counting Two Short Lines

Trace the dictionary while processing the source example lines: But soft what and light through yonder.

Start: The counts dictionary is empty: {}.

Process But: But is not present, so its count becomes 1.

Process soft: soft is new, so its count becomes 1.

Process what: what is new, so its count becomes 1.

Process light: light is new, so its count becomes 1.

Process through: through is new, so its count becomes 1.

Process yonder: yonder is new, so its count becomes 1.

The final dictionary contains each of the six words with a value of 1.

process Butprocess softprocess whatprocess lightprocess throughprocess yonder{}before any wordBut: 1after Butsoft: 1But remains 1what: 1three keyslight: 1four keysthrough: 1five keysyonder: 1six keys
What happens to the word-count dictionary after each word is processed?

Reading and Updating Counts

Each line supplied by the web source arrives as bytes. Bytes must be decoded into a string before the line can be split into words. The processing sequence is therefore: obtain a line, decode it, split it into words, then process each word.

python

The expression reads the current count with get. If the word is already a key, counts.get(word, 0) returns its current value. If the word is not a key, the method returns the default value 0. Adding 1 then produces the correct next count in either case, and assigning the result back to counts[word] stores it.

look uppresentabsentadd 1wordlookup keyexisting countif key is present0if key is absentcount plus 1new stored value
How does dictionary.get(word, 0) choose between an existing count and the default value 0?
Word statusValue returned by getValue stored after adding 1
Word is already in the dictionaryThe current countThe current count increased by 1
Word is not in the dictionary01

The two cases handled by dictionary.get(word, 0).

Tracing Repeated Words

The important state change appears when a word repeats. A first occurrence creates a key with value 1. A later occurrence does not create a separate entry; it reads the existing value and replaces it with that value plus 1.

Following One Repeated Word

Trace the count for the word soft while processing the generated sequence: soft, light, soft.

Read soft first time: soft is absent, so get returns 0. Adding 1 stores soft with count 1.

Read light: light is absent, so light receives count 1. The existing soft count remains 1.

Read soft second time: soft is present with count 1. get returns 1, adding 1 produces 2, and soft is stored with count 2.

The final counts assign soft the value 2 and light the value 1.

process soft againunchangedsoft: 1first occurrencelight: 1different wordsoft: 2second occurrencelight: 1same count
How does the dictionary change when a word appears for a second time?

counts = {} counts[word] = counts.get(word, 0) + 1

Debugging the Loop

When the result looks wrong, inspect the program while it runs instead of looking only at the final dictionary. Print the word being processed and the dictionary after each update. This exposes whether the input is being decoded correctly, whether the loops are iterating as expected, and whether counts are accumulating.

python
add 1resetword: old countexisting dictionary valueword: old countexisting dictionary valueword: old plus 1accumulated valueword: 1previous value discarded
What is the difference between accumulating a word's count and losing its previous state?

Mistakes in Frequency Updates

  • Resetting a word to 1 every time it appears

    A repeated word loses its previous count, so the dictionary cannot accumulate the total number of occurrences.

    Fix: Read the existing value or use 0, then add 1 with counts[word] = counts.get(word, 0) + 1.

  • Forgetting to store the updated value

    The expression calculates a number but does not assign that number back to the dictionary key.

    Fix: Assign the result to counts[word].

  • Processing bytes as though they were already readable text

    Each line from the web source arrives as bytes and must be decoded to a string before it can be split into words.

    Fix: Decode each line first, then split the resulting string.

  • Inspecting only the final dictionary

    The final state does not show when an unexpected word or incorrect count entered the result.

    Fix: Print the current word and dictionary inside the processing loop to trace state changes.

EASY

A word is processed three times. Describe the dictionary value after the first, second, and third occurrences when the update is counts[word] = counts.get(word, 0) + 1. Then explain what would happen if the update were replaced by counts[word] = 1.

Hints
  • For the first occurrence, the word is absent and get returns 0.
  • For later occurrences, get returns the value already stored for the word.
  • Compare accumulating the old value plus 1 with replacing the value by 1.

A Reliable Processing Pattern

  1. Open the remote URL with urllib.request.urlopen().
  2. Iterate over the returned file-like object one line at a time.
  3. Decode each bytes line into a string.
  4. Split the decoded string into individual words.
  5. For each word, use counts[word] = counts.get(word, 0) + 1.
  6. Inspect the word and dictionary during debugging when the final counts are unexpected.
MEDIUM

Suppose the next three processed words are soft, soft, and what. Write the dictionary state after each update, starting with an empty dictionary. In a second attempt, describe which stage you would inspect first if an unexpected bytes value appeared as a word.

Hints
  • The first soft creates a key with value 1.
  • The second soft increases that value.
  • what is a separate key.
  • An unexpected bytes value points you toward the decoding stage.

Key Takeaways

  1. urllib.request.urlopen() returns a file-like object that can be iterated line by line.
  2. Web lines arrive as bytes, so they must be decoded into strings before word processing.
  3. A word-frequency dictionary maps each unique word to its total occurrence count.
  4. counts.get(word, 0) handles both a word's first occurrence and its later occurrences.
  5. Printing the current word and dictionary during the loop reveals decoding, iteration, and update errors.

Key Takeaways

  • Remote text can be opened as a file-like object and processed one line at a time.
  • Each bytes line must be decoded before it is split into words.
  • The dictionary stores words as keys and total frequencies as values.
  • The expression counts[word] = counts.get(word, 0) + 1 accumulates counts correctly for new and repeated words.
  • Tracing each word and dictionary state is an effective way to find word-counting mistakes.