Dictionaries in Python
Use a dictionary to count word frequencies by checking if each word exists and incrementing its count
From Text to Useful Results
Imagine a text file containing thousands of words. Printing every word and its count would produce a large, unsorted result. A more useful goal is to identify the words that appear most often. The solution has two phases: first, build a frequency dictionary; second, transform and sort its contents so the largest frequencies appear first.
The dictionary handles counting. The list of tuples handles sorting by frequency.
Building the Frequency Map
A frequency dictionary uses each unique word as a key and its number of appearances as the value. The text is processed line by line. Each line is prepared by removing punctuation and converting it to lowercase, then each word is examined. If the word is not already in the dictionary, its count is initialized to 1. If it is already present, its existing count is increased by 1.
Counting a Short Word Sequence
Track the frequencies in the sequence: love, hope, love, hope, love.
Read love: The dictionary does not contain love, so create the entry love: 1.
Read hope: The dictionary does not contain hope, so create the entry hope: 1.
Read love again: The word already exists, so increase its count from 1 to 2.
Read hope again: The word already exists, so increase its count from 1 to 2.
Read love once more: The existing love count increases from 2 to 3.
The resulting frequency dictionary is {love: 3, hope: 2}. The words are keys, and their occurrence totals are values.
The Membership Decision
The counting loop depends on one membership check: does the current word already exist as a key in the dictionary? A new word receives a count of 1. An existing word receives its previous count plus 1. Repeating this decision for every word builds a complete frequency map of the document.
Reordering Entries for Sorting
The frequency dictionary has the right information, but the desired ordering is by value rather than by key. To make that possible, process each word-count pair and create a tuple with the count first and the word second. For example, the dictionary entry love: 47 becomes the tuple (47, love). Collect these tuples in a list.
This reordering is the central technique. Sorting a sequence of tuples compares the first tuple element by default. By placing the frequency first, sorting compares frequencies. The word is placed second because it is still needed for the final output and can be considered after the frequency when appropriate.
Finding the Most Common Words
After the frequency-first tuples have been collected, sort the list in reverse order. Reverse sorting places the highest frequency at the beginning. The first ten entries can then be selected with a slice representing the first ten positions. Iterating through that slice produces the most common words and their counts.
Selecting the Top Three
Suppose the transformed tuples are [(3, love), (7, hope), (5, peace), (2, light)]. Sort them in reverse order and select the top three.
Sort in reverse order: Because the frequency is the first tuple element, reverse sorting places 7 first, then 5, then 3, then 2.
Read the leading entries: The sorted sequence is [(7, hope), (5, peace), (3, love), (2, light)].
Take the first three: A first-three slice selects (7, hope), (5, peace), and (3, love).
The three most common words are hope with frequency 7, peace with frequency 5, and love with frequency 3.
The final tuple order remains frequency first, word second. Therefore, the output loop reads the frequency as the first item and the word as the second item.
Practical Boundaries
The complete process is a clear pipeline: prepare and read the text, update the dictionary for every word, convert each word-count entry into a count-word tuple, sort the tuple list in reverse order, slice the leading results, and print or otherwise process those results.
| Situation | Result |
|---|---|
| Fewer unique words than the requested top count | The slice returns all available tuples without an error |
| Two words have the same frequency | The stable sort preserves their relative order from before sorting |
| Empty file | The dictionary and tuple list are empty, so the final loop prints nothing |
| Word appears once | Its dictionary value is initialized to 1 |
Predictable behavior at the boundaries of the word-frequency process
Mistakes with Frequency Sorting
Sorting the dictionary representation directly when the goal is frequency order
The key-value structure does not place the frequency first for sequence sorting.
Fix:
Transform each entry into a tuple with the count first and the word second.Creating tuples in word-count order
The word is the first tuple element, so normal tuple sorting compares the word before the frequency.
Fix:
Create (47, love) so the frequency controls the primary sort.Sorting in ascending order when looking for the most common words
The leading entries would represent the least frequent words rather than the most frequent ones.
Fix:
Sort in reverse order so the highest frequencies appear at the beginning.Forgetting that the final tuple is frequency first
The transformation deliberately reversed the original dictionary entry order.
Fix:
Interpret the first tuple item as the frequency and the second as the word.
Practice the Pipeline
A document produces this frequency dictionary: {river: 4, stone: 9, cloud: 2, field: 6}. Describe the tuple list you would create, the order after reverse sorting, and the result of selecting the top two entries.
Hints
- Put each frequency before its word.
- Place the largest frequency first after reverse sorting.
- The top-two selection uses the first two sorted tuples.
What do you think happens?
What happens when the document has fewer than ten unique words but the process selects the first ten entries?
Reveal answer
Answer: The available entries are returned.
Selecting the first ten positions returns all available tuples when the list contains fewer than ten entries.
Summary
- Use a dictionary to map each word to the number of times it appears.
- For every word, initialize a missing entry to 1 or increment an existing entry.
- Convert each word-count pair into a frequency-word tuple so the frequency becomes the primary sort element.
- Sort the tuple list in reverse order to place the most common words first.
- Slice the beginning of the sorted list to select the top results.
Key Takeaways
- Dictionaries efficiently build a word-to-frequency map.
- The membership check distinguishes a new word from a word already counted.
- Reversing each entry into a frequency-word tuple makes sorting operate on counts.
- Reverse sorting places the most frequent words at the front.
- Slicing the sorted list extracts the requested number of leading results.