Concepts / Dictionaries in Python

Dictionaries in Python

Use a dictionary to count word frequencies by checking if each word exists and incrementing its count

  • Programming

From Text to Useful Results

Imagine a text file containing thousands of words. Printing every word and its count would produce a large, unsorted result. A more useful goal is to identify the words that appear most often. The solution has two phases: first, build a frequency dictionary; second, transform and sort its contents so the largest frequencies appear first.

The dictionary handles counting. The list of tuples handles sorting by frequency.

Building the Frequency Map

A frequency dictionary uses each unique word as a key and its number of appearances as the value. The text is processed line by line. Each line is prepared by removing punctuation and converting it to lowercase, then each word is examined. If the word is not already in the dictionary, its count is initialized to 1. If it is already present, its existing count is increased by 1.

readinitializeread againincrementEmpty counts{}lovenew wordlove1loveexisting wordlove2
How does the dictionary change as each word is read, especially when a word is new versus when it has already been counted?

Counting a Short Word Sequence

Track the frequencies in the sequence: love, hope, love, hope, love.

Read love: The dictionary does not contain love, so create the entry love: 1.

Read hope: The dictionary does not contain hope, so create the entry hope: 1.

Read love again: The word already exists, so increase its count from 1 to 2.

Read hope again: The word already exists, so increase its count from 1 to 2.

Read love once more: The existing love count increases from 2 to 3.

The resulting frequency dictionary is {love: 3, hope: 2}. The words are keys, and their occurrence totals are values.

maps tomaps tolovekey3valuehopekey2value
What does the dictionary contain, and how is each word connected to its frequency count?

The Membership Decision

The counting loop depends on one membership check: does the current word already exist as a key in the dictionary? A new word receives a count of 1. An existing word receives its previous count plus 1. Repeating this decision for every word builds a complete frequency map of the document.

inspectnoyesCurrent wordWord in counts?1new countcount + 1existing count
How does the membership check determine whether to create a new count or increment an existing one?

Reordering Entries for Sorting

The frequency dictionary has the right information, but the desired ordering is by value rather than by key. To make that possible, process each word-count pair and create a tuple with the count first and the word second. For example, the dictionary entry love: 47 becomes the tuple (47, love). Collect these tuples in a list.

reverse orderappendlove: 47word: count47, lovecount, word[(47, love)]list of tuples
How do dictionary key-value pairs become a list of frequency-first tuples that can be sorted?

This reordering is the central technique. Sorting a sequence of tuples compares the first tuple element by default. By placing the frequency first, sorting compares frequencies. The word is placed second because it is still needed for the final output and can be considered after the frequency when appropriate.

sortsort(love, 47)word firstalphabeticalkey controls comparison(47, love)count firstfrequency ordercount controls comparison
What changes when the sort uses each tuple's frequency value rather than its word key?

Finding the Most Common Words

After the frequency-first tuples have been collected, sort the list in reverse order. Reverse sorting places the highest frequency at the beginning. The first ten entries can then be selected with a slice representing the first ten positions. Iterating through that slice produces the most common words and their counts.

Selecting the Top Three

Suppose the transformed tuples are [(3, love), (7, hope), (5, peace), (2, light)]. Sort them in reverse order and select the top three.

Sort in reverse order: Because the frequency is the first tuple element, reverse sorting places 7 first, then 5, then 3, then 2.

Read the leading entries: The sorted sequence is [(7, hope), (5, peace), (3, love), (2, light)].

Take the first three: A first-three slice selects (7, hope), (5, peace), and (3, love).

The three most common words are hope with frequency 7, peace with frequency 5, and love with frequency 3.

reverse sortsliceFrequency tuples(3, love), (7, hope), (5,peace), (2, light)Sorted tuples(7, hope), (5, peace), (3,love), (2, light)Leading resultstop three
How does sorting the frequency tuples and taking the leading entries identify the most common words?

The final tuple order remains frequency first, word second. Therefore, the output loop reads the frequency as the first item and the word as the second item.

Practical Boundaries

The complete process is a clear pipeline: prepare and read the text, update the dictionary for every word, convert each word-count entry into a count-word tuple, sort the tuple list in reverse order, slice the leading results, and print or otherwise process those results.

SituationResult
Fewer unique words than the requested top countThe slice returns all available tuples without an error
Two words have the same frequencyThe stable sort preserves their relative order from before sorting
Empty fileThe dictionary and tuple list are empty, so the final loop prints nothing
Word appears onceIts dictionary value is initialized to 1

Predictable behavior at the boundaries of the word-frequency process

Mistakes with Frequency Sorting

  • Sorting the dictionary representation directly when the goal is frequency order

    The key-value structure does not place the frequency first for sequence sorting.

    Fix: Transform each entry into a tuple with the count first and the word second.

  • Creating tuples in word-count order

    The word is the first tuple element, so normal tuple sorting compares the word before the frequency.

    Fix: Create (47, love) so the frequency controls the primary sort.

  • Sorting in ascending order when looking for the most common words

    The leading entries would represent the least frequent words rather than the most frequent ones.

    Fix: Sort in reverse order so the highest frequencies appear at the beginning.

  • Forgetting that the final tuple is frequency first

    The transformation deliberately reversed the original dictionary entry order.

    Fix: Interpret the first tuple item as the frequency and the second as the word.

Practice the Pipeline

MEDIUM

A document produces this frequency dictionary: {river: 4, stone: 9, cloud: 2, field: 6}. Describe the tuple list you would create, the order after reverse sorting, and the result of selecting the top two entries.

Hints
  • Put each frequency before its word.
  • Place the largest frequency first after reverse sorting.
  • The top-two selection uses the first two sorted tuples.

What do you think happens?

What happens when the document has fewer than ten unique words but the process selects the first ten entries?

  • The process raises an error
  • The available entries are returned
  • The dictionary is filled with empty entries
Reveal answer

Answer: The available entries are returned.

Selecting the first ten positions returns all available tuples when the list contains fewer than ten entries.

Summary

  1. Use a dictionary to map each word to the number of times it appears.
  2. For every word, initialize a missing entry to 1 or increment an existing entry.
  3. Convert each word-count pair into a frequency-word tuple so the frequency becomes the primary sort element.
  4. Sort the tuple list in reverse order to place the most common words first.
  5. Slice the beginning of the sorted list to select the top results.

Key Takeaways

  • Dictionaries efficiently build a word-to-frequency map.
  • The membership check distinguishes a new word from a word already counted.
  • Reversing each entry into a frequency-word tuple makes sorting operate on counts.
  • Reverse sorting places the most frequent words at the front.
  • Slicing the sorted list extracts the requested number of leading results.