Concepts / Building Dictionaries from Text

Building Dictionaries from Text

Convert a dictionary to a list of (value, key) tuples so that sorting operates on the values (frequencies) rather than the keys (words).

  • Programming

From Words to Rankings

A word-frequency dictionary answers how often each word appears, but it does not directly answer which words are most common. The dictionary uses words as keys and counts as values. To rank the words by frequency, restructure each entry so the count comes first. This changes the sorting problem from comparing words first to comparing counts first.

entryrestructureentryrestructureword-frequencydictionaryword → countthe5(5, the)count → wordto5(5, to)count → word
How does each word-frequency entry change when the count becomes the first tuple element?

Counting Words First

The complete workflow begins with a text file. As words are read, the program builds a dictionary: each word becomes a key, and its frequency becomes the associated value. The first encounter with a word establishes its count, while later encounters increase that count. After the dictionary is complete, the program creates a separate list of tuples, placing each count before its word. The original dictionary is not modified.

first encounterencounter againcontinue readingempty dictionary{}wordcount establishedsame wordcount increasedfrequency dictionarywords → counts
How does the dictionary change when a word appears for the first time and when it appears again?

A Small Frequency Dictionary

Suppose the completed dictionary contains the entries the: 5, to: 5, and: 4.

Read the dictionary: The keys are words and the values are their counts.

Create tuples: Rearrange each entry so the count is first: (5, the), (5, to), and (4, and).

Prepare for sorting: Because the count is the first tuple element, it becomes the primary value used during sorting.

The new list contains sortable count-and-word tuples while the original dictionary remains separate.

Reverse Sorting and Ties

Python compares tuple elements from left to right. With count-and-word tuples, the count is compared first. Calling sort with reverse=True places larger counts before smaller counts. If two tuples have the same count, Python compares their second elements, the words. Because the whole sort is reversed, the alphabetically later word comes first among tied entries.

count descendsequal count; word descendscount descends(8, word)higher count(5, to)tie: word lateralphabetically(5, the)tie: word earlieralphabetically(4, word)lower count
After reverse sorting, how are equal-frequency words ordered?

Resolving an Equal-Frequency Tie

Compare the tuples (5, the) and (5, to) when the list is sorted in reverse order.

Compare the first elements: Both first elements are 5, so neither tuple wins on frequency.

Compare the second elements: Python compares the words. The word to comes later alphabetically than the word the.

Apply reverse ordering: Reverse ordering places the alphabetically later tied word first.

The tied entries appear as (5, to), then (5, the).

Selecting the Top Results

Once the tuple list is sorted, list slicing selects the results to display. The slice lst[:10] starts at index 0 and stops before index 10, selecting indices 0 through 9. In a loop, tuple unpacking assigns the first tuple element to count and the second to word. Changing the slice limit changes how many leading results are displayed without changing the sorted list itself.

next resultnext resultstop before index 3not selectedindex 0(12, alpha)index 1(10, beta)index 2(9, gamma)index 3(8, delta)slice boundaryN = 3
Which portion of the sorted tuple list is selected by a top-N slice?
first elementsecond elementformatformat(61, i)one selected resultcount61display resultcount and wordwordi
How does each two-item tuple become separate count and word values?

Changing N

A sorted list contains twelve count-and-word tuples. Compare a top-3 selection with a top-10 selection.

Use the top-3 slice: The first three tuples, at indices 0, 1, and 2, are selected.

Use the top-10 slice: The first ten tuples, at indices 0 through 9, are selected.

Keep the ranking unchanged: Both selections come from the same sorted list; only the endpoint of the slice changes.

The slice controls the number of displayed leading results.

Following the Full Workflow

count wordsrestructuresort reverseselect Nunpack and displaytext fileinput wordsfrequency dictionaryword → countcount-word tuples(count, word)reverse-sorted listhighest counts firsttop-N sliceleading resultsranked outputcount and word
How does data move through the complete word-frequency workflow?

The full program combines several separate operations. It reads a text file, counts the words into a dictionary, converts the dictionary entries into count-and-word tuples, sorts those tuples in reverse order, selects the leading results, and displays each selected count with its word. The sorting stage ranks by frequency, while the tuple's second element provides deterministic ordering for ties.

In the Romeo and Juliet text example, the output lists the ten most frequent words with their counts. The word i appears 61 times and and appears 42 times. These results are ordered from the highest frequency downward because the count was placed first in each tuple before reverse sorting.

Keep the stages conceptually separate: build the dictionary, create the sortable tuple list, sort it, select the desired prefix, and unpack the selected tuples for display. This makes the transformation easy to trace and leaves the original dictionary unchanged.

Mistakes to Avoid

  • Leaving the word first in each tuple.

    Python compares the first tuple element first, so the words become the primary sorting values instead of the frequencies.

    Fix: Use (count, word) so frequency controls the main ordering.

  • Expecting equal-frequency words to keep an arbitrary order.

    Python compares the second tuple elements when the counts match.

    Fix: Expect reverse alphabetical ordering for tied words when reverse sorting is used.

  • Reading lst[:10] as ten through twenty.

    The slice begins at index 0 and stops before index 10.

    Fix: Interpret it as indices 0 through 9.

  • Reversing the unpacking order.

    The tuple is organized as count first and word second.

    Fix: Assign the first element to count and the second element to word.

  • Assuming sorting changes the original dictionary into a ranked dictionary.

    The tuple list is a new sortable data structure.

    Fix: Remember that the dictionary and the sorted list are separate structures.

Apply the Pattern

MEDIUM

A frequency dictionary contains the entries alpha: 7, beta: 7, gamma: 10, and delta: 3. Describe the count-and-word tuple list, its order after reverse sorting, and the entries selected by a top-2 slice.

Hints
  • Put each count before its word.
  • Place the largest count first.
  • For the two entries with count 7, compare alpha and beta as the secondary values.
  • A top-2 slice selects indices 0 and 1.

What do you think happens?

Which tied entry appears first after reverse sorting: (5, the) or (5, to)?

  • (5, the)
  • (5, to)
  • They remain in dictionary order
Reveal answer

Answer: (5, to)

The counts tie, so Python compares the words. Since to comes later alphabetically than the, reverse sorting places to first.

Key Takeaways

  1. A word-frequency dictionary stores words as keys and counts as values.
  2. Convert entries to (count, word) tuples so sorting uses frequency as the primary criterion.
  3. Reverse tuple sorting places larger counts first and uses the word as a reverse-alphabetical tiebreaker.
  4. The slice lst[:10] selects the first ten sorted tuples, while tuple unpacking separates each count from its word.
  5. The complete workflow moves from file input to counting, tuple conversion, sorting, slicing, and formatted output.

Key Takeaways

  • Restructure dictionary entries as (value, key), or specifically (count, word), before sorting.
  • Reverse sorting compares counts first and uses words to resolve equal-count ties.
  • Use a top-N slice to select only the leading results.
  • Use tuple unpacking to handle each count and word separately during output.
  • The pattern provides deterministic word-frequency rankings without changing the original dictionary.