List Methods and Sorting
Convert a dictionary to a list of (value, key) tuples so that sorting operates on the values (frequencies) rather than the keys (words).
From Counts to Rankings
A word-frequency dictionary answers the question “How many times did each word appear?” Its words are keys and its counts are values. To answer a different question—“Which words appeared most often?”—you need the counts to control the ordering. The useful pattern is to build a list of tuples in the form (count, word), sort that list, and then display the first few tuples.
Put the frequency first in each tuple. The first tuple element becomes the primary sorting criterion.
The Data Flow
The complete process has distinct stages. The program reads text, builds a dictionary of word counts, restructures each dictionary entry as a (count, word) tuple, sorts the tuple list in reverse order, selects a prefix of that list, and unpacks each selected tuple for display. Keeping these stages separate makes the final output easier to understand and debug.
Rebuilding the Sortable Structure
A dictionary stores words as keys and frequencies as values. Sorting the dictionary by its keys would organize words, not frequencies. Instead, create a separate list and place each value before its key. A dictionary entry such as word: frequency becomes the tuple (frequency, word). The original dictionary remains unchanged; the list is a new sortable representation.
Dictionary Entries Become Sortable Tuples
Restructure the frequency entries the: 5, to: 5, and i: 61 so that the counts become the primary sort elements.
Start with key-value entries: Each entry has a word as its key and a frequency as its value.
Reverse the positions: Represent each entry as (count, word), producing (5, the), (5, to), and (61, i).
Prepare for sorting: The count is now the first tuple element, so tuple sorting compares frequency before the word.
The sortable list is [(5, the), (5, to), (61, i)] before ordering.
Reverse Order and Ties
Calling sort(reverse=True) sorts the entire tuple in descending order. Python first compares the tuple's first element, so larger counts come first. If two counts match, Python compares the second elements. Because the whole tuple is being sorted in reverse order, the words are then ordered in reverse alphabetical order.
A Tie Uses the Word
Order the tuples (5, the), (5, to), and (61, i) with reverse sorting.
Compare counts: The tuple containing 61 comes before the tuples containing 5 because 61 is the larger first element.
Resolve the equal counts: The remaining tuples both begin with 5, so Python compares the words to and the.
Apply reverse alphabetical order: to comes after the alphabetically earlier word the, so to appears first when the order is reversed.
The order is (61, i), (5, to), (5, the).
Selecting the Top Results
After sorting, the most frequent words are at the beginning of the list. The slice lst[:10] starts at index 0 and stops before index 10, so it selects indices 0 through 9. In a loop written as for count, word in lst[:10]:, each tuple is unpacked: its first element is assigned to count and its second element to word.
61 i
42 and
5 toTracing a Complete Program
A complete word-frequency program combines file reading, dictionary building, tuple construction, sorting, slicing, and output formatting. The important point is not one particular file-reading implementation; it is the order in which the data changes shape. Text begins as input, becomes word counts, becomes sortable tuples, and finally becomes a short ranked display.
counts = {"i": 61, "and": 42, "the": 5, "to": 5} ranked = [] for word, count in counts.items(): ranked.append((count, word)) ranked.sort(reverse=True) for count, word in ranked[:3]: print(count, word)
In the source program's Romeo and Juliet result, i appears 61 times and and appears 42 times. The displayed words are ordered with the highest frequency first. The smaller example above follows the same ranking logic while using fewer entries so that every stage can be traced directly.
Mistakes with Ranked Lists
Sorting the word keys instead of the frequency values
The question is which words have the largest counts, so alphabetical word order is not the primary criterion.
Fix:
Create (count, word) tuples so the count is the first element.Putting the word before the count
Tuple sorting would compare the word first, making the frequency secondary.
Fix:
Use the (count, word) arrangement.Expecting ties to remain in an unspecified order
Python compares the second tuple elements when the counts match.
Fix:
Expect reverse alphabetical ordering of the words when sort(reverse=True) is used.Reading lst[:10] as including index 10
The ending index of this slice is excluded.
Fix:
Interpret lst[:10] as indices 0 through 9.Forgetting the tuple's element order during unpacking
The first loop variable receives the first tuple element.
Fix:
Use for count, word in ranked[:10]: when the tuples are (count, word).
Practice the Transformation
Given the frequency entries {"blue": 3, "red": 7, "green": 3}, describe the tuple list you would build, the order produced by reverse sorting, and the output produced by displaying the first two tuples.
Hints
- Place each frequency before its word.
- Compare the frequencies first.
- For the equal frequencies, compare the words in reverse alphabetical order.
- The first two positions are selected by the slice [:2].
What do you think happens?
After converting the entries to tuples and sorting in reverse order, which word comes first among the two words with frequency 3?
Reveal answer
Answer: green
Both tuples have the same first element, 3. Python compares the words next, and reverse alphabetical order places green before blue.
Key Takeaways
- A frequency dictionary stores words as keys and counts as values, but ranking by frequency requires a new sortable structure.
- Represent each entry as (count, word) so the count is the primary tuple-sorting element.
- sort(reverse=True) puts larger counts first and uses reverse alphabetical word order to resolve equal counts.
- The slice lst[:10] selects the first ten positions, and tuple unpacking assigns each tuple's count and word to loop variables.
- The pattern separates counting, restructuring, sorting, selecting, and displaying into clear stages.
Key Takeaways
- Convert dictionary entries from word: count into (count, word) tuples.
- Sort the tuple list in reverse order to place the highest frequencies first.
- When counts tie, Python compares the words and reverse sorting produces reverse alphabetical order.
- Use slicing to choose the top N tuples and tuple unpacking to display their elements.
- The complete word-frequency workflow moves through counting, tuple conversion, sorting, slicing, and output.