Understanding List Methods and Mutability
Finding unique words involves a pipeline: read file → split lines into words → check membership to avoid duplicates → sort alphabetically
From Text to Vocabulary
Imagine trying to discover how many distinct words Shakespeare used across all his works. Reading thousands of pages and manually tracking every new word would be impractical. A program can perform this task by moving the text through a sequence of list and string operations. The file is read, each line is split into words, each word is checked against a growing list, duplicates are skipped, and the final list is sorted alphabetically.
The important idea is not one isolated list operation. It is the order of the stages and the way the collection changes as each word is processed.
Splitting Lines into Words
A line begins as one string. Calling split() without arguments transforms that string into a list of separate word strings. The split occurs at whitespace boundaries, including spaces, tabs, and newlines. This transformation changes the unit you are processing: instead of handling one complete line, the program can examine each resulting word individually.
One Line Becomes Several Words
Suppose one line contains the words sun moon sun separated by spaces. What does the splitting stage provide for the next stage?
Start with the line: The input is one line treated as a single string.
Apply split(): The line is divided at its whitespace boundaries.
Prepare word-by-word processing: The result is a list containing the separate word strings sun, moon, and sun. The duplicate has not been removed yet; that happens during membership checking.
The splitting stage creates separate word strings so that each word can be checked individually.
Filtering with Membership Tests
After a line has been split, the program examines each word and compares it with the growing list of unique words. The central check is whether the word is not already in that list. If the word is absent, it is added. If the word is already present, the program skips it. This check-before-adding pattern is the mechanism that prevents duplicates from entering the collection.
Tracing Repeated Words
Process the word sequence sun, moon, sun and build a list containing only unique words.
Read sun: The list is empty, so sun is not present. Add it to the list.
Read moon: The list contains sun but not moon. Add moon.
Read sun again: The list already contains sun. Skip this occurrence rather than adding another copy.
The resulting unique list contains sun and moon in their first-appearance order.
Following List State
The growing list is the program's changing state. It begins empty, gains a word when membership testing finds a new word, and stays the same when a repeated word is skipped. This is why tracing a small input is useful: you can identify exactly which word caused the list to change and which word was filtered out.
Sorting the Final Collection
After every line has been processed, the list contains unique words, but its order reflects the order in which those words were first encountered. Sorting is a separate final stage. It changes the arrangement of the already filtered collection, making the result easier to search, verify, and understand.
The Romeo Output
What does the complete pipeline produce for romeo.txt according to the expected output?
Read and split: The file content is processed line by line, and each line is split into separate word strings.
Filter duplicates: Each word is checked against the growing unique list before it is added.
Sort: The final collection is arranged alphabetically after duplicate filtering is complete.
The expected output contains 26 unique words: Arise, But, It, Juliet, Who, already, and, breaks, east, envious, fair, grief, is, kill, light, moon, pale, sick, soft, sun, the, through, what, window, with, yonder. Capitalized words appear before lowercase words because uppercase letters come before lowercase letters in ASCII ordering.
Mistakes in the Pipeline
Adding every word without checking membership
The resulting list contains duplicates, so it is not a collection of unique words.
Fix:
Check whether the word is already in the unique list before adding it.Forgetting the final sort
An unsorted list is harder to search, verify, and understand.
Fix:
Sort the completed unique list after the file has been fully processed.Not reading the entire file
Words from the unread portion cannot enter the unique list, so the output is incomplete.
Fix:
Trace the file-reading stage and confirm that all lines are processed.Ignoring case sensitivity or punctuation
The source identifies case sensitivity and punctuation as issues that can affect expected output.
Fix:
When the output differs from expectations, inspect how capitalization and punctuation are represented in the extracted word strings.
- Confirm that the file is being read line by line.
- Check that split() produces separate word strings.
- Trace each word through the membership test.
- Record the unique list after additions and skipped duplicates.
- Verify that sorting occurs after all lines have been processed.
- Compare capitalization and punctuation when the output differs from the expected words.
Practice the Trace
Trace the sequence moon, sun, moon, light. Write the unique list after each word is examined, then write the list after alphabetical sorting. Identify which word causes no state change.
Hints
- Begin with an empty list.
- Add a word only when it is not already present.
- Sorting is a separate final step.
To implement the complete process, organize the work into five stages: open and read the file, process each line, split each line into words, check each word before adding it to the unique list, and sort and display the completed list. Keeping these stages distinct makes the program easier to trace and makes it clearer where an unexpected result began.
Key Takeaways
- The word-extraction task is a pipeline: read text, split lines, check membership, and sort the final collection.
- The membership test prevents duplicates by adding a word only when it is not already in the unique list.
- The growing list changes when a new word is added and stays the same when a duplicate is skipped.
- Sorting occurs after duplicate filtering and changes first-appearance order into alphabetical order.
- Tracing the list after each word helps locate missing duplicate checks, incomplete file reading, missing sorting, and case or punctuation differences.