Concepts / Understanding List Methods and Mutability

Understanding List Methods and Mutability

Finding unique words involves a pipeline: read file → split lines into words → check membership to avoid duplicates → sort alphabetically

  • Programming

From Text to Vocabulary

Imagine trying to discover how many distinct words Shakespeare used across all his works. Reading thousands of pages and manually tracking every new word would be impractical. A program can perform this task by moving the text through a sequence of list and string operations. The file is read, each line is split into words, each word is checked against a growing list, duplicates are skipped, and the final list is sorted alphabetically.

The important idea is not one isolated list operation. It is the order of the stages and the way the collection changes as each word is processed.

readsplitcheck membershipsortText fileraw textLinesone line at a timeWord stringssplit piecesUnique wordsduplicates filteredAlphabetical listfinal output
How does data move from the text file to the final list of unique words?

Splitting Lines into Words

A line begins as one string. Calling split() without arguments transforms that string into a list of separate word strings. The split occurs at whitespace boundaries, including spaces, tabs, and newlines. This transformation changes the unit you are processing: instead of handling one complete line, the program can examine each resulting word individually.

One Line Becomes Several Words

Suppose one line contains the words sun moon sun separated by spaces. What does the splitting stage provide for the next stage?

Start with the line: The input is one line treated as a single string.

Apply split(): The line is divided at its whitespace boundaries.

Prepare word-by-word processing: The result is a list containing the separate word strings sun, moon, and sun. The duplicate has not been removed yet; that happens during membership checking.

The splitting stage creates separate word strings so that each word can be checked individually.

splitsplitsplitOne linesun moon sunsunword stringmoonword stringsunword string
How is a line transformed into individual words before duplicate checking begins?

Filtering with Membership Tests

After a line has been split, the program examines each word and compares it with the growing list of unique words. The central check is whether the word is not already in that list. If the word is absent, it is added. If the word is already present, the program skips it. This check-before-adding pattern is the mechanism that prevents duplicates from entering the collection.

inspectnot presentalready presentsuncandidate wordMembership checkalready in list?Unique listadd sunExisting listskip sun
What happens to each word when it is checked against the list?

Tracing Repeated Words

Process the word sequence sun, moon, sun and build a list containing only unique words.

Read sun: The list is empty, so sun is not present. Add it to the list.

Read moon: The list contains sun but not moon. Add moon.

Read sun again: The list already contains sun. Skip this occurrence rather than adding another copy.

The resulting unique list contains sun and moon in their first-appearance order.

Following List State

The growing list is the program's changing state. It begins empty, gains a word when membership testing finds a new word, and stays the same when a repeated word is skipped. This is why tracing a small input is useful: you can identify exactly which word caused the list to change and which word was filtered out.

add sunadd moonskip duplicateEmpty list[]sunafter sunsun, moonafter moonsun, moonafter repeated sun
How does the collection change after each word is examined?

Sorting the Final Collection

After every line has been processed, the list contains unique words, but its order reflects the order in which those words were first encountered. Sorting is a separate final stage. It changes the arrangement of the already filtered collection, making the result easier to search, verify, and understand.

sortsun, moon, lightfirst-appearance orderlight, moon, sunalphabetical order
At what point does alphabetical sorting occur, and how does it change the final collection?

The Romeo Output

What does the complete pipeline produce for romeo.txt according to the expected output?

Read and split: The file content is processed line by line, and each line is split into separate word strings.

Filter duplicates: Each word is checked against the growing unique list before it is added.

Sort: The final collection is arranged alphabetically after duplicate filtering is complete.

The expected output contains 26 unique words: Arise, But, It, Juliet, Who, already, and, breaks, east, envious, fair, grief, is, kill, light, moon, pale, sick, soft, sun, the, through, what, window, with, yonder. Capitalized words appear before lowercase words because uppercase letters come before lowercase letters in ASCII ordering.

Mistakes in the Pipeline

  • Adding every word without checking membership

    The resulting list contains duplicates, so it is not a collection of unique words.

    Fix: Check whether the word is already in the unique list before adding it.

  • Forgetting the final sort

    An unsorted list is harder to search, verify, and understand.

    Fix: Sort the completed unique list after the file has been fully processed.

  • Not reading the entire file

    Words from the unread portion cannot enter the unique list, so the output is incomplete.

    Fix: Trace the file-reading stage and confirm that all lines are processed.

  • Ignoring case sensitivity or punctuation

    The source identifies case sensitivity and punctuation as issues that can affect expected output.

    Fix: When the output differs from expectations, inspect how capitalization and punctuation are represented in the extracted word strings.

  1. Confirm that the file is being read line by line.
  2. Check that split() produces separate word strings.
  3. Trace each word through the membership test.
  4. Record the unique list after additions and skipped duplicates.
  5. Verify that sorting occurs after all lines have been processed.
  6. Compare capitalization and punctuation when the output differs from the expected words.

Practice the Trace

EASY

Trace the sequence moon, sun, moon, light. Write the unique list after each word is examined, then write the list after alphabetical sorting. Identify which word causes no state change.

Hints
  • Begin with an empty list.
  • Add a word only when it is not already present.
  • Sorting is a separate final step.

To implement the complete process, organize the work into five stages: open and read the file, process each line, split each line into words, check each word before adding it to the unique list, and sort and display the completed list. Keeping these stages distinct makes the program easier to trace and makes it clearer where an unexpected result began.

Key Takeaways

  • The word-extraction task is a pipeline: read text, split lines, check membership, and sort the final collection.
  • The membership test prevents duplicates by adding a word only when it is not already in the unique list.
  • The growing list changes when a new word is added and stays the same when a duplicate is skipped.
  • Sorting occurs after duplicate filtering and changes first-appearance order into alphabetical order.
  • Tracing the list after each word helps locate missing duplicate checks, incomplete file reading, missing sorting, and case or punctuation differences.