Concepts / String Methods: split(), strip(), and lower()

String Methods: split(), strip(), and lower()

Finding unique words involves a pipeline: read file → split lines into words → check membership to avoid duplicates → sort alphabetically

  • Programming

From File to Vocabulary

Finding unique words is easier when the task is treated as a pipeline rather than one large operation. Text is read from a file, each line is divided into words, each word is checked against a growing collection, and the final collection is sorted. At every stage, the data changes form. Tracing those changes makes it easier to understand both the result and the location of an error.

Imagine wanting to identify how many distinct words Shakespeare used across his works. Reading thousands of pages and manually recording every new word would be impractical. A word-extraction program can read the text, separate it into words, prevent repeats, and present the vocabulary in an organized form.

readsplitcheck membershipsortText fileraw textLinestext read from fileWordsseparate stringsUnique listduplicates excludedSorted listalphabetical order
How does text move from a file through reading, splitting, duplicate filtering, and sorting to become the final list of unique words?

Splitting Lines into Words

The split() function changes one line string into a list of separate word strings. When called without arguments, it breaks the string at whitespace boundaries such as spaces, tabs, and newlines. This is the first major change in representation: instead of processing a complete line as one unit, the program can process each resulting word separately.

One Line Becomes Several Words

Trace the first transformation in the pipeline for a line containing several words separated by whitespace.

Read a line: The program receives a line as text from the file.

Apply split(): The line is transformed into a list of separate word strings.

Process each word: The program can now examine words one at a time and decide whether each belongs in the growing unique list.

The program has moved from line-by-line text to word-by-word processing.

split()A line of textone stringWord stringsa list of separate pieces
How does a line change when it is divided into individual word strings?

Filtering Repeated Words

After splitting, the program examines each word against the unique-word list built so far. The membership test is the central duplicate-filtering mechanism. If the word is already in the list, the program skips it. If the word is not in the list, the program adds it. The list therefore changes only when a newly encountered word passes the membership check.

inspectyesnoExtracted wordMembership testalready in list?Skip wordduplicateAdd wordnew word
What happens to each extracted word when it is checked against the existing list?
new wordnew wordduplicateUnique listemptyFirst wordone itemNew wordlist growsRepeated wordlist unchanged
How does the collection of unique words change after each word is processed?

What do you think happens?

A word is encountered for the second time. What should happen to the unique-word list?

  • The word is added again
  • The word is skipped
  • The entire list is cleared
Reveal answer

Answer: The word is skipped.

The membership test checks whether the word is already in the list. A word that is already present is not added again.

Normalizing Word Text

The complete word-extraction task also requires attention to the form of each word. The source identifies case sensitivity and punctuation as issues that can affect the result. The methods named in this article, strip() and lower(), belong to this normalization stage, while split() performs the separation into word strings. Before deciding that two extracted strings are duplicates, inspect whether differences in capitalization or punctuation are being treated as meaningful in the intended output.

split()normalizenormalizeRaw linefile textWord stringsseparate piecesStripped textwhitespace consideredLowercase textcase considered
How can a raw line or token pass through separation and text-normalization stages before duplicate comparison?

Sorting the Final Collection

Sorting occurs after the program has finished reading the file and filtering duplicates. This timing matters: the program first builds the collection of unique words, then orders that collection alphabetically. An unsorted list reflects the order in which words were first encountered. A sorted list is easier to search, verify, and understand because its order is based on the words themselves rather than the source file's sequence.

From First Appearance to Alphabetical Order

Trace the final stages of a word-extraction pipeline.

Finish reading: The program has processed the file's lines and considered the extracted words.

Complete membership filtering: Words already in the unique list have been skipped, while new words have been added.

Sort the unique list: The remaining words are reordered alphabetically.

Display the result: The ordered list can now be searched and checked more easily.

Sorting is a final organization step, not the step that removes duplicates.

sortFirst appearancefile orderAlphabetical ordersorted unique list
When does sorting occur, and how does the order of the unique-word list change?
StageWhat the data representsMain question
ReadingText from the fileWas the entire file read?
SplittingSeparate word stringsDid each line become individual words?
Membership testingThe growing unique listHas this word already been added?
SortingThe completed unique listAre the words in alphabetical order?

Each stage has a different responsibility in the pipeline.

Expected Romeo Output

For romeo.txt, the expected result is a list of 26 unique words in sorted order. The source output places capitalized words such as Arise and But before lowercase words such as already. This illustrates that the displayed alphabetical order is affected by capitalization rather than being a case-insensitive vocabulary order.

Output
Arise, But, It, Juliet, Who, already, and, breaks, east, envious, fair, grief, is, kill, light, moon, pale, sick, soft, sun, the, through, what, window, with, yonder

Mistakes in the Pipeline

  • Adding every extracted word without a membership test

    The collection is intended to contain unique words, so repeated entries should not accumulate.

    Fix: Check whether the word is already in the unique list before adding it.

  • Sorting too early or forgetting to sort

    First-appearance order does not provide the searchable, verifiable alphabetical organization expected from the final result.

    Fix: Sort after the full file has been read and duplicate filtering is complete.

  • Reading only part of the file

    Words in unread portions of the file never reach the splitting or membership-testing stages.

    Fix: Trace whether the program processes the entire file.

  • Ignoring case sensitivity or punctuation

    Differences in capitalization or punctuation can affect which strings are compared as duplicates.

    Fix: Inspect the extracted strings before membership testing and decide how text normalization should be handled.

When the result is unexpected, do not begin by inspecting only the final list. Check the pipeline in order: confirm that the file was read completely, inspect the pieces produced from each line, observe whether the unique list changes only for new words, and verify that sorting happens at the end. This narrows the problem to a specific stage instead of treating the whole program as one unexplained operation.

Pipeline Practice

EASY

Describe what should happen to a word as it moves through the pipeline. Your answer should name the stage where a line becomes separate words, the test that prevents duplicates, and the stage that changes first-appearance order into alphabetical order.

Hints
  • The line-to-word transformation uses split().
  • Duplicate prevention depends on checking membership before adding.
  • Sorting belongs after the unique list has been built.
MEDIUM

A program returns an unsorted list containing repeated words. Identify two separate pipeline failures that could explain the result, then state which stage should be inspected for each failure.

Hints
  • Repeated words point to the membership-testing stage.
  • An unsorted result points to the final sorting stage.
  • Consider whether the entire file was read as a separate possibility.

Key Takeaways

  1. The word-extraction pipeline reads file text, splits lines into words, filters duplicates, and sorts the final list.
  2. split() changes one line string into separate word strings so that words can be processed individually.
  3. Membership testing is the step that prevents a word already in the unique list from being added again.
  4. Sorting occurs after duplicate filtering and changes first-appearance order into alphabetical order.
  5. Case sensitivity, punctuation, incomplete file reading, missing membership checks, and missing sorting are important debugging points.

Key Takeaways

  • A word-extraction program is easiest to understand as a sequence of data transformations.
  • split() enables word-by-word processing by turning a line into separate strings.
  • The membership test filters duplicates while the unique list is being built.
  • Sorting is a final step that makes the completed vocabulary easier to search and verify.
  • Tracing the state after each stage reveals where missing words, repeated words, or unexpected ordering originate.