Concepts / Processing Structured Text Files

Processing Structured Text Files

startswith() checks if a string begins with a specific prefix and returns True or False—it does not search for the substring anywhere in the string.

  • Programming

A Line Has More Than You See

A line read from a text file usually contains an invisible newline character at its end. That character is part of the string even though it is not visible on the screen. If you pass the line directly to print(), the existing newline and print()'s automatic newline work together, creating an empty line between displayed lines.

What do you think happens?

Suppose line contains the text From: sender followed by a newline character. What happens when print(line) runs?

  • The text appears with no newline
  • The text appears with normal single spacing
  • The text is followed by an extra blank line
Reveal answer

Answer: The text is followed by an extra blank line.

The newline already stored in the file line moves to the next line, and print() adds another newline after printing the string.

print()print()file lineFrom: sender + newlineclean lineFrom: senderprinted textFrom: sender + extranewlinesingle-spaced textFrom: sender + printnewline
What changes when a file line is printed before and after its trailing newline is removed?

Prefix Checks with startswith()

The startswith() method checks whether a string begins with a specified prefix. It returns True when the beginning matches and False otherwise. It does not search for the prefix anywhere else in the string.

python
Output
True
False

This distinction matters when a file uses prefixes to identify different kinds of lines. For example, a structured text file might use From:, Subject:, or Date: at the beginning of lines. A conditional using startswith() can select only the kind of line you need instead of matching the same text when it appears in the middle or at the end.

Removing Trailing Characters

The rstrip() method removes trailing whitespace and newline characters from a string. It operates at the end of the string, so the meaningful text at the beginning remains unchanged. It can remove multiple trailing whitespace characters, including spaces, tabs, or newlines.

python
rstrip()raw lineSubject: Report + spaces +newlineclean lineSubject: Report
Which characters does rstrip() remove, and which part of the line remains?

The Filtering Pipeline

The useful workflow combines both methods. For each line, first remove trailing whitespace with rstrip(). Then check the cleaned line with startswith(). If the Boolean result is True, display the line. If it is False, skip it. Calling rstrip() first creates a clean value for both the test and the output, even though a newline at the end would not change a prefix check at the beginning.

readcleanTrueFalsefile lineone line from the fileclean linerstrip() appliedprefix resultstartswith("From:")display lineonly when result is Trueskipped linewhen result is False
How does each file line move through cleaning, prefix checking, and display?

for line in file: clean_line = line.rstrip() if clean_line.startswith("From:"): print(clean_line)

Output
From: sender@example.com
From: archive@example.com

In this generated example, lines beginning with Subject: or other text are not displayed. The selected From: lines appear with single spacing because rstrip() removed their stored newline before print() added its own.

Structured Text in Practice

Structured text files often give different meanings to lines based on their prefixes. Email files can use prefixes such as From:, Subject:, To:, and Date:. Log files can use ERROR:, WARNING:, or INFO: to mark severity levels. The same processing pattern works in both cases: clean each line, test its prefix, and display or process only the lines that match.

Selecting Warning Lines

Process lines from a log-like file and display only lines that begin with WARNING:.

Clean: Call rstrip() on the line so the trailing newline and other trailing whitespace do not remain in the displayed result.

Check: Call startswith("WARNING:") on the cleaned line. The result is True only when WARNING: is at the beginning.

Display: Use print() inside the conditional so only matching lines are displayed.

The workflow extracts WARNING: lines and displays them without the extra blank lines caused by printing the original file lines.

Mistakes to Avoid

  • Treating startswith() as a general search method

    The target text is not at the beginning of the string, so startswith() returns False.

    Fix: Use startswith() when the prefix position matters. It checks only the beginning of the string.

  • Printing the original file line directly

    The file line already contains a newline, and print() adds another newline.

    Fix: Call line.rstrip() before printing.

  • Using rstrip() after the line has already been printed

    The extra spacing has already occurred by the time the cleaned value is created.

    Fix: Clean the line first, then pass the cleaned value to startswith() and print().

  • Checking one value but displaying a different uncleaned value

    The condition uses the cleaned line, but the output still contains the original trailing newline.

    Fix: Print clean_line so the value tested and displayed has the same cleaned form.

Practice the Pattern

EASY

Write a loop that processes each line in file and displays only lines beginning with Subject:. Remove trailing whitespace before checking and printing the line.

Hints
  • Create a cleaned value by calling rstrip() on each line.
  • Use startswith("Subject:") as the conditional test.
  • Pass the cleaned value to print().
python

Key Takeaways

  1. startswith() checks only whether a string begins with a specified prefix and returns True or False.
  2. A line read from a file includes a trailing newline, while print() adds another newline, which can create double spacing.
  3. rstrip() removes trailing whitespace and newline characters from the end of a string.
  4. A reliable filtering pattern is to call rstrip(), test the result with startswith(), and print the cleaned line when the test is True.
  5. Prefix-based filtering is useful for structured text such as email-style files and log files.

Key Takeaways

  • Use startswith() to recognize lines by their beginning, not by text appearing anywhere.
  • Remove a file line's trailing newline with rstrip() before displaying it.
  • Clean each line first, then check its prefix and print only matching lines.
  • This pattern supports practical filtering of structured email and log text files.